How and why we built Managed PostgreSQL for AI workloads

The Nebius Managed PostgreSQL team has crafted a solution tailored specifically for AI workloads, leveraging the capabilities of CloudNativePG and Kubernetes. This innovative approach has led to significant improvements in backup and restore times, particularly through the transition from Barman to WAL-G, which reduced a 1.5 TB backup duration from over a day to just two hours. Additionally, features such as PgBouncer, point-in-time recovery, pgvector, and pgvectorscale have been integrated to accommodate the unique demands of bursty AI workloads.

Since its general availability in 2025, Nebius has established a fully managed, one-click PostgreSQL solution that now supports hundreds of production clusters and manages tens of tebibytes of data. This article delves into the underlying architecture, the engineering choices made during development, and the optimizations implemented for AI workloads.

In designing the Nebius Managed PostgreSQL service, the goal was to merge the best practices from general cloud services with cutting-edge technologies tailored for AI applications. The service is utilized not only by external clients but also by various Nebius services, setting a high standard for performance and reliability. Key expectations include:

  • High availability
  • Point-in-time recovery (PITR)
  • Automated failover
  • Periodic physical backups that are compressed and encrypted
  • Rapid restoration capabilities
  • Connection pooling
  • Comprehensive metrics and logging
  • Zero data loss through strict synchronous replication

To meet these requirements, Kubernetes was chosen in conjunction with CloudNativePG (CNPG), an open-source operator for PostgreSQL. Despite its relative newness, CNPG offered several advantages:

  • Development by the experienced EnterpriseDB team
  • Provision of most required features out of the box
  • Open-source distribution under the Apache License 2.0

After a thorough evaluation of available operators, the decision to adopt CNPG has proven to be beneficial, as it has matured into the most popular Kubernetes operator for PostgreSQL.

Each cloud project is allocated a separate MK8s cluster, allowing multiple instances of Managed PostgreSQL and other Nebius cloud services to operate within the same Kubernetes environment while maintaining customer isolation and security. This architecture ensures fair resource utilization, with each PostgreSQL instance assigned to a distinct compute node within the MK8s cluster.

Backup Innovations

Persistent data is stored on Nebius network SSD drives, facilitating swift migration of cluster instances in the event of compute instance failures. These drives are designed for fault tolerance, ensuring they can withstand physical device failures.

Regular physical backups are conducted for each cluster, with WAL segments continuously archived to object storage. Customers can restore their entire cluster to any point within a seven-day recovery window with a single click in the console. This rapid backup and restore capability is crucial for disaster recovery and experimentation in AI applications, enabling teams to quickly test new datasets and models while ensuring safe rollback options.

From Barman to WAL-G

Initially, Barman was employed for backups due to its integration with the CNPG operator. However, two significant limitations were identified:

  1. Slow backup and restoration: Barman’s reliance on pg_basebackup, which operates in a single-threaded manner, resulted in prolonged backup and restoration times for large clusters. For instance, a Barman backup could exceed a day and restoration could take over 15 hours, which is impractical for teams needing rapid access to production-like data.
  2. Resource control limitations: The need for precise control over resource allocation during backup processes was not adequately addressed by CNPG, leading to competition for disk and network bandwidth that slowed down queries.

To overcome these challenges, the team transitioned to WAL-G for backup and restoration, achieving remarkable results. A 1.5 TB cluster that previously required over a day for backup and more than 20 hours for restoration can now be backed up in just two hours and restored in one hour, allowing for smoother operations and peace of mind for the SRE team.

AI workloads often generate unpredictable and bursty database traffic, with various processes accessing PostgreSQL simultaneously. Without a connection pooler, each AI worker or service replica may attempt to open its own database connections, leading to connection storms that hinder query execution. By implementing PgBouncer as a connection pooler, the team effectively manages database connections, reducing churn and allowing PostgreSQL to focus on executing queries efficiently.

Customers can select their preferred connection mode—session or statement pooling—and connection pooler pods are allocated on the same nodes as PostgreSQL pods to minimize latency and potential disruptions during compute node maintenance.

Two vector search extensions, pgvector and pgvectorscale, are readily available in every cluster, catering to different memory and dataset requirements:

  • pgvector: Offers memory-optimized data structures and index types suitable for datasets that fit within RAM, providing fast and efficient performance.
  • pgvectorscale: Designed for lower RAM usage and large SSD-backed datasets, this extension optimizes for disk and cache efficiency, albeit with a trade-off in throughput and recall performance.

Benchmarking conducted using VectorDBBench on Nebius and a comparable AWS instance demonstrated competitive performance, with Nebius delivering similar query per second (QPS) rates at a significantly lower cost. This reinforces Nebius’s commitment to providing a robust and cost-effective solution for AI workloads.

At Nebius, the integration of general cloud engineering principles with specialized expertise allows for a tailored platform that meets the demands of AI workloads. The infrastructure supporting the Managed PostgreSQL service exemplifies this synergy, ensuring that features like PITR and connection pooling are not just standard practices but essential tools for enhancing experimentation and operational efficiency.

To experience the capabilities of Nebius Managed PostgreSQL, users are encouraged to spin up a cluster, load their embeddings, and run the provided benchmarks.

Tech Optimizer
How and why we built Managed PostgreSQL for AI workloads