The Nebius Managed PostgreSQL team developed a solution for AI workloads using CloudNativePG and Kubernetes, resulting in improved backup and restore times by transitioning from Barman to WAL-G, reducing a 1.5 TB backup duration from over a day to two hours. Since its launch in 2025, the service supports hundreds of production clusters and manages tens of tebibytes of data. Key features include high availability, point-in-time recovery, automated failover, compressed and encrypted backups, rapid restoration, connection pooling, and strict synchronous replication for zero data loss.
Kubernetes and CloudNativePG were chosen for their advantages, including out-of-the-box features and open-source distribution. Each cloud project has a separate MK8s cluster for customer isolation and security. Persistent data is stored on fault-tolerant SSD drives, enabling swift instance migration during failures. Regular physical backups are archived, allowing restoration within a seven-day window with a single click.
The transition from Barman to WAL-G addressed slow backup and restoration times and resource control limitations. With WAL-G, backups and restorations for a 1.5 TB cluster improved to two hours and one hour, respectively. PgBouncer was implemented to manage database connections effectively, reducing connection storms during peak AI workload traffic.
Two vector search extensions, pgvector and pgvectorscale, are available to cater to different memory and dataset needs. Benchmarking showed that Nebius performs competitively with AWS at a lower cost. The infrastructure integrates cloud engineering principles with specialized expertise to meet AI workload demands.