Lakebase is a fully managed Postgres database designed for modern application development, featuring an architecture that separates storage and compute with a serverless compute layer. This design allows for cost efficiencies through various mechanisms:
1. **Branching**: Developers can create isolated environments for development, testing, or experimentation without duplicating storage costs, as branches share the same underlying storage.
2. **Autoscaling**: Lakebase adjusts compute resources based on activity levels, allowing users to pay only for the compute they use. It can scale down during low activity and can be suspended entirely after inactivity, reducing costs to zero.
3. **Read Replicas and High Availability**: The separation of storage and compute allows for adding read replicas and high availability without incurring additional storage costs, as these instances utilize the same storage layer.
4. **Synced Tables**: Integration with the Databricks Intelligence Platform enables efficient data syncing, allowing users to sync only the necessary working set of data, which helps avoid unnecessary storage and sync costs.
5. **Sync Modes**: There are three sync modes (Snapshot, Triggered, Continuous) for transferring data from the Lakehouse to Lakebase, each with different cost implications based on data freshness requirements.
6. **Right-sizing Compute**: Users can set the initial compute range during project provisioning to avoid over-provisioning and unnecessary costs.
7. **Monitoring and Metrics**: The Lakebase Metrics dashboard provides insights into working set size and cache utilization, helping users optimize their compute sizing and performance.
8. **Point-in-Time Restore (PITR) and Snapshots**: PITR maintains history for recovery within a configurable window, while snapshots offer discrete recovery points. Both are priced lower than standard storage, making them cost-effective options for data recovery.
Lakebase costs are categorized into compute, storage, and serverless pipeline compute for syncing data. Compute is measured by CU usage, while storage includes branch storage, PITR history, and snapshot storage, all tracked separately. Users can access detailed usage and cost information through system billing tables.