Amazon Web Services Inc. has enhanced its Aurora PostgreSQL database management system to allow users to query the Apache Iceberg data lake directly. This feature integrates the DuckDB analytical engine, enabling seamless access to live transactions and historical records without data duplication or complex ETL processes. Customers can use existing PostgreSQL applications to query data in Amazon S3, including S3 Tables. The integration aims to reduce the engineering workload associated with data pipelines and supports real-time applications.
DuckDB processes analytical scans within Aurora, eliminating extra network hops and allowing single queries to access both live and archived data. The feature supports external catalogs compliant with the Iceberg REST Catalog specification and can create foreign tables referencing data across multiple catalogs. Aurora optimizes query performance by filtering records and caching frequently accessed data.
Customers can combine recent transactions in Aurora with historical data from S3 without manual schema definitions. For low-latency workloads, selected data can be copied into native Aurora tables using standard SQL. This capability requires the aurora_analytics extension and an AWS IAM role for S3 and Glue access, and is compatible with Aurora PostgreSQL versions 17.11 and 18.6. The feature is available across all AWS regions at no additional charge, with costs incurred only for Aurora computing resources and S3 requests.