From Live Transactions to In-Database Analytics and Machine Learning
Skip the CDC and ETL step between a transaction and your analysis.
Getting a live transaction in front of an analytical query or a machine learning workload usually means extra steps: a separate CDC tool captures the change, a custom ETL job transforms and loads it, and by the time the data lands in the warehouse, the moment to act on it has often passed.
This session shows how to skip that path: a health insurer's claim, written to EDB Postgres Distributed (PGD), replicates into Apache Iceberg the instant it's written, and WarehousePG reads that same Iceberg data directly—no external CDC tool, no custom code. From there, WarehousePG runs SQL analytics against the full claims history and scores each member, using in-database machine learning, for their likelihood of having a non-preventative claim in the future—while the data is still fresh enough for a care team to act on. It's one instance of a bigger idea: transactional, analytical, and real-time workloads converging on a single, open copy of your data.
What you will learn:
- Why capturing, transforming, and loading a change into a warehouse doesn't have to be a separate, delayed step
- How EDB Postgres Distributed (PGD) replicates live transactional data into Apache Iceberg the moment it's written, no external tooling required
- How WarehousePG reads that same Iceberg data directly to run SQL analytics and in-database machine learning, without moving the data anywhere else
- How that same in-database machine learning scores members for their risk of a future non-preventative claim
- What it looks like when transactional, analytical, and real-time workloads all read from one open copy of data instead of three
- How the same architecture extends to real-time serving with ClickHouse, no new platform and no new copy of the data