How Matia Made Postgres CDC Faster, Without Making Your Database Pay for It

Matia now streams the Postgres WAL over the replication protocol: up to 4x faster CDC on Aurora, with less load on your production database.
Benjamin Segal, Co-Founder & CEO

If you run change data capture (CDC) on Postgres, you already know the tradeoff everyone quietly accepts: real-time sync is great until it starts competing with your production database for memory, I/O, and CPU. The faster you pull changes out of the Write-Ahead Log (WAL), the more pressure you put back on the database your app depends on.

We shipped a change to how Matia's Postgres connector reads the WAL that breaks that tradeoff. It's faster, a lot faster, but the bigger story is that it's lighter. It asks less of your database while doing more work. Here's the plain version of what changed and why it matters.

A Quick Primer: How CDC actually reads postgres

Postgres doesn't hand you a live feed of every insert, update, and delete out of the box. Every change first lands in the WAL, an append-only log Postgres already uses internally for durability and crash recovery. CDC tools like Matia "tail" that log, translate each entry into a structured event, and ship it downstream to your warehouse, lake, or wherever the data needs to go.

There's more than one way to read that log, and the method you choose has a big effect on both speed and how much load you put on the source database.

The Old Way: Querying the WAL through a cursor

The traditional, widely used approach treats the WAL like a big table you can query. It repeatedly asks Postgres, "Give me the next chunk of changes," and Postgres opens a cursor to hand them back page by page.

It works, and it's the industry-standard approach for good reason: it's simple and predictable. But it has a structural cost. To serve each request, Postgres has to decode a meaningful portion of the log on its own CPU, hold that decoded data in memory, and manage the cursor's state for as long as you're reading. On a quiet database, that's barely noticeable. On a busy, high-write production database, it adds real memory pressure and I/O load to the same machine your application is using.

It can also get in the way of routine maintenance, like building a new index, while a long read is in progress. This is manageable. With careful attention to lock behavior, you can reduce how much a long-running read interferes with other database activity. But it's an extra layer of operational care that the cursor-based approach requires and the streaming approach doesn't.

The New Way: Streaming the WAL directly

Instead of repeatedly asking Postgres for the next batch, Matia's new approach opens a single continuous stream using Postgres's built-in replication protocol. That's the same core mechanism Postgres uses to replicate data to standby servers. Changes flow to Matia continuously as they happen, rather than being fetched in a request/response loop.

This is a fundamentally different relationship with the source database. Postgres decodes and forwards each change once, as part of a lightweight, ongoing conversation, rather than through a series of discrete, resource-intensive queries. The database isn't asked to materialize a large batch of decoded changes and hold it while a client slowly works through them. It just keeps handing off what's new.

We built it as a custom streaming plugin, and it's now available in the Matia Postgres connector.

The Results: Faster, and lighter on the database

We benchmarked both approaches in two environments, a local Postgres instance and Aurora (Postgres-compatible, cloud-managed), each processing change events.

On a local database:

  • Cursor-based reads: ~124,000 events/sec
  • Streaming: ~247,000 events/sec, roughly 2x faster

On Aurora, a more realistic, production-like environment:

  • Cursor-based reads: ~3,000 events/sec (about 5.5 minutes to process 1 million changes)
  • Streaming: ~12,800 events/sec (about 1.3 minutes for the same volume), over 4x faster

The gap widens in the cloud-managed environment because that's where the old approach's resource overhead hurts most. Network round trips and per-batch decoding compound each other under real-world conditions.

Why This Is More Than a Speed Story

It would be easy to stop at "4x faster" and call it a day. But the more meaningful shift is how that speed is achieved.

The cursor-based approach is fast the way a delivery truck making a hundred separate trips is fast. It gets there, but it burns fuel on every trip. Because it works by having Postgres decode and stage batches of changes for a client to page through, the database has to keep more decoded data in memory and manage more open state, for longer, no matter how quickly the client on the other end consumes it.

For large transactions or backfills, Postgres can end up holding a substantial chunk of decoded WAL in memory at once, just to serve the read.

The streaming approach never asks Postgres to do that. There's no large decoded batch sitting in memory waiting to be paged through. Changes are handed off as a continuous flow and don't need to be staged. That means:

  • Lower memory footprint on the database, even under sustained high-write load.
  • Reduced I/O overhead, because the read pattern lines up with how Postgres already produces WAL data, instead of layering extra query-and-fetch cycles on top.
  • No interference with concurrent database maintenance, like index builds, which the old approach could block during long-running reads.
  • Faster recovery from replication lag, so if your sync ever falls behind (a deploy, a network blip, a traffic spike), it catches up quickly without leaning harder on the database.

In other words, this isn't just the same job done quicker. It's the same job with a meaningfully smaller footprint on the production system you care most about protecting. For teams running Postgres as a live, business-critical database, not just a batch source, that distinction matters as much as raw throughput, if not more.

The Bigger Picture: Postgres CDC That Respects Your Database

This work is part of a broader investment in how Matia handles Postgres, historically one of the trickiest sources to replicate reliably at scale without disrupting the source system. We've seen customers cut full-database sync times from days to hours by parallelizing large syncs. This streaming change follows the same philosophy for ongoing, real-time CDC: move data fast, but never at the database's expense.

If you're running Postgres CDC today and watching replication lag or database load creep up under write-heavy workloads, this change was built for you.

Want to see how Matia's Postgres connector performs on your workload? Get in touch. We're happy to walk through the benchmarks in more detail.