
CDC Crash Test
A reproducible lab for a quiet PostgreSQL CDC failure that filled a disk while Debezium reported healthy.
I rebuilt a production outage in a local lab and tested three heartbeat configurations. Without an effective heartbeat, PostgreSQL filled its disk while Debezium reported healthy in every sample. A heartbeat that wrote to a published table kept the same workload below the cap. Two independently reset pairs reproduced the result.
The incident behind the lab
A staging PostgreSQL database stopped accepting connections after write-ahead log (WAL) retained by a stalled Debezium replication slot filled its disk.
Other environments ran the same CDC setup without this failure. Their captured tables were busy enough to keep the slot moving. Staging was quiet, although other database writes continued. The setup lacked an effective heartbeat for that condition.
Read the incident analysis. The lab asks which heartbeat configuration would have prevented the failure, and which signal would have warned us early.
Recreate the quiet failure
PostgreSQL 17 publishes changes from orders and cdc_heartbeat to Debezium Server 3.6, which sends events to a local HTTP receiver. A third table, noise, generates WAL but is deliberately excluded from the publication. Writing only to noise makes the database busy while the captured tables stay quiet.
orders carries application changes; cdc_heartbeat moves only when an action query updates it.noise produces WAL that Debezium does not see as captured row changes.Everything runs locally in Docker with images pinned by digest. Each scenario starts from a reset database, replication slot, and connector offsets so inherited backlog cannot distort the comparison.
Healthy until the disk filled
I capped PostgreSQL's data directory at 256 MiB and generated 8 MiB of unpublished writes every 10 seconds, for up to 30 batches. Debezium stayed running in both scenarios.
| Measure | No heartbeat | Published-table heartbeat |
|---|---|---|
| Batches completed | 23 of 30 | 30 of 30 |
| Outcome | Disk full on batch 24 | Stayed below the cap |
| Slot position | Never moved | Advanced |
| Peak disk use | 100% | About 68% |
| Debezium health | UP in every sample | UP in every sample |
The slot was active, but its confirmed position never moved without the published-table heartbeat. A second independently reset pair reproduced the same outcome.
A timer alone did not move the slot
I compared three heartbeat configurations on freshly reset databases, each over 10 minutes of busy, unpublished writes. The timer alone did not advance the slot in this setup. It moved only when heartbeat.action.query updated a table in the publication.
| Configuration | Slot position | Retained-WAL growth |
|---|---|---|
| Disabled | Never moved | +1.63 MB |
| 10-second timer only | Never moved | +1.92 MB |
| Timer plus published-table action query | Advanced | +0.33 MB |
Limit: the heartbeat bounded the backlog but did not eliminate it. Retained WAL peaked at 99 MB and 108 MB across the two disk-fill runs, around 40% of the disk. A faster write rate or longer interval could raise that peak.
Keep the experiment safe and honest
The database runs on a dedicated 256 MiB in-memory filesystem, never a host directory. The runner checks the cap before generating load and requires a typed YES. Cleanup removes the temporary filesystem and connector offsets while preserving the CSV, summary, and PostgreSQL log.
Early runs exposed two measurement traps. A back-to-back heartbeat comparison inherited WAL backlog, so I reset the database, slot, and offsets before each scenario. Another repeat had a 9-minute-38-second sampling gap; the runner marked it inconclusive, and I reran it.
One first-run health sample started roughly 84 ms after PostgreSQL logged its disk-full PANIC and still read UP. The repeat's final sample was 1.5 seconds before the crash, so a post-crash UP reading is an observation from one run, not a repeated result. Stopping Debezium was also a different failure: its health failed within seconds.
Operational lessons and limits
For this workload, configure heartbeat.action.query to update a table in the connector's publication. Alert on retained WAL per slot from pg_replication_slots and make sure disk alerts reach someone who can act. max_slot_wal_keep_size can protect the database as a last line of defense, at the cost of possible CDC gaps.
The lab uses self-hosted PostgreSQL 17.11, Debezium Server 3.6.3.Final, and a non-durable HTTP receiver. It does not test managed databases, Kafka Connect, failover, other CDC tools, production-scale timing, or whether every event reaches a durable destination. It isolates one connector and does not reproduce every contributing condition of the original incident.
Tradeoff: the small disk cap makes failure reproducible in minutes. The useful result is the shape of the failure and the slot's behavior, not the exact time or byte count at production scale.