← Project index
CDC Crash Test mark showing data flow reaching a stop

CDC Crash Test

A reproducible lab for a quiet PostgreSQL CDC failure that filled a disk while Debezium reported healthy.

Outcome

I rebuilt a production outage in a local lab and tested three heartbeat configurations. Without an effective heartbeat, PostgreSQL filled its disk while Debezium reported healthy in every sample. A heartbeat that wrote to a published table kept the same workload below the cap. Two independently reset pairs reproduced the result.

The incident behind the lab

A staging PostgreSQL database stopped accepting connections after write-ahead log (WAL) retained by a stalled Debezium replication slot filled its disk.

Other environments ran the same CDC setup without this failure. Their captured tables were busy enough to keep the slot moving. Staging was quiet, although other database writes continued. The setup lacked an effective heartbeat for that condition.

Read the incident analysis. The lab asks which heartbeat configuration would have prevented the failure, and which signal would have warned us early.

Recreate the quiet failure

PostgreSQL 17 publishes changes from orders and cdc_heartbeat to Debezium Server 3.6, which sends events to a local HTTP receiver. A third table, noise, generates WAL but is deliberately excluded from the publication. Writing only to noise makes the database busy while the captured tables stay quiet.

Published tablesorders carries application changes; cdc_heartbeat moves only when an action query updates it.
Unpublished loadnoise produces WAL that Debezium does not see as captured row changes.
ObserverA Go process samples replication-slot positions, retained WAL, WAL directory size, checkpoints, and Debezium health, then writes CSV and Markdown summaries.

Everything runs locally in Docker with images pinned by digest. Each scenario starts from a reset database, replication slot, and connector offsets so inherited backlog cannot distort the comparison.

Healthy until the disk filled

I capped PostgreSQL's data directory at 256 MiB and generated 8 MiB of unpublished writes every 10 seconds, for up to 30 batches. Debezium stayed running in both scenarios.

Paired disk-fill runs under the same workload
MeasureNo heartbeatPublished-table heartbeat
Batches completed23 of 3030 of 30
OutcomeDisk full on batch 24Stayed below the cap
Slot positionNever movedAdvanced
Peak disk use100%About 68%
Debezium healthUP in every sampleUP in every sample

The slot was active, but its confirmed position never moved without the published-table heartbeat. A second independently reset pair reproduced the same outcome.

A timer alone did not move the slot

I compared three heartbeat configurations on freshly reset databases, each over 10 minutes of busy, unpublished writes. The timer alone did not advance the slot in this setup. It moved only when heartbeat.action.query updated a table in the publication.

Heartbeat comparison; a repeat gave the same result
ConfigurationSlot positionRetained-WAL growth
DisabledNever moved+1.63 MB
10-second timer onlyNever moved+1.92 MB
Timer plus published-table action queryAdvanced+0.33 MB

Limit: the heartbeat bounded the backlog but did not eliminate it. Retained WAL peaked at 99 MB and 108 MB across the two disk-fill runs, around 40% of the disk. A faster write rate or longer interval could raise that peak.

Keep the experiment safe and honest

The database runs on a dedicated 256 MiB in-memory filesystem, never a host directory. The runner checks the cap before generating load and requires a typed YES. Cleanup removes the temporary filesystem and connector offsets while preserving the CSV, summary, and PostgreSQL log.

Early runs exposed two measurement traps. A back-to-back heartbeat comparison inherited WAL backlog, so I reset the database, slot, and offsets before each scenario. Another repeat had a 9-minute-38-second sampling gap; the runner marked it inconclusive, and I reran it.

One first-run health sample started roughly 84 ms after PostgreSQL logged its disk-full PANIC and still read UP. The repeat's final sample was 1.5 seconds before the crash, so a post-crash UP reading is an observation from one run, not a repeated result. Stopping Debezium was also a different failure: its health failed within seconds.

Operational lessons and limits

For this workload, configure heartbeat.action.query to update a table in the connector's publication. Alert on retained WAL per slot from pg_replication_slots and make sure disk alerts reach someone who can act. max_slot_wal_keep_size can protect the database as a last line of defense, at the cost of possible CDC gaps.

The lab uses self-hosted PostgreSQL 17.11, Debezium Server 3.6.3.Final, and a non-durable HTTP receiver. It does not test managed databases, Kafka Connect, failover, other CDC tools, production-scale timing, or whether every event reaches a durable destination. It isolates one connector and does not reproduce every contributing condition of the original incident.

Tradeoff: the small disk cap makes failure reproducible in minutes. The useful result is the shape of the failure and the slot's behavior, not the exact time or byte count at production scale.

Evidence in the repository