Adding DMS Change Capture Without Letting an Abandoned Slot Fill the Source Database
After deploying the PostgreSQL prerequisites for DMS change capture, I still couldn’t answer one question: were changes reaching S3? Meanwhile, a stalled replication consumer threatened to fill the staging database’s 100 GB disk. There was no storage autoscaling to buy time. Protecting that source and proving delivery were independent obligations. I had configured a safeguard for the first; I still owed a test of the second.
TL;DR
DMS change capture can leave PostgreSQL retaining write-ahead log when a replication consumer stops advancing. On this 100 GB staging instance, the design used a dedicated parameter group with logical replication enabled and max_slot_wal_keep_size set to 10240 MB. That limit prioritizes source availability over indefinite replication recovery, while delivery to S3 requires a separate end-to-end check.
I was using AWS Database Migration Service to copy staging changes into object storage. I gave the database a dedicated parameter group, enabled logical replication, and set a roughly 10 GB allowance for slot-retained write-ahead log. None of those changes proved that a record had arrived downstream.
I used the following distinctions to judge what each check would establish. These are evidence requirements, not a list of completed tests:
| Evidence | What it establishes | What remains unproven |
|---|---|---|
| Runtime database settings | Logical replication is enabled | The task can read selected tables |
| Successful endpoint connection | Connectivity and authentication work | A selected change reaches the destination |
| Advancing replication position | The consumer is progressing | A particular change is visible in S3 |
| A matching target record | That committed change reached S3 | Continued operation during future failures |
What I was willing to retain for DMS change capture
I wanted changes available downstream without giving an unattended staging task an unlimited claim on database storage. Letting the application database run out of disk to preserve an export pipeline’s recovery position is a poor exchange. This was a preventive change; I was not recovering from a documented disk incident.
PostgreSQL records changes in its write-ahead log, usually called WAL. Logical decoding interprets those records for a replication consumer. A replication slot tracks the position the consumer still needs. That allows PostgreSQL to retain relevant WAL while the consumer catches up.
Stopping the consumer does not remove its slot. If writes continue while the slot remains, the database generates more WAL. It also preserves older records for the stalled reader. PostgreSQL explicitly warns that replication slots can retain enough WAL to fill the disk. PostgreSQL replication guidance
I set max_slot_wal_keep_size to 10240 MB. The setting limits slot-driven WAL retention at checkpoint time, not total WAL directory size. It reserves no space for application data. Once a stalled consumer falls beyond the allowance, required WAL can be removed at a checkpoint. The consumer then cannot resume from its old position. Restarting the task does not bring that history back. PostgreSQL replication settings
I accepted the prospect of rebuilding downstream state rather than retaining an abandoned consumer’s history indefinitely. Once that history is removed, recovery means establishing a fresh replication starting point. Where needed, it also means reloading affected data to restore downstream state.
Warning
A 10 GB slot retention setting does not guarantee 90 GB remains available on a 100 GB instance. Existing data, other WAL requirements, and checkpoint timing still matter. Free storage needs its own monitoring.
The 10 GB allowance was provisional. I had not measured WAL generation rates in this work. I therefore had no basis for treating it as a particular number of hours of recoverable downtime. As hypothetical arithmetic, 10 GB at roughly 1 GB per hour suggests ten hours of history; at 5 GB per hour, two.
To justify the setting, I would measure peak WAL generation over representative busy periods. That would tell me how much history a tolerated consumer outage needs. I would compare that with available storage, allowing for data growth and other WAL requirements.
I would also measure how long a downstream rebuild takes. If rebuilding takes longer than the downstream service can tolerate, a small retention allowance carries a cost. Addressing that cost would require more storage, faster intervention, or a different recovery plan. The source’s spare disk and the acceptable rebuild time both constrain the decision.
Unlimited retention would buy a longer recovery opportunity with source disk space. More storage would add time without resolving the abandoned-consumer problem. With no storage autoscaling enabled here, I chose a finite allowance and left its sizing subject to measurement.
What the running database needs to confirm
The AWS-managed default parameter group could not carry the required changes. I used a dedicated PostgreSQL 17 parameter group so the staging instance’s replication settings were explicit and reviewable.
These anonymized commands express the configuration, not a complete infrastructure deployment. They assume an existing RDS instance named stage-app-db and AWS credentials authorized to manage it:
aws rds create-db-parameter-group \
--region us-west-2 \
--db-parameter-group-name stage-app-pg17 \
--db-parameter-group-family postgres17 \
--description 'Staging logical replication with bounded WAL retention'
aws rds modify-db-parameter-group \
--region us-west-2 \
--db-parameter-group-name stage-app-pg17 \
--parameters \
'ParameterName=rds.logical_replication,ParameterValue=1,ApplyMethod=pending-reboot' \
'ParameterName=max_slot_wal_keep_size,ParameterValue=10240,ApplyMethod=pending-reboot'
aws rds modify-db-instance \
--region us-west-2 \
--db-instance-identifier stage-app-db \
--db-parameter-group-name stage-app-pg17rds.logical_replication is static and requires a reboot. A successful configuration API response does not establish that logical decoding is active. The interruption belongs in the deployment plan. Once the group association is ready, reboot and wait for the instance: AWS PostgreSQL source configuration
aws rds reboot-db-instance \
--region us-west-2 \
--db-instance-identifier stage-app-db
aws rds wait db-instance-available \
--region us-west-2 \
--db-instance-identifier stage-app-dbThen check the effective settings through an administrative SQL connection:
SHOW rds.logical_replication;
SHOW wal_level;
SHOW max_slot_wal_keep_size;I expect logical replication enabled, wal_level set to logical, and a retention value equivalent to 10 GB. The running database’s answer establishes the first row of the evidence table; the intended parameter values do not. That gap between configuration and observed behavior also matters in Why Configured Telemetry Does Not Prove Production Tracing.
What a connection leaves unproven
The setup created a dedicated replication login, retrieved its password from Secrets Manager, and granted replication and schema access. Table-level SELECT grants were started, but I do not have the completed table list to verify their coverage. I cannot claim the permissions were complete.
Replication access permits reading the change stream; table access permits reading selected source data. DMS also has supporting database object and setup requirements. Those depend on the configuration, including use of a non-master account. A pair of grants is not a universal installation recipe. AWS PostgreSQL source prerequisites
The infrastructure work included the standard dms-vpc-role and its managed VPC policy. That role supports DMS networking. The S3 target needs its own appropriately configured write access. Even a successful endpoint connection would establish only connectivity and authentication. Delivery of a selected change would remain unproven. AWS PostgreSQL-to-S3 walkthrough
What the slot tells me
For the evidence table’s third row, I need to inspect progress over time. An active connection alone says little about whether the consumer is advancing. This PostgreSQL 17 query exposes slot state and estimates the WAL span it requires:
SELECT
slot_name,
active,
restart_lsn,
confirmed_flush_lsn,
wal_status,
safe_wal_size,
pg_size_pretty(
pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)
) AS retained_wal_span
FROM pg_replication_slots
WHERE slot_type = 'logical'
ORDER BY slot_name;restart_lsn identifies the oldest WAL the slot may still need. confirmed_flush_lsn records the consumer’s acknowledged position. The span between the current WAL position and restart_lsn is a retention indicator, not a measurement of total disk usage. Retention and acknowledgment positions answer different questions.
wal_status reports the state of the slot’s required WAL, including whether it has been lost. safe_wal_size estimates how much more WAL can be written before the slot risks becoming lost. A null value is possible, including for an already lost slot. PostgreSQL slot monitoring reference
I would monitor these alongside source free storage so there is time to intervene before required history disappears. Progress here still does not establish that a particular change is visible in S3.
The record I still needed to find
I had deployed logical-replication prerequisites and done DMS setup work. I had not established that a specific committed change reached S3. The answer to my opening question remained unverified.
My acceptance check would use an already selected table and a disposable staging record. I would commit a recognizable insert, record its key, and find the matching target record. I would repeat the check with an update and delete if those operations belong in the intended feed. For each, I would check its representation against the configured target format.
A fresh object timestamp is weaker evidence because it might reflect another table or an initial load. The useful proof connects a known committed source operation to its target record. That check needs to allow enough time for configured output batching. DMS S3 output supports file-size and time thresholds, so source consumption and object visibility need not happen together. AWS S3 target settings
Until I could trace that committed operation into the target output, I would report the safeguard as configured and delivery as unverified.
FAQ
Why can a stopped DMS task fill PostgreSQL storage?
A replication slot can remain while its consumer stops advancing, causing PostgreSQL to retain WAL needed for recovery. Without a finite retention allowance, that history can grow until storage becomes a problem.
Does max_slot_wal_keep_size strictly cap total WAL usage?
No. It limits slot-driven retention at checkpoint time; other WAL requirements and checkpoint timing still affect disk usage.
Does enabling rds.logical_replication require a reboot?
Yes, it is a static setting. Associate the custom parameter group, schedule the reboot, and verify the effective settings through SQL afterward.
How do I prove PostgreSQL changes reached S3?
Commit a recognizable change in a selected table and find the matching record in the target output, allowing for configured batching. Infrastructure deployment establishes readiness; an observed destination record establishes delivery.
