← all posts

Moving a Production Database: Pressing a Button You Can't Undo

aws rds delete-db-instance --skip-final-snapshot ...

I stopped just before pressing Enter.

If this were code, I'd have pushed it. A bad deployment can be rolled back; a bug can get a hotfix. Most of our work is reversible, so ‘try it and fix it if it breaks’ works. This command was different. Press Enter, and the instance disappears for good.

The hard part of this production database migration wasn't moving the data. The dump took 26 seconds and the restore 0.75 seconds. The difficulty was the irreversible moments scattered through the process, like that final Enter. ‘Fix it if it breaks’ doesn't work there. You need reasons to believe the action is right before taking it.

139 ms: the reason to move

The application was in Seoul, but RDS was in Sydney, for assorted reasons. Before settling for ‘distance makes it slow,’ I measured it. The TCP round trip from Seoul to Sydney RDS was 139 ms, versus 0.2 ms in the same region. Every query started with that 139 ms cost.

The dump took 26 seconds for the same reason. pg_dump queried the catalogs 184 times to discover the schema, waiting for each answer before asking the next question. 184 × 139 ms = 25.7 seconds. Physical distance was the bottleneck; having only 12 MB of data didn't help. After the move, the same query round trips fell to 0.2–0.9 ms.

One minute: the price of downtime

A zero-downtime approach would have required PgBouncer between the app and DB, plus multiple application instances. We had one server with the app connected directly to the DB. Supporting that approach meant rebuilding the setup—a substantial project.

All that work would buy us a reduction from one minute of downtime to zero. The rehearsal gave us the breakdown: 26 seconds to dump, 0.75 to restore, and 30 to restart the app. I wasn't going to spend weeks eliminating one quiet minute overnight. With 12 GB, a dump alone could take tens of minutes and change the decision. We had 12 MB, so we chose to stop and copy.

836 MB: the memory budget

Would adding a DB overwhelm a small server with 1.9 GB of RAM already running a JVM application? I checked in a rehearsal. The app used 490 MB, the PostgreSQL container 33 MB, and 836 MB remained available. Since both fit, I kept the instance as it was.

PostgreSQL used only 33 MB because its working set was small: 12 MB of data. Even with shared_buffers set to 128 MB, it touched only a small number of pages. The data fit in memory, leaving few disk reads; the cache hit rate after migration was 99.5%. On this server, the JVM was the bigger memory concern, not PostgreSQL.

1.7 GB: room for swap

The server had no swap. Without it, a brief memory spike can trigger OOM, so I added a safety net. Disk space chose the size: with 1.7 GB free, a 2 GB swap file wouldn't fit. I made it 1 GB; afterward, reported free space was 1.4 GB at 80% utilization. On a small server, even a safety net consumes a scarce resource.

Cutover: one minute overnight

Night came. From this point on, irreversible actions would appear one by one, so I defined a gate for each step.

First, I stopped the app to freeze writes to RDS. Comparing two databases while both are changing doesn't work. I assumed the app was the only writer, but strictly speaking I should have checked pg_stat_activity for zero active connections. Stopping the app and assuming nobody else was writing was the least comfortable part of this migration. After stopping it, the dump took 26 seconds, the restore 0.75, with no warnings.

I didn't validate the restore using row counts. Two tables can both have 35 rows and still contain different values. Instead, I hashed all concatenated rows and compared the two databases.

            New container       RDS
users       8479836e179a    =    8479836e179a
blog_posts  5b2b243ef8ac    =    5b2b243ef8ac

I checked that both sides used pg_trgm 1.6. With the big-bang approach, sequences came with the dump, preserving users_id_seq=1649. With logical replication, that would have needed manual adjustment.

Once the hashes matched, I changed DB_HOST to the container and restarted the app. It's configuration, not code, so a restart was enough—but there's a trap. Without also updating Secrets Manager, the next deployment would overwrite .env with the old value and point at the deleted database. I ran ANALYZE because statistics were empty after restoration. Health was UP, there were 11 connections, and reads worked. One minute had passed since stopping the app.

The new database was now authoritative. We'd crossed one point of no return.

Deletion, and hesitation

I returned to the delete command. I could run it because every earlier gate was green: matching hashes, health UP, and successful reads. If any had failed, I'd have stopped and pointed DB_HOST back to RDS. I'd also taken a manual snapshot to leave a recovery path. That's why I could use --skip-final-snapshot.

status: deleting

The gates passed, and a snapshot existed. There wasn't much left to lose. Still, my finger paused. delete-db-instance works in one direction: that instance doesn't come back. Restoring a snapshot creates a new instance; it's a different operation. This hesitation was almost reflexive. But checking gates before small irreversible actions builds a habit that can survive much larger deletions. What got my hand moving again wasn't courage. It was those gates.

After the move

Only after leaving RDS did I understand what the bill had bought: automated backups, monitoring, and recovery. The snapshot was a picture of the old RDS instance, not a backup of the new container database. I now owned all of that work. I'd configured Discord alerts below 100 MB of free memory, but shipping pg_dump backups to S3 was still pending. Even then, a backup deserves that name only after a restore test. Leaving a managed service wasn't simply saving money. It traded money for labor and risk.

Debt remained: self-managed backups and monitoring, and a single server whose failure would take both the app and DB down. The 8 GB disk was already 90% full.

Moving the data took 27 seconds. Everything else was about crossing irreversible boundaries. 139 ms justified the move; one minute justified big-bang cutover; 836 MB justified keeping the instance; 1.7 GB determined swap size; matching hashes justified pressing Enter. Without those numbers, the decisions would have rested on intuition. That matters even more when the work can't be undone.

More work lacks git revert than we think: deleting a production table, migrating data, sending a payment that can't be recalled. Pausing before those commands isn't cowardice. Use the pause to check that every reason to proceed is green.