On this page
Key takeaways
- Pin versions, run as non-root, and keep configuration and secrets outside the image.
- Health checks, structured logs, and a few well-chosen metrics catch most problems early.
- Practise restores and upgrades in staging; a backup is a hope until it has been restored.
- Keep migrations small, reversible where possible, and separate from the application's runtime role.
Self-hosting a service means you own its operation. For a container plus PostgreSQL, that is a manageable job with a small set of habits. These tips are what we hand to teams deploying StablePay-style software, and they apply to most containerised web services with a relational database.
Images and containers
Small decisions in the image pay off for years.
1. Pin image versions, never `latest`
A moving tag deploys different code without a decision. Pin a version and update deliberately, ideally by digest [1].
2. Run as a non-root user
Add USER in the Dockerfile and drop capabilities you do not need.
3. Make the root filesystem read-only where you can
Mount a tmpfs for scratch space. A compromised process then cannot alter the image at runtime.
4. Add a real health check
Check that the app can serve requests and reach its database, not merely that the process exists.
healthcheck:
test: ["CMD", "node", "-e", "fetch('http://localhost:3000/healthz/ready').then(r=>process.exit(r.ok?0:1)).catch(()=>process.exit(1))"]
interval: 30s
timeout: 3s
retries: 3
start_period: 20s5. Set resource limits
Memory and CPU limits stop a runaway process from taking neighbours down with it.
6. Keep the image build reproducible
Lockfiles, a frozen install, and a multi-stage build make deployments predictable.
Configuration and secrets
Configuration belongs in the environment, and secrets belong in a store [2].
7. Never bake secrets into images
Images are copied, cached, and pushed. Mount secrets at runtime from a secret store or files.
8. Validate configuration at startup
Fail fast with a clear message when a required setting is missing or malformed, rather than failing mysteriously later.
9. Separate configuration per environment
Staging and production must not share credentials, databases, or webhook secrets.
10. Log the version and config hash at startup
When debugging, knowing exactly what is running is half the answer.
Migrations and upgrades
Schema changes are where deployments hurt.
11. Run migrations with a separate role
The application's runtime role should not be able to alter schema. Use a dedicated migrator credential.
12. Keep migrations small and additive
Add columns and tables first; remove them in a later release after code stops using them. This lets old and new code coexist during a rolling deploy.
13. Avoid long locks
Adding an index on a large table can block writes. Use concurrent index creation where supported.
CREATE INDEX CONCURRENTLY payments_open_by_age
ON payments (created_at)
WHERE state IN ('created','seen','confirming','underpaid');14. Rehearse upgrades in staging
Apply the same patch release to a copy of production data first. Time it and note anything surprising.
15. Have a rollback story
Know whether you can roll back the image alone, or whether the migration makes that unsafe.
Backups and recovery
A backup you have never restored is a hope [3].
16. Automate backups and alert when they fail
A silent failure is discovered when you need the backup.
17. Use continuous archiving for point-in-time recovery
Base backups plus archived write-ahead logs let you rebuild to a moment just before a bad change [4].
18. Store backups elsewhere, encrypted
A different account or region, with separate credentials, so one compromise cannot destroy both.
19. Restore into a scratch environment on a schedule
Time it, verify the data, and run your reconciliation checks against the restored copy.
20. Write down the recovery procedure and who decides
During an incident you do not want to invent it.
Observability and capacity
Know it is unwell before your users tell you.
21. Log as structured events to stdout
One JSON object per line, with a request or event id. Let the platform collect and ship the stream [2].
22. Watch a handful of database signals
Connection count, replication lag if any, slow queries, and disk growth [6]. Alert on trends, not single spikes.
SELECT pid, now() - query_start AS running_for, state, left(query, 120) AS query
FROM pg_stat_activity
WHERE state <> 'idle'
ORDER BY running_for DESC
LIMIT 5;23. Keep vacuum healthy
Autovacuum normally handles it; watch for tables with heavy churn and bloat, and tune deliberately [5].
24. Track disk growth and plan capacity
Append-only ledgers grow steadily. Project when you will need more space, and act early.
25. Alert on expiring certificates and credentials
Preventable outages are the most embarrassing kind.
Monthly operations routine
- Confirm backups ran and restore one into scratch.
- Review image and dependency scan results; apply pending patch releases in staging first.
- Check disk growth, slow queries, and connection usage.
- Verify certificate and secret expiry dates.
- Run one drill from the runbook.
- Review alerts: silence noisy ones, add any that were missing in the last incident.
None of this is glamorous. That is the goal: when the containerised service and its database are boring, you get to spend your attention on the product.
References & further reading
- 1Docker Security Cheat Sheet — OWASP Cheat Sheet Series
- 2The Twelve-Factor App — Adam Wiggins, 12factor.netConfig, logs, and disposability.
- 3PostgreSQL Documentation: Backup and Restore — PostgreSQL Global Development Group
- 4PostgreSQL Documentation: Continuous Archiving and Point-in-Time Recovery (PITR) — PostgreSQL Global Development Group
- 5PostgreSQL Documentation: Routine Vacuuming — PostgreSQL Global Development Group
- 6PostgreSQL Documentation: Monitoring Database Activity — PostgreSQL Global Development Group
The product behind this post
StablePay
Self-hosted stablecoin payment infrastructure.
From $4,800 one-time license
About the author
Product behaviour described here reflects what is implemented and tested; anything else is marked as planned. Code samples are illustrative.
All writing