01Services02Process03Projects04About05FAQ06Blog07Hire Me

25+ products shipped · $3.8M+ raised by clients

Back to Blog
InfrastructurePart 8 of 9September 28, 202615 min read

Backups you have actually restored

Subhankar Denria

Subhankar Denria

Software Architect · Product Engineer

~18 min

What this part does

If the disk dies, the server is deleted by mistake, or a bad update wrecks the data, rebuild the database as it was last night. Everything else on the server can be rebuilt from scripts in an hour. The database — people, their contacts, their history — can't.

Cost
$0 — up to 10 GB
Kept
30 nights, encrypted, on another continent
If you skip a step
You find out the backups don't work during the disaster

In post 4 I turned off Google's disk snapshots, because they're billed. That wasn't "no backups". It was "back up the one thing that matters, somewhere free, in a way I can prove works".

The design, in one breath

Every night at 02:30 UTC, the server dumps the database, encrypts it with a key only my laptop can open, and uploads it to Cloudflare R2. R2 deletes anything older than 30 days by itself, and until then a lock stops anyone erasing them — the server included. The alarm service from post 7 emails me if a night fails or is missed. And I have restored one of those backups, end to end, to prove the whole chain.

Every night at 02:30 UTC
  1. Postgres

    the database: people, contacts, history

  2. pg_dump

    a consistent copy, while the app keeps running

  3. age · encrypted

    with the public key only — the server can lock, never open

  4. Cloudflare R2 · Western Europe

    another company, another continent; locked for 29 days, deleted after 30

The private key lives in two places

the laptop · and a password manager

Lose both, and every backup is unreadable.

The rehearsal

decrypt → throwaway Postgres 18 in Docker → restore → count every table

✅ RESTORE WORKS

The real rehearsal restored all 21 tables and 14 migrations. They're empty only because nobody uses the API yet.

Cloudflare R2

What it is
File storage for servers
Why this one
10 GB free, and downloading is free. A backup you'd have to pay to restore isn't a backup. Reachable over IPv6. And it's a different company from Google — one account problem can't take out both the server and its backups

pg_dump

What it is
Postgres's own export tool
Why this one
Makes a consistent copy while the app keeps running

age

What it is
A small, modern encryption tool
Why this one
The server only gets the public key: it can lock backups but never open them

rclone

What it is
Uploads files to cloud storage
Why this one
Standard, in Ubuntu's own packages, speaks R2

systemd timer

What it is
Runs it nightly
Why this one
Same pattern as the every-minute jobs

Why encrypt, if R2 is private anyway?

Because of what's inside: names, phone numbers, and the fact that someone lives alone and is being looked out for. The server needs a key to upload to R2, and that key lives on the server. If the server were ever compromised, the attacker would have the key.

With encryption, what they'd get is locked files. The server holds only the public half of an age key — like a padlock anyone can snap shut but only the key-holder can open. The private half never touches the server.

Step 1 — Switch R2 on

Cloudflare dashboard (account level) → R2 Object Storage. The first time, it asks for a payment method — a card or PayPal — even for the free plan. (From India, your bank may send another e-mandate message, like Google's in post 2. Same answer: it's a ceiling, not a charge.)

Why it stays $0: the free tier is 10 GB of storage and a million uploads a month. Thirty nightly backups of a few MB each is well under 1% of that.

Step 2 — The bucket

A bucket is a top-level folder in R2. Create bucket:

Name

Value
lampsill-backups
Why
One bucket, one purpose, so the key can be limited to it. Permanent

Location

Value
Automatic, then Provide a location hint → Western Europe
Why
See below

Storage class

Value
Standard
Why
Infrequent Access isn't in the free tier

The location hint. Left alone, Cloudflare placed my bucket in Asia Pacific — because I was signing up from India. But the backups hold UK users' data. UK law already treats the EU as a safe place for personal data, so a European location keeps the list of countries short (for the privacy notice), and puts the backups on a different continent from the US server.

Don't choose "Specify jurisdiction" unless you need a legal guarantee: it changes the bucket's web address, and scripts written for the normal address stop working.

Step 3 — Keep 30 nights, and make them impossible to erase

In the bucket → Settings → Object lifecycle rules → Add rule:

Rule name

Value
delete-after-30-days

Prefix

Value
db/ (with the slash)

✅ Delete uploaded objects after

Value
30 Days

✅ Abort incomplete multipart uploads after

Value
7 Days

Transition to Infrequent Access

Value
unticked (not free)

Why R2 does the deleting, not the server: if the server did the deleting, a broken or hacked server could quietly delete every backup.

That only covers half of it, though. The server's key (step 4) can write to the bucket, and a key that can write can also overwrite: a hacked server could list every backup and replace each one with junk, and the lifecycle rule would faithfully keep thirty nights of nothing.

A bucket lock closes that gap. Same bucket → Settings → Bucket lock → Add rule:

Rule name

Value
keep-29-days

Prefix

Value
db/
Why
The same files the lifecycle rule looks after

Retention

Value
29 days
Why
For 29 days nothing — not the server, not a leaked key, not me by mistake — can delete or overwrite a backup

Why 29, not 30: a lock outranks a lifecycle rule. Ending the lock a day before the delete date means the two never argue over a file.

(Worth remembering for a privacy notice: someone who deletes their account still exists in the backups for up to 30 days.)

Step 4 — A key for the server, limited to one bucket

R2 overview → Manage API tokens → Create Account API token:

Token name

Value
backup-server
Why
So you know what it is in a year

Permissions

Value
Object Read & Write
Why
Can add and read files — and could overwrite them, which is what the lock in step 3 is for. Can't create, delete or reconfigure buckets

Specify bucket(s)

Value
Apply to specific buckets only → lampsill-backups
Why
A leaked key reaches this bucket and nothing else

TTL

Value
Forever
Why
An expiring key would silently stop the backups one night

→ Create. The next page shows an Access Key ID, a Secret Access Key and an S3 endpoint URL. The secret is shown once. Keep the tab open until step 7, and don't screenshot it.

Step 5 — The encryption key, on your laptop

bash
brew install age      # macOS; on Linux: apt install age
age-keygen -o ~/backup-key.txt && chmod 600 ~/backup-key.txt && grep 'public key' ~/backup-key.txt

It prints # public key: age1… — that's the only part the server gets.

Then save a second copy of the whole file in your password manager. This is the most important sentence in the post: if the private key is lost, every backup is unreadable. A dead laptop without that second copy means no backups at all.

How I leaked the key, and why it didn't matter

My first attempt: I opened the key file in a text editor to copy it into my password manager — and took a screenshot of the screen to show someone what I was doing. The private key (AGE-SECRET-KEY-…) was right there in the image.

It didn't matter, because nothing had been encrypted with it yet. I deleted it and made a new one. Replacing a key before first use costs nothing; after a month of backups, it would mean those backups were exposed.

This time I copied the file to the password manager without it ever appearing on screen:

bash
rm ~/backup-key.txt && age-keygen -o ~/backup-key.txt 2>/dev/null && chmod 600 ~/backup-key.txt && grep 'public key' ~/backup-key.txt
pbcopy < ~/backup-key.txt      # straight to the clipboard; paste into the password manager
pbcopy < /dev/null             # then clear the clipboard
age-keygen -y ~/backup-key.txt # later: shows only the PUBLIC half, to check it's there

The same lesson appeared three times in this project — a tunnel token, alarm URLs, and this key all showed up in screenshots. Type clear and look before you screenshot.

Step 6 — The backup script

The nightly job, trimmed to its working parts:

bash
#!/usr/bin/env bash
set -euo pipefail
umask 077

# rclone configured entirely from environment variables — no config file
# holding the key. NO_CHECK_BUCKET: the token may only touch objects, so
# rclone must not try to check or create the bucket.
export RCLONE_CONFIG_R2_TYPE=s3
export RCLONE_CONFIG_R2_PROVIDER=Cloudflare
export RCLONE_CONFIG_R2_ACCESS_KEY_ID="$R2_ACCESS_KEY_ID"
export RCLONE_CONFIG_R2_SECRET_ACCESS_KEY="$R2_SECRET_ACCESS_KEY"
export RCLONE_CONFIG_R2_ENDPOINT="https://${R2_ACCOUNT_ID}.r2.cloudflarestorage.com"
export RCLONE_CONFIG_R2_NO_CHECK_BUCKET=true

NAME="lampsill-$(date -u +%Y%m%d-%H%M%S).dump.age"
WORK="$(mktemp -d)"; trap 'rm -rf "$WORK"' EXIT

# Dump and encrypt in one stream. pipefail: if pg_dump fails, the whole
# line fails — even though age would happily encrypt half a dump.
pg_dump --format=custom --dbname=lampsill \
  | age --encrypt --recipient "$AGE_RECIPIENT" --output "$WORK/$NAME"

# An empty database still dumps to a few KB; under 1 KB means a broken run.
BYTES="$(stat -c %s "$WORK/$NAME")"
[ "$BYTES" -gt 1024 ] || { echo "Backup is only $BYTES bytes — not uploading." >&2; exit 1; }

rclone copyto "$WORK/$NAME" "r2:$R2_BUCKET/db/$NAME" --retries 5
echo "Backed up → r2:$R2_BUCKET/db/$NAME ($((BYTES / 1024)) KB, encrypted)"

Who runs it — and why not root

A detail that's easy to miss. The app's code folder belongs to the app's user, lampsill — the same user PHP runs as. If the nightly backup ran as root from that folder, then anyone who ever broke into the PHP app could edit the backup script and have it run as root that night.

So:

  • The script is copied to /usr/local/sbin/, owned by root, where the app user can't change it.
  • It runs as the postgres user — enough to dump the database, and nothing more.
  • Its settings (the R2 key) live in /etc/lampsill/backup.env, readable by root only. systemd reads that file as root, then starts the job as postgres.

The systemd service:

ini
[Service]
Type=oneshot
User=postgres
Group=postgres
EnvironmentFile=/etc/lampsill/backup.env
EnvironmentFile=-/etc/lampsill/healthchecks.env
ExecStart=/usr/local/sbin/lampsill-backup
# success: the normal ping
ExecStartPost=/bin/sh -c 'if [ -n "$HC_BACKUP_URL" ]; then curl -fsS -m 10 --retry 3 -o /dev/null "$HC_BACKUP_URL" || true; fi'
# failure: tell healthchecks.io straight away, instead of waiting for the grace period
ExecStopPost=/bin/sh -c 'if [ "$SERVICE_RESULT" != success ] && [ -n "$HC_BACKUP_URL" ]; then curl -fsS -m 10 --retry 3 -o /dev/null "$HC_BACKUP_URL/fail" || true; fi'
Nice=10                    # gentle on a 1 GB machine
IOSchedulingClass=idle
PrivateTmp=yes
NoNewPrivileges=yes
ProtectSystem=full

And the timer:

ini
[Timer]
OnCalendar=*-*-* 02:30:00 UTC   # the quietest hour for UK users
Persistent=true                 # if the server was off at 02:30, run as soon as it's back

Add a third healthchecks.io check, lampsill-backup, with Period 1 day and Grace 2 hours. A backup that quietly stopped in November should be found in November, not during the disaster.

Step 7 — Switching it on (safely)

A setup script asks for the values one at a time — the secret is typed hidden — and then does something I'd recommend for any scheduled job:

That way a timer is never running on settings that don't work. It also refuses obvious mistakes: pasting the private key where the public one belongs ("That's the PRIVATE key. It must never be on the server"), a bucket name that can't be valid, or the same value pasted twice.

My first real run:

text
Backed up lampsill → r2:lampsill-backups/db/lampsill-20260928-144636.dump.age (48 KB, encrypted)
✅ Backup works. Every night at 02:30 UTC from now on:
NEXT                        LEFT  UNIT
Tue 2026-09-29 02:30:00 UTC 11h   lampsill-backup.timer

Step 8 — The rehearsal: prove it restores

"Backed up" only means a file was uploaded. Whether that file can turn back into a database is a separate question — and the only way to answer it is to do it.

A small script on my laptop:

  1. 1Decrypts the downloaded backup with the private key.
  2. 2Starts a throwaway Postgres 18 in Docker.
  3. 3Restores the backup into it.
  4. 4Prints every table with its row count, and the last database migration.
  5. 5Deletes everything.

The core of it:

bash
age --decrypt --identity ~/backup-key.txt --output db.dump "$BACKUP"
docker run -d --name restore-test -e POSTGRES_PASSWORD=x postgres:18-alpine
# wait until it answers over TCP — the image's first-start runs a temporary
# server on the socket only, which would say "ready" and then shut down
until docker exec restore-test pg_isready -h 127.0.0.1 -q; do sleep 1; done
docker cp db.dump restore-test:/tmp/db.dump
docker exec restore-test createdb -U postgres lampsill
docker exec restore-test pg_restore -U postgres -d lampsill --no-owner --no-privileges --exit-on-error /tmp/db.dump

Download the file from the R2 dashboard (it saves as db_lampsill-….dump.age — Cloudflare adds the folder name to the front, which confused my first attempt), then run it:

text
==> What came back
 table                  | rows
------------------------+------
 contacts               |    0
 escalations            |    0
 households             |    0
 migrations             |   14
 users                  |    0
 … 21 tables in all
Migrations: 14 (last: 2026_09_06_140000_add_away_and_sleep_to_monitored_people)

✅ RESTORE WORKS. This backup can rebuild the database.

The tables are empty because nobody uses the API yet — but every table and every migration is there. When there's data, the same backup carries it.

Repeat the rehearsal every few months, and after any big change. It takes two minutes.

Restoring for real

If the live database is ever lost:

bash
# on the laptop
age -d -i ~/backup-key.txt -o ~/Desktop/lampsill.dump ~/Downloads/db_lampsill-<date>.dump.age
# upload lampsill.dump to the server (rebuilt first if needed), then on the server:
sudo systemctl stop lampsill-tick.timer lampsill-silence-tick.timer lampsill-backup.timer
sudo -u postgres pg_restore --clean --if-exists --no-owner --role=lampsill -d lampsill ~/lampsill.dump
sudo systemctl start lampsill-tick.timer lampsill-silence-tick.timer lampsill-backup.timer
rm ~/lampsill.dump   # and delete the unencrypted copy on the laptop

Stop the jobs first: they mustn't run against a half-restored database, and tonight's backup mustn't copy one.

How I tested all this before trusting it

Before touching the real server, I ran the whole chain in Docker: Ubuntu 26.04 with Postgres 18, and a local stand-in for R2. Every one of these was checked:

Pasting the private key into setup

Result
Refused; nothing saved

Setup with an unreachable R2

Result
Showed the error; timer left off

A successful backup

Result
Uploaded; the file starts age-encryption.org/v1 — no readable data

A failed pg_dump

Result
Nothing uploaded

Restore with the right key

Result
Every row back (5,000 / 250 / 14 in the test data)

Restore with the wrong key

Result
Refused with a clear error

"Restoring for real" over a damaged database

Result
Fully repaired, tables owned by the app user

What you should see

  • A db/ folder in the bucket with one .dump.age file per night.
  • lampsill-backup green on healthchecks.io.
  • "✅ RESTORE WORKS" from the rehearsal, at least once.
  • The private key in two places: the laptop, and your password manager.

Let's connect

Choose your preferred way

Available for new projects