The alarm that notices silence

Subhankar Denria
Software Architect · Product Engineer
What this part does
If the every-minute jobs ever stop — server down, code broken, disk full, Google switched it off — get an email within about six minutes, instead of finding out when someone wasn't checked on. It's the most important post in the series.
- Time
- 15 minutes
- Cost
- $0 — the free Hobbyist plan
- If you skip it
- You hear about a dead server 25 hours late — or never
Why normal monitoring doesn't work here
Most monitoring watches for errors: a crash, a failed request, a full disk. But the failure that matters most for Lampsill makes no error at all.
If the server is switched off, nothing runs, so nothing fails. If the timer is accidentally disabled, the job simply never starts. From the inside, a stopped job and a quiet night look identical. And any monitor running on the same server dies along with it — it can't report that the server is gone.
The answer is an old idea from trains: the dead man's switch. The driver has to keep pressing a pedal; if they stop — for any reason at all — the train brakes. You don't detect the failure. You detect the absence of "I'm fine".
For a server:
- 1After every successful run, the job sends a tiny "I'm alive" ping to an outside service.
- 2The outside service expects a ping every minute.
- 3If the pings stop, for any reason, it emails you.
It doesn't need to know why the pings stopped. That's its strength: it catches failures you haven't imagined yet.
Step 1 — healthchecks.io
healthchecks.io does exactly this, and its free Hobbyist plan includes 20 checks with email alerts. (Paid plans add SMS and phone calls; email is enough for this.)
Sign up with an address you actually read.
Why outside? The alarm has to survive the thing it's watching. healthchecks.io runs on completely different infrastructure from Google Cloud.
Step 2 — One check per job
Create two checks — one per every-minute job, because they can fail separately:
| Setting | Value | Why |
|---|---|---|
| Name | lampsill-tick (and lampsill-silence-tick) | |
| Schedule | Simple | |
| Period | 1 minute | How often a ping should arrive |
| Grace | 5 minutes | How long to wait before raising the alarm |
Name
- Value
lampsill-tick(andlampsill-silence-tick)
Schedule
- Value
- Simple
Period
- Value
- 1 minute
- Why
- How often a ping should arrive
Grace
- Value
- 5 minutes
- Why
- How long to wait before raising the alarm
⚠️ The form defaults to a period of 1 day and a grace of 1 hour. Leave those and you'd hear about a dead server 25 hours later.
The form's defaults · period 1 day + grace 1 hour
25 hoursWhat it should be · period 1 minute + grace 5 minutes
6 minutesChanging two fields hears about it 250× sooner.
Why a 5-minute grace? One slow run, or a server restart for an update, shouldn't wake you. A real stop lasts longer than five minutes.
Delete the "My First Check" it creates for you. Copy each check's ping URL — it looks like https://hc-ping.com/1a2b3c4d-….
Treat those URLs as mildly secret: anyone who has one can send fake "all fine" pings and hide a real failure.
Step 3 — Ping only after success
This is the heart of it — the ExecStartPost line in the job's systemd service from post 5:
[Service]
Type=oneshot
EnvironmentFile=-/etc/lampsill/healthchecks.env
ExecStart=/usr/bin/php artisan lampsill:tick
ExecStartPost=/bin/sh -c 'if [ -n "$HC_TICK_URL" ]; then curl -fsS -m 10 --retry 3 -o /dev/null "$HC_TICK_URL" || true; fi'systemd runs ExecStartPost only if ExecStart succeeded. So:
| What happens | Ping? | healthchecks.io |
|---|---|---|
| The job runs and succeeds | ✅ Yes | Stays green |
| The job crashes or errors | ❌ No | Goes red after 5 min |
| The timer is disabled | ❌ No (nothing runs) | Goes red after 5 min |
| The server is off / deleted / out of memory | ❌ No | Goes red after 5 min |
| Google stops the server on day 91 | ❌ No | Goes red after 5 min |
The job runs and succeeds
- Ping?
- ✅ Yes
- healthchecks.io
- Stays green
The job crashes or errors
- Ping?
- ❌ No
- healthchecks.io
- Goes red after 5 min
The timer is disabled
- Ping?
- ❌ No (nothing runs)
- healthchecks.io
- Goes red after 5 min
The server is off / deleted / out of memory
- Ping?
- ❌ No
- healthchecks.io
- Goes red after 5 min
Google stops the server on day 91
- Ping?
- ❌ No
- healthchecks.io
- Goes red after 5 min
Why not ping from inside the PHP code? A ping inside the job could fire halfway through a run that then fails. Pinging from systemd, after the job exits successfully, means a ping really does mean "the whole run worked".
The URLs live in a separate settings file, /etc/lampsill/healthchecks.env, readable only by root — systemd reads it as root before starting each run, so the job still gets it. To save them without typing a URL into a file by hand, this one line asks for each URL, checks they look right, saves them, and runs both jobs once:
read -p "tick URL: " T; read -p "silence-tick URL: " S; case "$T$S" in *https://hc-ping.com/*https://hc-ping.com/*) printf 'HC_TICK_URL=%s\nHC_SILENCE_URL=%s\n' "$T" "$S" | sudo tee /etc/lampsill/healthchecks.env >/dev/null && sudo chmod 600 /etc/lampsill/healthchecks.env && sudo systemctl start lampsill-tick.service lampsill-silence-tick.service && clear && echo "saved; pinged both";; *) echo "Those don't look like hc-ping.com URLs — nothing saved";; esacMind the order. Paste the URLs the wrong way round and each alarm watches the other job.
Within two minutes, both checks turn green.
Step 4 — Test the alarm. Actually test it.
Stop one of the jobs on purpose:
sudo systemctl stop lampsill-tick.timer; echo "STOPPED at $(date -u +%H:%M) UTC"Then watch:
| Time (UTC) | What happened |
|---|---|
| 07:40 | Timer stopped |
| ~07:41 | Check turns amber — "late" |
| 07:46:15 | Check turns red — "down". The email "DOWN | lampsill-tick" is sent |
| The other check stays green: the alarms are independent |
07:40
- What happened
- Timer stopped
~07:41
- What happened
- Check turns amber — "late"
07:46:15
- What happened
- Check turns red — "down". The email "DOWN | lampsill-tick" is sent
- What happened
- The other check stays green: the alarms are independent
healthchecks.io
The other check stays green: the alarms are independent.
Your mail
- 07:40Timer stopped on purpose.
- ~07:41lampsill-tick turns amber: late.
- 07:46:15Red: down. The email "DOWN | lampsill-tick" is sent.
- ↳…and lands in the Newsletter folder.
- afterA filter rule on the exact sender address → Inbox, flagged.
- testThe test email arrives in the Inbox.
Then start it again straight away:
sudo systemctl start lampsill-tick.timer; systemctl list-timers 'lampsill*' --no-pagerIt runs at once, the check turns green, and an "UP" email follows.
(If list-timers shows a dash under NEXT right after starting, you caught it mid-run. Check again a minute later.)
Step 5 — Make sure the email lands where you'll see it
The test found a real problem: my email provider filed the DOWN alert in the "Newsletter" folder. Alert emails contain an "unsubscribe" link, which looks like a mailing list to a spam filter.
An alarm filed with the newsletters is as good as no alarm.
Moving the email to the inbox by hand only teaches the filter — it can still guess wrong next time. A filter rule makes it certain. In most mail providers it's Settings → Filters → New filter:
- Name:
Server alarms - Condition: the sender's exact address, copied from a real alert email — not "From contains healthchecks.io"
- Action: Move to folder → Inbox (and Flag it)
Why the exact address: "contains" also matches a stranger who just puts "healthchecks.io" in their sender name. A rule that lifts mail out of spam and flags it is exactly what a phishing email would love to ride.
Every mail provider has an equivalent (Gmail: Filter messages like these → Never send it to Spam, Categorize as: Primary, Always mark it as important).
One habit to go with it: never act on the link inside an alert email. Open healthchecks.io yourself and look. A real alarm will be there; a fake one won't.
Then prove it: healthchecks.io → Integrations → your email → Test! The test email must arrive in the Inbox. Mine did.
And put the mail app on your phone, with notifications on for the inbox. The point is to find out wherever you are.
What went wrong (or nearly did)
- The default period (1 day) and grace (1 hour) would have made the alarm 25 hours late.
- The first alarm went to the Newsletter folder. Without the test, I'd have found that out during a real outage.
What you should see
- Both checks green on healthchecks.io, pinged within the last minute.
- A DOWN email, then an UP email, from your test — in your inbox.
Next up · Part 8 of 9 · 15 min read
Backups you have actually restored
If the disk dies, the server is deleted by mistake, or a bad update wrecks the data, rebuild the database as it was last night. Everything else on the server can be rebuilt from scripts in an hour. The database — people, their contacts, their history — can't.
Keep going Part 6: A public HTTPS address with no open portsWritten by
Subhankar Denria
Software Architect · 25+ products shipped