01Services02Process03Projects04About05FAQ06Blog07Hire Me

25+ products shipped · $3.8M+ raised by clients

Back to Blog
InfrastructurePart 7 of 9September 28, 20269 min read

The alarm that notices silence

Subhankar Denria

Subhankar Denria

Software Architect · Product Engineer

~11 min

What this part does

If the every-minute jobs ever stop — server down, code broken, disk full, Google switched it off — get an email within about six minutes, instead of finding out when someone wasn't checked on. It's the most important post in the series.

Time
15 minutes
Cost
$0 — the free Hobbyist plan
If you skip it
You hear about a dead server 25 hours late — or never

Why normal monitoring doesn't work here

Most monitoring watches for errors: a crash, a failed request, a full disk. But the failure that matters most for Lampsill makes no error at all.

If the server is switched off, nothing runs, so nothing fails. If the timer is accidentally disabled, the job simply never starts. From the inside, a stopped job and a quiet night look identical. And any monitor running on the same server dies along with it — it can't report that the server is gone.

The answer is an old idea from trains: the dead man's switch. The driver has to keep pressing a pedal; if they stop — for any reason at all — the train brakes. You don't detect the failure. You detect the absence of "I'm fine".

For a server:

  1. 1After every successful run, the job sends a tiny "I'm alive" ping to an outside service.
  2. 2The outside service expects a ping every minute.
  3. 3If the pings stop, for any reason, it emails you.

It doesn't need to know why the pings stopped. That's its strength: it catches failures you haven't imagined yet.

Step 1 — healthchecks.io

healthchecks.io does exactly this, and its free Hobbyist plan includes 20 checks with email alerts. (Paid plans add SMS and phone calls; email is enough for this.)

Sign up with an address you actually read.

Why outside? The alarm has to survive the thing it's watching. healthchecks.io runs on completely different infrastructure from Google Cloud.

Step 2 — One check per job

Create two checks — one per every-minute job, because they can fail separately:

Name

Value
lampsill-tick (and lampsill-silence-tick)

Schedule

Value
Simple

Period

Value
1 minute
Why
How often a ping should arrive

Grace

Value
5 minutes
Why
How long to wait before raising the alarm

⚠️ The form defaults to a period of 1 day and a grace of 1 hour. Leave those and you'd hear about a dead server 25 hours later.

How long until you hear about a dead server

The form's defaults · period 1 day + grace 1 hour

25 hours

What it should be · period 1 minute + grace 5 minutes

6 minutes

Changing two fields hears about it 250× sooner.

Why a 5-minute grace? One slow run, or a server restart for an update, shouldn't wake you. A real stop lasts longer than five minutes.

Delete the "My First Check" it creates for you. Copy each check's ping URL — it looks like https://hc-ping.com/1a2b3c4d-….

Treat those URLs as mildly secret: anyone who has one can send fake "all fine" pings and hide a real failure.

Step 3 — Ping only after success

This is the heart of it — the ExecStartPost line in the job's systemd service from post 5:

ini
[Service]
Type=oneshot
EnvironmentFile=-/etc/lampsill/healthchecks.env
ExecStart=/usr/bin/php artisan lampsill:tick
ExecStartPost=/bin/sh -c 'if [ -n "$HC_TICK_URL" ]; then curl -fsS -m 10 --retry 3 -o /dev/null "$HC_TICK_URL" || true; fi'

systemd runs ExecStartPost only if ExecStart succeeded. So:

The job runs and succeeds

Ping?
✅ Yes
healthchecks.io
Stays green

The job crashes or errors

Ping?
❌ No
healthchecks.io
Goes red after 5 min

The timer is disabled

Ping?
❌ No (nothing runs)
healthchecks.io
Goes red after 5 min

The server is off / deleted / out of memory

Ping?
❌ No
healthchecks.io
Goes red after 5 min

Google stops the server on day 91

Ping?
❌ No
healthchecks.io
Goes red after 5 min

Why not ping from inside the PHP code? A ping inside the job could fire halfway through a run that then fails. Pinging from systemd, after the job exits successfully, means a ping really does mean "the whole run worked".

The URLs live in a separate settings file, /etc/lampsill/healthchecks.env, readable only by root — systemd reads it as root before starting each run, so the job still gets it. To save them without typing a URL into a file by hand, this one line asks for each URL, checks they look right, saves them, and runs both jobs once:

bash
read -p "tick URL: " T; read -p "silence-tick URL: " S; case "$T$S" in *https://hc-ping.com/*https://hc-ping.com/*) printf 'HC_TICK_URL=%s\nHC_SILENCE_URL=%s\n' "$T" "$S" | sudo tee /etc/lampsill/healthchecks.env >/dev/null && sudo chmod 600 /etc/lampsill/healthchecks.env && sudo systemctl start lampsill-tick.service lampsill-silence-tick.service && clear && echo "saved; pinged both";; *) echo "Those don't look like hc-ping.com URLs — nothing saved";; esac

Mind the order. Paste the URLs the wrong way round and each alarm watches the other job.

Within two minutes, both checks turn green.

Step 4 — Test the alarm. Actually test it.

Stop one of the jobs on purpose:

bash
sudo systemctl stop lampsill-tick.timer; echo "STOPPED at $(date -u +%H:%M) UTC"

Then watch:

07:40

What happened
Timer stopped

~07:41

What happened
Check turns amber — "late"

07:46:15

What happened
Check turns red — "down". The email "DOWN | lampsill-tick" is sent

What happened
The other check stays green: the alarms are independent
Stopping a job on purpose

healthchecks.io

lampsill-tickup
lampsill-silence-tickup

The other check stays green: the alarms are independent.

Your mail

Inbox
Newsletter
  1. 07:40Timer stopped on purpose.
  2. ~07:41lampsill-tick turns amber: late.
  3. 07:46:15Red: down. The email "DOWN | lampsill-tick" is sent.
  4. ↳…and lands in the Newsletter folder.
  5. afterA filter rule on the exact sender address → Inbox, flagged.
  6. testThe test email arrives in the Inbox.
The real test, times in UTC, replayed quickly. Without it, the Newsletter folder would have been found during a real outage.

Then start it again straight away:

bash
sudo systemctl start lampsill-tick.timer; systemctl list-timers 'lampsill*' --no-pager

It runs at once, the check turns green, and an "UP" email follows.

(If list-timers shows a dash under NEXT right after starting, you caught it mid-run. Check again a minute later.)

Step 5 — Make sure the email lands where you'll see it

The test found a real problem: my email provider filed the DOWN alert in the "Newsletter" folder. Alert emails contain an "unsubscribe" link, which looks like a mailing list to a spam filter.

An alarm filed with the newsletters is as good as no alarm.

Moving the email to the inbox by hand only teaches the filter — it can still guess wrong next time. A filter rule makes it certain. In most mail providers it's Settings → Filters → New filter:

  • Name: Server alarms
  • Condition: the sender's exact address, copied from a real alert email — not "From contains healthchecks.io"
  • Action: Move to folder → Inbox (and Flag it)

Why the exact address: "contains" also matches a stranger who just puts "healthchecks.io" in their sender name. A rule that lifts mail out of spam and flags it is exactly what a phishing email would love to ride.

Every mail provider has an equivalent (Gmail: Filter messages like these → Never send it to Spam, Categorize as: Primary, Always mark it as important).

One habit to go with it: never act on the link inside an alert email. Open healthchecks.io yourself and look. A real alarm will be there; a fake one won't.

Then prove it: healthchecks.io → Integrations → your email → Test! The test email must arrive in the Inbox. Mine did.

And put the mail app on your phone, with notifications on for the inbox. The point is to find out wherever you are.

What went wrong (or nearly did)

  • The default period (1 day) and grace (1 hour) would have made the alarm 25 hours late.
  • The first alarm went to the Newsletter folder. Without the test, I'd have found that out during a real outage.

What you should see

  • Both checks green on healthchecks.io, pinged within the last minute.
  • A DOWN email, then an UP email, from your test — in your inbox.

Let's connect

Choose your preferred way

Available for new projects