A lane of its own

Subhankar Denria
Software Architect · Product Engineer
What this part does
Give a three-minute job timing rules of its own, so it runs exactly once, start to finish. Queue defaults are tuned for two-second jobs; a long, paid one earns its own lane.
- The change
- One new queue connection, and timeouts that agree with each other
- Costs to run
- Nothing — the same table, different timing rules
- What you get
- Every run executes exactly once, and a deploy never interrupts one
A real moment: the week a second worker is added
Adding a worker is the most ordinary scaling step there is. For a long, paid job it's also the moment a queue's default timing starts to matter.
The moment
The feature is popular, so a second queue worker is added to keep up. An organisation’s research run is ninety seconds in and going well.
A typical first build
- The job sits on the shared queue connection, with its 90-second limit
- The queue assumes the first worker has gone and offers the job to the new one
- The same research ends up paid for twice
This build
- The job has a connection of its own, with eleven minutes of patience
- The second worker picks up the next organisation instead
- One run, finished once, paid for once
How: a dedicated queue connection with its visibility timeout set above the job timeout, and a test that keeps it there.
The 90-second assumption inside every default queue
Every job queue has a setting for how long a job may be held before the queue assumes its worker has gone. Depending on the stack it's called a visibility timeout, a retry-after or a redelivery timeout. It answers one question: how long may a job be held before we assume the worker died?
Past that point, the queue hands the job to another worker. This is a good feature. A server really can lose power mid-job, and without it that job would be stuck forever.
The app's default connection uses 90 seconds, which is generous for emails and push notifications. A research run takes three to five minutes.
Those two numbers were never meant for each other.
Worker A
started at 0:00 · still payingWorker B
never wokenOne run, one bill
The run finishes long before the queue would ever think of handing it on.
On a 90-second connection, the queue would look at a research run a minute and a half in, see only a job that's been reserved a long time, and offer it to a second worker. Depending on how many attempts the job allows, that worker would either run the research again or close the job as failed while the first is still working. Either way the same research gets paid for twice.
So this job never runs on that connection.
It's worth knowing early, because the behaviour only shows up once there is more than one worker — which is exactly when a product starts to grow.
Why a new queue name doesn't fix it
The instinct is to put the job on its own named queue and move on. Here that doesn't help, because the visibility timeout belongs to the connection, not the queue name. Every queue on that connection shares it.
Raising the whole connection to 11 minutes isn't the answer either. Then an email job that really does need a retry waits eleven minutes for it.
So the feature gets its own connection, pointed at the same table:
queue connection "ai":
driver: database
table: jobs # the same table as everything else
queue: ai-jobs
visibility timeout: 660 seconds # the default connection uses 90Same database, same jobs table, same driver. Only the timing rules differ. It's a five-line change, and it's the one the rest of the feature stands on.
Four numbers, and which of them are load-bearing
There are four timeouts in play, owned by four different systems that have never heard of each other. Two relationships between them matter, and one number matters less than it looks.
Everything keys off the job timeout
Job timeout · The job class
600sThe one that actually stops this job, and failed() tidies up
Worker default timeout · The worker
620sOnly for jobs with no timeout of their own — ignored here
Visibility timeout · The connection
660sOutlasts the job timeout, so a running job is never handed on
Stop grace period · The process manager
660sOutlasts the job timeout, so a deploy always lets a run finish
job < visibility · job ≤ stop grace
Two rules, and every run finishes exactly once.
| The rule | What it guarantees |
|---|---|
| Visibility timeout > job timeout | A job that is still running is never offered to a second worker |
| Stop grace period ≥ job timeout | The process manager always waits for a run to finish, so a deploy never interrupts paid research |
Visibility timeout > job timeout
- What it guarantees
- A job that is still running is never offered to a second worker
Stop grace period ≥ job timeout
- What it guarantees
- The process manager always waits for a run to finish, so a deploy never interrupts paid research
The worker's default timeout is the one that surprises people. A job that declares its own timeout ignores it. The queue uses the job's value when there is one, and the worker's default only covers jobs that don't set their own. It's kept above the job's 600 seconds so the two never disagree, but it isn't what stops this job.
Four numbers, two rules, and nothing in the framework that checks them for you. So the one a test can reach is asserted:
test "the visibility timeout outlasts the job timeout":
expect connection("ai").visibility_timeout
to be greater than GenerateIdeas.timeoutA config change that drops the visibility timeout below the job's timeout fails CI, so the guarantee holds through every future change. This is the single highest-value test in the suite, and it asserts nothing about AI at all.
The stop grace period — how long the process manager waits for a worker to finish before stopping it — lives in server config, outside the codebase. It's written down next to the other three, so the four are always reviewed together.
The workers
Three worker processes are dedicated to this connection, separate from the general worker:
worker pool "ai":
connection: ai
processes: 3
max attempts: 1
default timeout: 620 seconds
recycle after: 50 jobs, or 1 hour
stop grace period: 660 secondsOne attempt is the setting worth stopping on. It's already this queue's default, and the job sets it too — but it's spelled out here because raising it is such a natural thing to do later. For most jobs a retry is exactly right. Here a retry means repeating paid web research. A failure is shown to the user, who can press the button again if they want to, and the failed run doesn't count against their allowance. The decision to spend money again belongs to the person, not the queue.
Recycling restarts each worker after fifty jobs or an hour, so every run starts on a fresh, lean process.
Three processes means three organisations can generate at once. That's up to 60 runs an hour, comfortably more than the feature is expected to need, so the real ceiling is the API provider's concurrency allowance, not the server.
Deploys use a graceful restart, which asks each worker to finish its current job and then exit. A running generation is never cut off mid-research.
The fourth person, and the five-second rule
Three workers means a fourth simultaneous request waits its turn. That's fine — and someone who is waiting deserves to be told they're in a queue.
Workers · dedicated to this connection
busy
Org A
busy
Org B
busy
Org C
Org D presses the button
status: queued
1sUnder five seconds, the page says nothing. A free worker usually grabs the job in half a second, and a message that flashes up and vanishes would only be noise.
Five seconds is the whole trick: it’s the difference between informative and twitchy.
So the run says so itself:
function isWaitingForWorker(run):
return run.status == "queued"
and run.created_at is more than 5 seconds agoThe five-second grace period is the whole trick. Without it, "Waiting to start" flashes up for the half-second it takes a free worker to grab the job — and a message that appears and vanishes is only noise, right at the moment you're asking someone to trust a three-minute wait.
With it, the message only appears when the wait is real.
Written by
Subhankar Denria
Software Architect · 25+ products shipped