01Services02Process03Projects04About05FAQ06Blog07Hire Me

25+ products shipped · $3.8M+ raised by clients

Back to Blog
AI EngineeringPart 2 of 6October 2, 20268 min read

A lane of its own

Subhankar Denria

Subhankar Denria

Software Architect · Product Engineer

~7 min

What this part does

Give a three-minute job timing rules of its own, so it runs exactly once, start to finish. Queue defaults are tuned for two-second jobs; a long, paid one earns its own lane.

The change
One new queue connection, and timeouts that agree with each other
Costs to run
Nothing — the same table, different timing rules
What you get
Every run executes exactly once, and a deploy never interrupts one

A real moment: the week a second worker is added

Adding a worker is the most ordinary scaling step there is. For a long, paid job it's also the moment a queue's default timing starts to matter.

The week a second worker is added

The moment

The feature is popular, so a second queue worker is added to keep up. An organisation’s research run is ninety seconds in and going well.

A typical first build

  • The job sits on the shared queue connection, with its 90-second limit
  • The queue assumes the first worker has gone and offers the job to the new one
  • The same research ends up paid for twice

This build

  • The job has a connection of its own, with eleven minutes of patience
  • The second worker picks up the next organisation instead
  • One run, finished once, paid for once

How: a dedicated queue connection with its visibility timeout set above the job timeout, and a test that keeps it there.

The 90-second assumption inside every default queue

Every job queue has a setting for how long a job may be held before the queue assumes its worker has gone. Depending on the stack it's called a visibility timeout, a retry-after or a redelivery timeout. It answers one question: how long may a job be held before we assume the worker died?

Past that point, the queue hands the job to another worker. This is a good feature. A server really can lose power mid-job, and without it that job would be stuck forever.

The app's default connection uses 90 seconds, which is generous for emails and push notifications. A research run takes three to five minutes.

Those two numbers were never meant for each other.

Who else is running this job

Worker A

started at 0:00 · still paying

Worker B

never woken
The 11-minute timeout is past the right-hand edge of this chart
0:001:002:003:004:005:00

One run, one bill

The run finishes long before the queue would ever think of handing it on.

The visibility timeout answers one question: how long may a job be held before the queue assumes the worker has gone? So it has to be longer than the job itself.

On a 90-second connection, the queue would look at a research run a minute and a half in, see only a job that's been reserved a long time, and offer it to a second worker. Depending on how many attempts the job allows, that worker would either run the research again or close the job as failed while the first is still working. Either way the same research gets paid for twice.

So this job never runs on that connection.

It's worth knowing early, because the behaviour only shows up once there is more than one worker — which is exactly when a product starts to grow.

Why a new queue name doesn't fix it

The instinct is to put the job on its own named queue and move on. Here that doesn't help, because the visibility timeout belongs to the connection, not the queue name. Every queue on that connection shares it.

Raising the whole connection to 11 minutes isn't the answer either. Then an email job that really does need a retry waits eleven minutes for it.

So the feature gets its own connection, pointed at the same table:

config
queue connection "ai":
    driver:              database
    table:               jobs          # the same table as everything else
    queue:               ai-jobs
    visibility timeout:  660 seconds   # the default connection uses 90

Same database, same jobs table, same driver. Only the timing rules differ. It's a five-line change, and it's the one the rest of the feature stands on.

Four numbers, and which of them are load-bearing

There are four timeouts in play, owned by four different systems that have never heard of each other. Two relationships between them matter, and one number matters less than it looks.

Four numbers, four owners

Everything keys off the job timeout

Job timeout · The job class

600s

The one that actually stops this job, and failed() tidies up

Worker default timeout · The worker

620s

Only for jobs with no timeout of their own — ignored here

Visibility timeout · The connection

660s

Outlasts the job timeout, so a running job is never handed on

Stop grace period · The process manager

660s

Outlasts the job timeout, so a deploy always lets a run finish

job < visibility · job ≤ stop grace

Two rules, and every run finishes exactly once.

Each number is set in a different file by a different system, and none of them validates the others. Only the visibility timeout can be tested from the codebase.

Visibility timeout > job timeout

What it guarantees
A job that is still running is never offered to a second worker

Stop grace period ≥ job timeout

What it guarantees
The process manager always waits for a run to finish, so a deploy never interrupts paid research

The worker's default timeout is the one that surprises people. A job that declares its own timeout ignores it. The queue uses the job's value when there is one, and the worker's default only covers jobs that don't set their own. It's kept above the job's 600 seconds so the two never disagree, but it isn't what stops this job.

Four numbers, two rules, and nothing in the framework that checks them for you. So the one a test can reach is asserted:

pseudocode
test "the visibility timeout outlasts the job timeout":
    expect connection("ai").visibility_timeout
        to be greater than GenerateIdeas.timeout

A config change that drops the visibility timeout below the job's timeout fails CI, so the guarantee holds through every future change. This is the single highest-value test in the suite, and it asserts nothing about AI at all.

The stop grace period — how long the process manager waits for a worker to finish before stopping it — lives in server config, outside the codebase. It's written down next to the other three, so the four are always reviewed together.

The workers

Three worker processes are dedicated to this connection, separate from the general worker:

config
worker pool "ai":
    connection:          ai
    processes:           3
    max attempts:        1
    default timeout:     620 seconds
    recycle after:       50 jobs, or 1 hour
    stop grace period:   660 seconds

One attempt is the setting worth stopping on. It's already this queue's default, and the job sets it too — but it's spelled out here because raising it is such a natural thing to do later. For most jobs a retry is exactly right. Here a retry means repeating paid web research. A failure is shown to the user, who can press the button again if they want to, and the failed run doesn't count against their allowance. The decision to spend money again belongs to the person, not the queue.

Recycling restarts each worker after fifty jobs or an hour, so every run starts on a fresh, lean process.

Three processes means three organisations can generate at once. That's up to 60 runs an hour, comfortably more than the feature is expected to need, so the real ceiling is the API provider's concurrency allowance, not the server.

Deploys use a graceful restart, which asks each worker to finish its current job and then exit. A running generation is never cut off mid-research.

The fourth person, and the five-second rule

Three workers means a fourth simultaneous request waits its turn. That's fine — and someone who is waiting deserves to be told they're in a queue.

When all three are busy

Workers · dedicated to this connection

busy

Org A

busy

Org B

busy

Org C

Org D presses the button

status: queued

1s

Under five seconds, the page says nothing. A free worker usually grabs the job in half a second, and a message that flashes up and vanishes would only be noise.

Five seconds is the whole trick: it’s the difference between informative and twitchy.

Telling someone they are in a queue turns a silent wait into an expected one.

So the run says so itself:

pseudocode
function isWaitingForWorker(run):
    return run.status == "queued"
       and run.created_at is more than 5 seconds ago

The five-second grace period is the whole trick. Without it, "Waiting to start" flashes up for the half-second it takes a free worker to grab the job — and a message that appears and vanishes is only noise, right at the moment you're asking someone to trust a three-minute wait.

With it, the message only appears when the wait is real.

Let's connect

Choose your preferred way

Available for new projects