Never paying twice

Subhankar Denria
Software Architect · Product Engineer
What this part does
Make sure one press of the button means one paid run — through double-clicks, second tabs, impatient refreshes, restarts and a monthly allowance that is counted fairly.
- The guard
- An in-flight check, plus an atomic lock the check can't replace
- Costs to run
- One cache key, held for ten seconds
- What you get
- One press, one paid run, and an allowance that only counts what was delivered
A real moment: a double-click, then a second tab
Nobody presses a button exactly once. With a two-a-month allowance, what happens to the extra presses decides whether the feature feels fair.
The moment
An admin double-clicks the button. Nothing seems to happen for a second, so they open the page in a second tab and press it there too.
A typical first build
- Each press starts its own research run
- Three runs are paid for, and all three count
- A two-a-month allowance is gone in four seconds
This build
- All three presses land on the same single run
- Both tabs show the same progress page
- One run is paid for, and one counts against the allowance
How: an in-flight check for the slow repeat, and an atomic per-organisation lock for the presses that arrive together.
Three clicks, one run
A user presses the button. Then they press it again, because nothing visibly happened for a second. Or they open a second tab. Or they refresh.
Every one of those sends another POST. All of them should land on the same single run, and what it takes depends on how many milliseconds apart they arrive.
- 0 msAReads: anything in flight?
- 3 msBAsks for the lock → sent straight to the run that already exists
- 6 msANothing found → insert run, queue job
One run · paid once
Scoped per organisation, so customers never block each other. Expires in ten seconds, so a crash can’t lock anyone out.
The slow second click — a few seconds or minutes later — is handled by an in-flight check. Before creating anything, look for a run that's already queued, researching or ideating, and redirect to it:
inFlight = first run for this organisation that is queued, researching or ideating
if inFlight exists:
redirect to its progress pageThe user lands on the progress page for the run they already started, which is what they wanted anyway.
The simultaneous double click needs something stronger. Two requests arriving within a millisecond can both run the check before either has inserted its row, so a check alone would let both through.
The check and the insert aren't atomic. A lock is:
lock = atomic lock "ideas:generate:{organisation.id}", expiring in 10 seconds
if lock could not be taken:
redirect to the landing page
try:
startRun(organisation) # the in-flight check lives inside
finally:
release lockTwo details make it safe to leave running in production:
- It's scoped per organisation. Different customers never block each other, which matters as soon as you have more than a handful.
- It expires in ten seconds. Long enough to cover check → insert → dispatch; short enough that a request crashing between
get()andrelease()can't lock a customer out for any length of time. Thefinallyhandles the ordinary case; the expiry handles the process being killed.
The one setting it relies on
An atomic lock only works if its store is shared between requests — the database, or a shared cache. With an in-process store, every request gets its own empty memory, so there is nothing shared for the lock to hold.
That's the one environment setting this feature depends on, so it's proven rather than assumed: a test holds the lock from outside, fires the request, and asserts that no run was created and no job was queued.
Every run reaches an ending
Servers restart. Deploys happen mid-run. Left alone, a row could go on saying researching with no process behind it.
Two mechanisms make sure every run is closed properly, and they cover different situations.
The job's failure hook runs when the queue gives up — including when the worker kills a job for running too long. It marks the run failed immediately. This covers almost everything.
The stale sweep covers the case where there's no process left to run any hook at all, like a server losing power. Any unfinished run older than twenty minutes is closed:
STALE_AFTER_MINUTES = 20There's no cron job for this, and that's deliberate. Instead, the check runs at the five points where it matters:
The landing page
So the user always lands somewhere useful
The in-flight check
So a new run can always start
The status endpoint
So the progress page always ends with a clear answer
Usage-limit queries
So an interrupted run never counts against an allowance
The start of the job
So a run already closed as stale is never researched
The one that protects the budget
| Where it runs | Why there |
|---|---|
| The landing page | So the user always lands somewhere useful |
| The in-flight check | So a new run can always start |
| The status endpoint | So the progress page always ends with a clear answer |
| Usage-limit queries | So an interrupted run never counts against an allowance |
| The start of the job itself | So a run already closed as stale is never researched |
The landing page
- Why there
- So the user always lands somewhere useful
The in-flight check
- Why there
- So a new run can always start
The status endpoint
- Why there
- So the progress page always ends with a clear answer
Usage-limit queries
- Why there
- So an interrupted run never counts against an allowance
The start of the job itself
- Why there
- So a run already closed as stale is never researched
That last row is the one that protects the budget. Under a very long queue, a run can be closed as stale before a worker ever reaches it. A guard at the top of handle() means the worker simply moves on, and money is only ever spent on research someone is waiting for.
job handle():
if run is already finished:
return # closed as stale while it waited; nobody is watching
...The twenty-minute constant lives on the model, so the controller, the progress page and the allowance queries all agree on what "stale" means. One definition, used everywhere.
Counting usage fairly
Each organisation gets two generations a month, with a daily guard on research runs underneath as a safety net. Counting that sounds like one line. It is five rules, and each one is there to keep the count fair.
count runs where
organisation = this one
and created this month
and reset_at is empty
and (
status = completed
or ( status in (queued, researching, ideating)
and created within the last 20 minutes )
)completed· last week
The allowance working as intended
researching· 2 min ago
Otherwise a second run walks past the limit
researching· 40 min ago
Interrupted — never the customer’s cost
failed· yesterday
Only delivered results are counted
completed· reset by support
reset_at is set; the research is kept
Every unticked row is a run the customer is never asked to spend allowance on.
- Completed runs count. That's the allowance working as intended.
- In-flight runs count. Otherwise someone starts a second run while the first is still going and walks straight past the limit.
- Stale in-flight runs don't count. An interrupted run is never the customer's cost.
- Failed runs don't count. The allowance is only spent on results that were delivered.
- Runs with
reset_atset don't count. Support can hand an organisation its allowance back without deleting anything — more on that in Part 6.
And one rule that isn't in the query at all: refining or dismissing an idea never creates a run. It's free by construction rather than by a special case in the counting, which means it can't drift out of sync with the limit later.
That's the pattern worth copying: the allowance only ever counts work the customer actually received.
Written by
Subhankar Denria
Software Architect · 25+ products shipped