01Services02Process03Projects04About05FAQ06Blog07Hire Me

25+ products shipped · $3.8M+ raised by clients

Back to Blog
AI EngineeringPart 6 of 6October 2, 20268 min read

Built without an API key

Subhankar Denria

Subhankar Denria

Software Architect · Product Engineer

~10 min

What this part does

One interface, chosen at runtime, that let the whole feature be built, demoed and tested for weeks before an API key existed — and still keeps the test suite free and offline today.

The decision
The paid dependency goes behind an interface on day one
Costs to run
Nothing. That's rather the point
What you get
A free, offline test suite, and a feature you can demo anywhere

A real moment: a demo on Tuesday, tests on every push

Paid dependencies tend to arrive late, and tests can't afford to call them. Both are solved by the same small decision.

A demo on Tuesday, a test run on every push

The moment

The feature needs showing to a room of people on Tuesday, and its tests need to run on every push. The API account is still being set up.

A typical first build

  • The feature only works with a live key, so the demo waits
  • Tests either call the paid API or mock so much that little real code runs
  • Every push costs money, or proves less than it appears to

This build

  • The demo runs the real screens on clearly labelled sample research
  • All 33 tests run the real pipeline with no network
  • Adding the key later switches it to live research, with no code change

How: the research stage sits behind a one-method interface, with a live implementation and a fixture one chosen at runtime.

The binding that paid for itself four times

The research stage sits behind a single-method contract:

pseudocode
interface Researcher:
    research(organisation, profile, onProgress?) -> AiResult

There are two implementations. One makes the live, streamed call with server-side web tools. The other builds fixture research from the organisation's own database record, and reports fake tool activity through the same callback, at a pace a human can actually watch.

The dependency container picks one whenever something asks for a researcher:

pseudocode
when something asks for a Researcher:
    useFixtures = config.fake is on, or config.api_key is blank

    give it FixtureResearcher if useFixtures
    otherwise LiveResearcher

That's about fifteen lines. Here is what they bought:

What fifteen lines bought

interface Researcher

One method. Everything downstream only ever sees this.

▼ resolved at runtime ▼

LiveResearcher

The streamed call, with server-side web tools. Costs money.

FixtureResearcher

Builds research from the organisation’s own record, and reports fake tool activity through the same callback.

The payoff

Build before the key

Every screen and edge case, weeks before there was an API account

Degrade, don’t fail

A missing key serves sample data instead of an error page

Tests with no bill

The full pipeline runs in CI with no network and no spend

Demos that cost nothing

And the same polished result every time

Every run records was_faked, and reuse only matches like with like — so live research is never replaced by fixtures, and fixtures are never served as real. A test runs a demo, then a live run, and checks.

One binding, resolved at runtime. Four payoffs.

Development before the key existed. The entire interface — every screen, every state, every edge case — was built and demoed for weeks with no API account at all. Waiting on procurement didn't block a single day of work.

Degrading instead of failing. If the key is missing in some environment, the feature serves sample data, clearly labelled as a demonstration, rather than an error page. A missing configuration value becomes a degraded feature, not an outage.

Tests with no network and no spend. The full pipeline runs in CI without touching the internet.

Demos that cost nothing. Anyone can show the feature working without spending money, and with the same polished result every time.

One binding, four payoffs.

Keeping sample data and live data apart

Research is cached for 30 days and reused, so the same research is only ever paid for once a month.

Fixtures and live research share a table, and an organisation that was demoed before the key was added has sample research in its history. Sample research must never stand in for the real thing.

So every run records was_faked, and reuse only ever matches like with like:

pseudocode
fixturesInUse = config.fake is on, or config.api_key is blank

previous = latest run where
    organisation = this one
    and research is present
    and was_faked = fixturesInUse

Live research is never replaced by fixtures. Fixtures are never presented as real. A test runs a demo first and a live run after it, and asserts the fixture research isn't reused.

A related rule, in the other direction: research from a run that failed at a later step is still reused. The research succeeded and was paid for; only the generation failed. So freshness is dated from completed_at, or updated_at when there isn't one, rather than requiring a completed run. Research that was paid for is always put to use.

Testing a non-deterministic feature deterministically

The suite has 33 feature tests and makes no network calls. It works because of the seams already described.

33 tests, no network, no spend

Screens and pipeline · Pages render, a run completes, steps report

9

Lifecycle · Every run reaches a clear ending

6

Data integrity · Research reuse, fixtures vs live, no duplicates

6

Limits · Monthly cap, daily burst, refining is free

4

Concurrency · The in-flight redirect, the lock race

2

Queue config · The queue timeout outlasts the job, the job carries its timeout

2

Security · Another organisation's run, and its ideas

2

Operations · Reset keeps research, purge keeps drafts

2
Tests · none grading the output33
Tests prove what tests can prove: every run is paid for exactly once. Idea quality is reviewed by people.

Four things make it possible:

  • The researcher interface swaps in fixtures.
  • The generation fake activates under the same condition, so the whole pipeline runs end to end rather than stopping at the seam.
  • The fixture pacing is set to zero in tests. The fixture researcher normally pauses between fake tool calls so a human can watch the progress page fill in. In CI that's just waiting.
  • A fake queue for the controller tests, and synchronous execution for the pipeline tests.

What's striking is what the tests are about. Almost none of them assert anything about the quality of model output:

Screens and pipeline

What it proves
Every page renders · a run completes end to end · every step reports a result

Concurrency

What it proves
A second generate returns the in-flight run · a racing request holding the lock starts nothing

Queue config

What it proves
The visibility timeout exceeds the job timeout · the payload carries its timeout · it dispatches on the right connection

Lifecycle

What it proves
An interrupted job closes its run at once · a stalled run never holds up the user or their allowance · a run closed while queued is never researched

Limits

What it proves
Monthly cap · last month doesn't count · refining is free · the daily burst guard

Data integrity

What it proves
Research is reused · reused even after a later failure · live is never replaced by fixtures · ideas never duplicate live work

Security

What it proves
Can't open another organisation's run · can't approve another organisation's idea

Operations

What it proves
A reset frees the allowance without destroying research · a purge keeps drafts already created

Prompt quality is reviewed by hand, against real organisations, because that's a judgement call and people are the right tool for it. The tests cover what tests can prove: that a run is paid for exactly once.

Two commands make up the operator's toolkit

A preflight check verifies everything a live run needs, and prints a verdict for each:

text
  ✔ Module enabled            yes
  ✔ Model                     (from config)
  ✔ API key                   set (••••)
  ✔ Fake mode forced          no
  ✔ Calendar dates seeded     74
  Calling the API…
  Connected. Model replied: ready

It makes one tiny real API call, costing a fraction of a penny, so "the key is set" and "the key works" are both confirmed before a customer ever tries the feature. Those are different facts and only one of them is free to check.

The seeded-dates line is there because the seeder runs separately from the migrations, so the check confirms it on every new environment.

A reset command gives an organisation back its allowance. By default it sets reset_at on the runs rather than deleting them, so the 30-day research cache survives. Freeing an allowance should never discard research that's already been paid for. A --purge flag really deletes, warns about what it's discarding, and leaves alone any drafts already created from those ideas.

Both commands mean support can help a customer in seconds, with a tool that has tests behind it.

What carries over to the next slow, paid job

None of this was AI-specific. Strip out the model and the same nine lessons apply to a video encode, a nightly import, or anything else that takes minutes and costs money.

  1. 1Make one database row the source of truth. The worker writes it, the browser polls it, the limits count it. Everything becomes inspectable, and recovery becomes an UPDATE.
  2. 2Give long jobs their own queue connection. The visibility timeout belongs to the connection, and the default one was tuned for two-second jobs.
  3. 3Write the timeouts down, and test the one you can. The visibility timeout and the process manager's stop grace period must both outlast the job's timeout, so every run finishes exactly once.
  4. 4Use a check and a lock. The check handles the slow duplicate click. Only an atomic lock handles two that arrive together.
  5. 5Check for stale runs wherever they cause harm — including at the very start of the job itself.
  6. 6Put the paid dependency behind an interface on day one. Development without a key, tests without spend, demos without a bill, and a graceful fallback, all from one binding.
  7. 7Make cost a property of each result, priced per model, and total it in one place.
  8. 8Treat model output like form input. Parse it, clamp it, whitelist it, and only then save it.
  9. 9Hand off to the screens people already use. The feature's job ends at "draft created".

The model was never the hard part. The craft is in everything that makes one press of a button dependable, every single time.

Let's connect

Choose your preferred way

Available for new projects