01Services02Process03Projects04About05FAQ06Blog07Hire Me

25+ products shipped · $3.8M+ raised by clients

Back to Blog
AI EngineeringPart 4 of 6October 2, 20269 min read

One door to the model

Subhankar Denria

Subhankar Denria

Software Architect · Product Engineer

~8 min

What this part does

Put every paid call behind one class, so that stop reasons, parsing and the bill are handled once, the same way, for every call the app makes.

The rule
No code anywhere else calls the API directly
Costs to run
Nothing — it's where the cost is finally measured
What you get
Every paid result is kept, and every cent is attributed to the run that spent it

A real moment: the right answer in the wrong wrapper

By the time an answer comes back, three minutes and the full cost of a run have been spent. What happens in the next few lines of code decides whether the customer ever sees it.

Three minutes of research, wrapped in a code fence

The moment

A run finishes its research and generation. The answer is exactly right, but it arrives with a Markdown code fence around the JSON.

A typical first build

  • One attempt to parse it, which fails
  • The run is marked as an error and the research is discarded
  • It was paid for, and nobody can say how much

This build

  • The fence is stripped and the answer is read
  • The run completes, and the admin sees four ideas
  • Its cost is on the run row, to the cent

How: every call goes through one gateway that checks the stop reason, decodes in three attempts, and returns the bill with the result.

Why a gateway, and not just a client

Every request to the model goes through one class with two entry points: call() for the blocking requests, and stream() for the research, which needs to report progress as it goes.

It would be easy to skip this. The SDK is already a client; wrapping a client in a client looks like architecture for its own sake.

It isn't, because there are four concerns that have to be handled identically every time, and none of them is obvious enough to be remembered twice:

  1. 1Read why the model stopped before reading what it said.
  2. 2Keep every result that was paid for.
  3. 3Price each call by the model that made it.
  4. 4Carry the bill through validation untouched.

Four jobs, one door.

1. A refusal arrives as a success

This is the detail worth knowing before you ship, and it costs nothing to handle once you do.

The 200 that isn't an answer
HTTP 200…with the real answer one field down
{
  "stop_reason": "refusal",
  "stop_details": { "category": … },
  "content": []
}

Checks the status code only

Would treat this as an answer and try to parse it. The status code alone can’t tell you.

Checks why it stopped, first

Throws a clear exception, and the run fails with something a human can act on.

max_tokens — a warning, not an error

The output may be complete or cut off mid-structure. The salvage chain decides which.

Check stop_reason before reading any content. It is the field that says whether the thing you're about to parse is a real answer.

When a safety classifier declines a request, you don't get a 4xx. You get HTTP 200, with stop_reason: "refusal" and a stop_details object saying which category fired. There is no usable answer in the content.

Code that only checks the HTTP status would go on to parse it as if it were an answer. So the gateway checks why the model stopped before it reads any content, and a declined request becomes a clear message a person can act on:

pseudocode
if response.stop_reason == "refusal":
    raise ModelRefused

if response.stop_reason == "max_tokens":
    log warning "response hit max_tokens; output may be truncated"

The max_tokens case is a warning rather than an exception on purpose. The output might be complete and might be cut off mid-structure — and the salvage chain below is the thing that decides which.

The general rule: check stop_reason first, every time. It's the field that tells you whether the thing you're about to parse is a real answer.

2. Salvaging JSON, because the run is already paid for

Both calls use JSON-schema structured output, which constrains the response to your schema:

pseudocode
request.output_config = {
    format: { type: "json_schema", schema: ideaSchema }
}

So the text should already be valid JSON. That's the point of the feature.

And the economics here are unusual: by the time you're parsing, the money is already spent. Three minutes of research is worth one more attempt to read it.

So decoding is a chain, not a single call:

Four lines that rescue a paid run
1

Decode as-is

Structured output means this almost always wins

2

Strip a Markdown code fence, decode again

Three backticks and a language tag

3

Take the first { to the last }, decode again

Rescues a stray sentence either side

4

Only then, stop

With a clear message, and the run closed properly

The general rule: the later a failure happens, the harder you should try to rescue it. A validation error on the way in is free to reject. A parse error on the way out is not.

By the time you're parsing, the money is already spent, so three minutes of research always gets one more attempt to be read.

Only after all three attempts does it throw. In practice step one almost always wins. The chain exists for the cases where it doesn't, and it costs four lines.

This is a general principle for anything expensive: the later a failure happens, the harder you should try to rescue it. A validation error on the way in is free to reject. A parse error on the way out is not.

3. The bill belongs to the result

Each call returns an immutable AiResult carrying the parsed data and a full billing breakdown: input tokens, output tokens, cache-read tokens, cache-write tokens, web searches and web fetches. It prices itself:

pseudocode
function costInUsdCents(result):
    rates = pricing[result.model] or pricing.default

    tokens = result.inputTokens      / 1,000,000 * rates.input
           + result.outputTokens     / 1,000,000 * rates.output
           + result.cacheReadTokens  / 1,000,000 * rates.cache_read
           + result.cacheWriteTokens / 1,000,000 * rates.cache_write

    searches = result.webSearches / 1,000 * pricing.web_search_per_1k

    return (tokens + searches) * 100
Six lines, two models, one total

What AiResult carries back

  • Input tokensper millionresearch
  • Output tokensper millionresearch
  • Cache-read tokensper million · cheaperresearch
  • Cache-write tokensper million · dearerresearch
  • Web searchesper thousand searchesresearch
  • Input + output tokensa different model, different ratesgeneration

Priced per model

Each stage is priced at its own model’s rate, so every line is right, not just the total.

Searches included

Billed on top of tokens, so counted on top of tokens. The number matches the invoice.

Because it lands on the run row, “what did this run cost?” and “what did we spend this month?” are SQL queries — not a trip to a dashboard, and not an estimate.

Stored as integer US cents, the currency the provider bills in, and as integer pence for display. Never a float.

Two details in there are what make the number match the invoice:

Rates are looked up per model. Research and generation run on different models at different prices, so each call is priced at its own model's rate and every line is right, not just the total.

Web searches are billed per search, on top of tokens. They're counted separately and added in, which matters most on exactly the research-heavy runs this feature exists for.

The job folds every result into running totals on the run row through a single addUsage() method, so research and generation can't record cost differently. Refining an idea later adds its cost to the original run, so the run's recorded cost is the true cost of everything done under it, not just the first pass.

Cost is stored twice: as integer US cents, the currency the provider bills in, and as integer pence for display. Integers both times, so totals always add up to the cent.

The payoff is that "what did this customer's run cost?" and "what did we spend this month?" are SQL queries. Not a trip to a provider's dashboard, and not an estimate.

4. Validation must not eat the bill

When the researcher strips invalid findings out of a response, it needs a new payload — but it must keep the original billing figures. Those searches happened. They were charged for. Whether a finding survived validation has nothing to do with it.

pseudocode
clean = result.withData(validFindings)     # new data, identical usage

withData() returns a copy with new data and untouched usage. Whatever validation removes, the recorded cost is always what was actually spent.

Schemas are prompts too

One last thing the gateway's schemas taught us, which wasn't obvious going in.

Descriptions carry instructions. The model reads each property's description, so the schema does part of the prompt's job — and it does it next to the field it applies to, which is far harder to get out of sync than a paragraph in a system prompt:

pseudocode
trigger_date:
    type: string
    description: "ISO date of the trigger, copied exactly from the data
                  given to you. Omit if there is none."

duration_weeks:
    type: integer
    description: "Length in weeks, between 1 and 26."

Enums close the vocabulary. trigger_type and confidence are enums, so the interface's badge colours and labels map straight from the value with no translation layer. The schema stops a sixth trigger type being invented. If a bad value somehow gets through, the code still falls back to a default — belt and braces, as Part 5 explains.

Everything rendered is required. Every field the templates draw is in required, with additionalProperties: false. That means the templates never need null guards on model output, which is a surprisingly large simplification once there are a dozen of them.

One thing the schema deliberately isn't asked to enforce: how many ideas come back. The list is cut to four in code. Counting is cheap, certain and testable in plain code, and it doesn't depend on schema features behaving a particular way across model versions.

There are two schemas — one returning a set of ideas, one returning a single idea for refinement — and the second reuses the first's idea definition, so the two can never drift apart.

Let's connect

Choose your preferred way

Available for new projects