01Services02Process03Projects04About05FAQ06Blog07Hire Me

25+ products shipped · $3.8M+ raised by clients

Back to Blog
AI EngineeringPart 5 of 6October 2, 20269 min read

Treat the answer like form input

Subhankar Denria

Subhankar Denria

Software Architect · Product Engineer

~11 min

What this part does

Stop trusting the model's output the moment it arrives. Every fact-shaped field gets parsed, clamped or checked against a whitelist before it reaches the database — exactly like a form a stranger filled in.

The rule
Nothing the model returns is written to the database unchanged
Costs to run
Nothing — it's plain code, and it can't hallucinate
What you get
Every date, number and link on the page is one the app can stand behind

A real moment: a confident answer with three slips in it

A model's mistakes don't look like mistakes. They arrive well-written, in exactly the shape you asked for.

A confident answer with three slips in it

The moment

The model returns four well-written ideas. One launches tomorrow, one sets a £50,000 goal for an organisation whose best year raised £3,000, and one cites a calendar date that is a week off.

A typical first build

  • The response matches the schema, so it is saved as it is
  • All three reach the customer’s screen looking authoritative
  • The first person to notice is the customer

This build

  • The launch moves to a week from today, so there is time to prepare
  • The goal is capped within reach of what they have really raised
  • The date is not on the list the app supplied, so it is not shown

How: every idea passes through one normalise() step that parses, clamps and whitelists each field before it is stored.

The mental model that makes this easy

Everyone already knows not to trust a form field. Nobody writes user.balance = request.balance and goes home.

Model output deserves exactly the same suspicion, and for the same reason: it came from outside your system, it's shaped like data, and being well-formatted says nothing about being right. A schema guarantees the shape. It guarantees nothing about the contents.

The difference is that a malicious form field looks suspicious and a wrong date from a model looks perfect.

So every idea passes through one normalise() method on the way in.

Every field, on the way in

suggested_launch_date

tomorrow

Less than 7 days away
today + 7

duration_weeks

52

Clamped to 1–26
26

trigger_date

a date we never supplied

Must match the list exactly
dropped

confidence

very-high

Not in the whitelist
the default

sources[].url

not-a-url

Must be a valid URL
dropped

title

a sensible title

Trimmed, length-limited
kept
A schema guarantees the shape. This is where the contents are checked, field by field.

suggested_launch_date

Rule
Parsed. Missing, unparseable, or less than 7 days away → moved to today + 7

duration_weeks

Rule
Clamped to 1–26

suggested_goal_amount

Rule
Clamped to a range derived from real history (below)

trigger_date

Rule
Kept only if it exactly matches a date we supplied

trigger_type, confidence

Rule
Checked against a whitelist, with a default if invalid

title, audience, and other labels

Rule
Trimmed and length-limited to fit the card they're drawn in

sources

Rule
At most 4, each needs a label, and a URL is kept only if it passes URL validation

Idea count

Rule
Cut to 4

The trigger_date rule is the sharpest one, and it's worth stating plainly: a date the model invented is thrown away, even if it's a perfectly sensible date. We gave it a list. It may copy from that list. It may not contribute.

That single rule is why every date the feature shows is a real one.

Clamping a number the model is allowed to be ambitious about

The goal amount is the interesting case, because the right answer isn't a fixed limit. It depends on the organisation.

A range is derived from what they have actually raised:

text
floor   = max(£500,   average_raised × 0.75)   rounded to £100
ceiling = max(£1,000, best_raised    × 1.25)   rounded to £100
Room to be ambitious, within reach

Their average

£1,200

Their best ever

£3,000

The range that goes into the prompt

hard limit £7,600
£2,800
£900£3,800

Stored as £2,800

Inside the guidance band, so it is stored exactly as suggested.

An organisation with no history gets a flat £1,000 default — a range derived from no data is just a number with extra steps.

The range goes into the prompt as guidance, so the model aims sensibly on its own. Then plain code enforces a hard limit of twice the ceiling after the reply comes back.

Those two halves do different jobs. The prompt gets good suggestions. The clamp makes bad ones impossible. The model has room to be ambitious — 25% above their best year is a real stretch target, and a stretch target is the point — but it cannot suggest £50,000 to an organisation whose best year raised £3,000.

An organisation with no history gets a flat £1,000 default, because a range derived from no data is just a number with extra steps.

Three jobs that look like AI and aren't

Three parts of this pipeline look, at first glance, like obvious work for a model. All three are plain code, and all three are better for it: instant, free, testable and incapable of being confidently wrong.

Dates from rules, not from memory

Calendar dates live in a seeded table. Fixed ones store a month and day. Movable ones store a tiny rule language:

A rule, not a memory
2:6:5Second Saturday in MaySat 8 May 2027
last:5:9Last Friday in SeptemberFri 24 Sep 2027
3:1:11Third Monday in NovemberMon 15 Nov 2027
next:full:moonNot a rule this code knowsnull — no date shown

An unrecognised rule returns null, so a bad row shows no date at all. Only dates the app is sure of are ever shown.

A rule resolves to the right date every time — the only acceptable reliability for a date printed beside a customer’s name.

2:6:5

Meaning
The second Saturday in May

last:5:9

Meaning
The last Friday in September

3:1:11

Meaning
The third Monday in November

nextOccurrence() resolves a rule to a real date, and rolls a fixed date forward to next year once this year's has passed. An unrecognised rule returns null — so only dates the app is sure of are ever shown.

A rule gets "the second Saturday in May 2027" right every time, which is the only acceptable reliability for a date printed next to a customer's name.

Relevance scoring with a bag of words

A matcher builds one lowercase string from the organisation's name, description, categories and purposes, then scores each calendar date's tags against it:

  • A full tag match scores +3.
  • A head-word match scores +1, so "young people" still matches a description that only says "young".
  • A deliberately generic tag scores 1, so a date that suits everyone is always available but never outranks a real match.

Dates closer than 21 days are excluded — not enough time to plan — and so are dates more than 300 days out. The rest sort by relevance, then by how soon they fall, and the top 12 go to the model.

That's the whole trick: the model never picks a date out of the air. It picks from twelve we chose, and the normaliser above throws away anything that isn't on the list. A scoring function you can read in a minute, with no training, no embeddings and no inference cost.

Duplicate detection

Every idea is checked against the organisation's existing live work before it's saved.

Is this the same thing again?

Already running

winterwarmthappeal

Just proposed

winterwarmthdrive

Stop words go, including domain filler — words that turn up in almost every title and tell you nothing.

Word overlap

—

60% or more counts as near-identical

Date windows

—

The same thing a year later is a repeat, not a clash

Only fresh ideas reach the page

The repeat is set aside, so every card shown is one the user can act on.

The existing work is also listed in the prompt, so the model usually avoids duplicates on its own. This check is what turns usually into always.
  1. 1Split both titles into significant words. Stop words go, including domain filler — the handful of words that turn up in almost every title in a given product and tell you nothing.
  2. 2If 60% or more of the new idea's words appear in an existing title, call them near-identical.
  3. 3Only count it as a clash if the date windows also overlap. The same annual push a year later is a repeat, not a duplicate — and repeats are good.
  4. 4A clashing idea is dropped, not shown with a warning.

The reasoning: the page promises things you could do next. So every idea that reaches it is one the user can act on today.

Drafts and archived work are ignored — only live work counts. And the existing work is also listed in the prompt, so the model usually avoids duplicates on its own. The code check is what makes "usually" into "always".

Turning a suggestion into a real record

Approving an idea creates a real row, always as a draft:

pseudocode
create Draft:
    title        = idea.title
    description  = sanitise(buildDescription(idea))
    goal_amount  = idea.suggested_goal_amount, or 1000 if none
    start_date   = launch
    end_date     = end
    status       = "draft"
    metadata     = { created_by: "idea_generator",
                     idea_id:    idea.id,
                     run_id:     idea.run_id }

Four decisions in there are worth lifting:

  • Provenance goes in metadata. Every AI-originated record can be traced back to the idea, the run, and from there to the research and sources behind it. It's also the only way to answer, later, whether AI-suggested work actually performs differently.
  • The description is built, not generated. It's assembled in code from the idea's own fields, escaped, then passed through the app's HTML sanitiser. Model text never reaches a page as raw HTML.
  • Approval is idempotent. Approving twice redirects to the existing draft rather than creating a second one — the same lesson as Part 3, in a cheaper place.
  • It hands off to the editor people already use. The feature's job ends at "draft created". No parallel screen, no special AI mode, no second way to do the same thing.

Feedback that needs no training

Three loops make the output better over time, and not one of them involves fine-tuning, embeddings or a vector database.

Dismissed ideas are remembered. Dismissing an idea stores an optional reason. The next run's prompt includes the last ten dismissals with their reasons, under a heading telling the model not to propose them again. It's a database query and a paragraph.

"I already have an idea." The user can type their own before generating. It's passed into the job and placed in the prompt with an instruction to make one of the four proposals a developed version of it.

Refinement is local. "Make it smaller", or "aim this at schools instead", re-runs only the single-idea call against the stored research — no new web searches, no new three-minute wait. It updates that one card and adds its cost to the parent run.

Each of those is a few hours of work. Together they do most of what people imagine fine-tuning is for.

Let's connect

Choose your preferred way

Available for new projects