Treat the answer like form input

Subhankar Denria
Software Architect · Product Engineer
What this part does
Stop trusting the model's output the moment it arrives. Every fact-shaped field gets parsed, clamped or checked against a whitelist before it reaches the database — exactly like a form a stranger filled in.
- The rule
- Nothing the model returns is written to the database unchanged
- Costs to run
- Nothing — it's plain code, and it can't hallucinate
- What you get
- Every date, number and link on the page is one the app can stand behind
A real moment: a confident answer with three slips in it
A model's mistakes don't look like mistakes. They arrive well-written, in exactly the shape you asked for.
The moment
The model returns four well-written ideas. One launches tomorrow, one sets a £50,000 goal for an organisation whose best year raised £3,000, and one cites a calendar date that is a week off.
A typical first build
- The response matches the schema, so it is saved as it is
- All three reach the customer’s screen looking authoritative
- The first person to notice is the customer
This build
- The launch moves to a week from today, so there is time to prepare
- The goal is capped within reach of what they have really raised
- The date is not on the list the app supplied, so it is not shown
How: every idea passes through one normalise() step that parses, clamps and whitelists each field before it is stored.
The mental model that makes this easy
Everyone already knows not to trust a form field. Nobody writes user.balance = request.balance and goes home.
Model output deserves exactly the same suspicion, and for the same reason: it came from outside your system, it's shaped like data, and being well-formatted says nothing about being right. A schema guarantees the shape. It guarantees nothing about the contents.
The difference is that a malicious form field looks suspicious and a wrong date from a model looks perfect.
So every idea passes through one normalise() method on the way in.
suggested_launch_date
tomorrow
duration_weeks
52
trigger_date
a date we never supplied
confidence
very-high
sources[].url
not-a-url
title
a sensible title
| Field | Rule |
|---|---|
suggested_launch_date | Parsed. Missing, unparseable, or less than 7 days away → moved to today + 7 |
duration_weeks | Clamped to 1–26 |
suggested_goal_amount | Clamped to a range derived from real history (below) |
trigger_date | Kept only if it exactly matches a date we supplied |
trigger_type, confidence | Checked against a whitelist, with a default if invalid |
title, audience, and other labels | Trimmed and length-limited to fit the card they're drawn in |
sources | At most 4, each needs a label, and a URL is kept only if it passes URL validation |
| Idea count | Cut to 4 |
suggested_launch_date
- Rule
- Parsed. Missing, unparseable, or less than 7 days away → moved to today + 7
duration_weeks
- Rule
- Clamped to 1–26
suggested_goal_amount
- Rule
- Clamped to a range derived from real history (below)
trigger_date
- Rule
- Kept only if it exactly matches a date we supplied
trigger_type, confidence
- Rule
- Checked against a whitelist, with a default if invalid
title, audience, and other labels
- Rule
- Trimmed and length-limited to fit the card they're drawn in
sources
- Rule
- At most 4, each needs a label, and a URL is kept only if it passes URL validation
Idea count
- Rule
- Cut to 4
The trigger_date rule is the sharpest one, and it's worth stating plainly: a date the model invented is thrown away, even if it's a perfectly sensible date. We gave it a list. It may copy from that list. It may not contribute.
That single rule is why every date the feature shows is a real one.
Clamping a number the model is allowed to be ambitious about
The goal amount is the interesting case, because the right answer isn't a fixed limit. It depends on the organisation.
A range is derived from what they have actually raised:
floor = max(£500, average_raised × 0.75) rounded to £100
ceiling = max(£1,000, best_raised × 1.25) rounded to £100Their average
£1,200
Their best ever
£3,000
The range that goes into the prompt
Stored as £2,800
Inside the guidance band, so it is stored exactly as suggested.
The range goes into the prompt as guidance, so the model aims sensibly on its own. Then plain code enforces a hard limit of twice the ceiling after the reply comes back.
Those two halves do different jobs. The prompt gets good suggestions. The clamp makes bad ones impossible. The model has room to be ambitious — 25% above their best year is a real stretch target, and a stretch target is the point — but it cannot suggest £50,000 to an organisation whose best year raised £3,000.
An organisation with no history gets a flat £1,000 default, because a range derived from no data is just a number with extra steps.
Three jobs that look like AI and aren't
Three parts of this pipeline look, at first glance, like obvious work for a model. All three are plain code, and all three are better for it: instant, free, testable and incapable of being confidently wrong.
Dates from rules, not from memory
Calendar dates live in a seeded table. Fixed ones store a month and day. Movable ones store a tiny rule language:
An unrecognised rule returns null, so a bad row shows no date at all. Only dates the app is sure of are ever shown.
| Rule | Meaning |
|---|---|
2:6:5 | The second Saturday in May |
last:5:9 | The last Friday in September |
3:1:11 | The third Monday in November |
2:6:5
- Meaning
- The second Saturday in May
last:5:9
- Meaning
- The last Friday in September
3:1:11
- Meaning
- The third Monday in November
nextOccurrence() resolves a rule to a real date, and rolls a fixed date forward to next year once this year's has passed. An unrecognised rule returns null — so only dates the app is sure of are ever shown.
A rule gets "the second Saturday in May 2027" right every time, which is the only acceptable reliability for a date printed next to a customer's name.
Relevance scoring with a bag of words
A matcher builds one lowercase string from the organisation's name, description, categories and purposes, then scores each calendar date's tags against it:
- A full tag match scores +3.
- A head-word match scores +1, so "young people" still matches a description that only says "young".
- A deliberately generic tag scores 1, so a date that suits everyone is always available but never outranks a real match.
Dates closer than 21 days are excluded — not enough time to plan — and so are dates more than 300 days out. The rest sort by relevance, then by how soon they fall, and the top 12 go to the model.
That's the whole trick: the model never picks a date out of the air. It picks from twelve we chose, and the normaliser above throws away anything that isn't on the list. A scoring function you can read in a minute, with no training, no embeddings and no inference cost.
Duplicate detection
Every idea is checked against the organisation's existing live work before it's saved.
Already running
Just proposed
Stop words go, including domain filler — words that turn up in almost every title and tell you nothing.
Word overlap
—
60% or more counts as near-identical
Date windows
—
The same thing a year later is a repeat, not a clash
Only fresh ideas reach the page
The repeat is set aside, so every card shown is one the user can act on.
- 1Split both titles into significant words. Stop words go, including domain filler — the handful of words that turn up in almost every title in a given product and tell you nothing.
- 2If 60% or more of the new idea's words appear in an existing title, call them near-identical.
- 3Only count it as a clash if the date windows also overlap. The same annual push a year later is a repeat, not a duplicate — and repeats are good.
- 4A clashing idea is dropped, not shown with a warning.
The reasoning: the page promises things you could do next. So every idea that reaches it is one the user can act on today.
Drafts and archived work are ignored — only live work counts. And the existing work is also listed in the prompt, so the model usually avoids duplicates on its own. The code check is what makes "usually" into "always".
Turning a suggestion into a real record
Approving an idea creates a real row, always as a draft:
create Draft:
title = idea.title
description = sanitise(buildDescription(idea))
goal_amount = idea.suggested_goal_amount, or 1000 if none
start_date = launch
end_date = end
status = "draft"
metadata = { created_by: "idea_generator",
idea_id: idea.id,
run_id: idea.run_id }Four decisions in there are worth lifting:
- Provenance goes in metadata. Every AI-originated record can be traced back to the idea, the run, and from there to the research and sources behind it. It's also the only way to answer, later, whether AI-suggested work actually performs differently.
- The description is built, not generated. It's assembled in code from the idea's own fields, escaped, then passed through the app's HTML sanitiser. Model text never reaches a page as raw HTML.
- Approval is idempotent. Approving twice redirects to the existing draft rather than creating a second one — the same lesson as Part 3, in a cheaper place.
- It hands off to the editor people already use. The feature's job ends at "draft created". No parallel screen, no special AI mode, no second way to do the same thing.
Feedback that needs no training
Three loops make the output better over time, and not one of them involves fine-tuning, embeddings or a vector database.
Dismissed ideas are remembered. Dismissing an idea stores an optional reason. The next run's prompt includes the last ten dismissals with their reasons, under a heading telling the model not to propose them again. It's a database query and a paragraph.
"I already have an idea." The user can type their own before generating. It's passed into the job and placed in the prompt with an instruction to make one of the four proposals a developed version of it.
Refinement is local. "Make it smaller", or "aim this at schools instead", re-runs only the single-idea call against the stored research — no new web searches, no new three-minute wait. It updates that one card and adds its cost to the parent run.
Each of those is a few hours of work. Together they do most of what people imagine fine-tuning is for.
Written by
Subhankar Denria
Software Architect · 25+ products shipped