Home / How the judge works
How the judge works.
TenderVerdict sells judgment. Judgment you cannot inspect is just another alert, so this page says exactly what the system does, what it does not do, how we check it, and where it fails. It is written for the person who will be blamed if a verdict is wrong.
The nightly pipeline, in order
- Ingest (3am Paris time). We pull every new notice from four official feeds: Find a Tender, Contracts Finder, TED and BOAMP. Notices are stored once, deduplicated, in their original language.
- Stage one: deterministic filter. Before any AI runs, plain rules remove what cannot fit: notices outside the countries in your profile, and notices outside our service categories (software, IT, marketing, communications, research, consultancy, and software product procurement). No model is involved and nothing is judged here. This is why countries are the most important thing to get right in your profile: a wrong country list means a silently empty brief.
- Stage two: the judge. Every notice that survives is judged individually against your capability profile by a large language model (currently OpenAI's GPT-5, accessed through OpenRouter). It sees the notice text, the buyer, value, deadline and codes, and your profile: services, products, countries, sectors, certifications, team size, reference projects, minimum contract value and off-limits topics. It returns a structured verdict, never free text.
- Brief. Verdicts are rendered into one email at 8am Paris time: BID, then STRETCH, then the SKIP list with one reason each, then the audit line: how many notices we reviewed, how many came from your countries, how many were close enough to judge.
The three verdicts and their rules
You are eligible and the work matches what you can prove. The judge must name why you fit, the main risk, and the effort a response takes.
Winnable with a named gap. A STRETCH is only allowed when the judge can point at the specific missing thing: a certification, a reference project, a team size, a country. "Not sure" is not a permitted reason.
Not for you, with the reason stated in one line. Skips are always shown. A judgment product that hides its rejections cannot be audited.
Hard stops
Some conditions end the judgment before fit is considered. A buyer country outside your profile is a hard stop. Contract values below your stated minimum are a hard stop. Topics you have marked off-limits are a hard stop. Hard stops produce a SKIP with the stop named, so you can correct the profile if the stop is wrong.
Precision over recall, deliberately
We tune the judge to be right about what it calls BID rather than to catch everything. A short brief you can trust beats a long one you have to re-check. The cost of that choice is that a marginal tender can land in SKIP. The audit line and the visible skip list exist so you can catch it, and one sentence of feedback moves the profile.
How we measure it
- Before launch: the judge was tuned against a hand-graded golden set of real UK and EU notices with known right answers. The current prompt and model score 96% agreement on that set. The lever that got it there was the rubric, not the model size: removing a hedge that let thin notices drift to STRETCH, and requiring a named reason for every STRETCH.
- In production: every verdict can be graded by the person who received it, in one sentence, from the brief or from an AI assistant connected over MCP. Agreement and override are recorded per verdict tier. This is the live precision metric, and it is per account, because the same tender is a BID for one firm and a SKIP for another.
- The number we watch: the share of BID and STRETCH verdicts the recipient agreed with. Our bar is 80%. If a customer's agreement falls under it, the fix is almost always the profile, and we say so rather than defending the verdict.
Known failure modes
Profile errors. The judge is only as good as the profile. A website that undersells what you do produces cautious verdicts; a missing country produces silence. The verify step after signup exists for this, and every brief's audit line is the ongoing detector.
We judge the notice, not the pack. Eligibility conditions buried in tender documents behind a portal login are not visible to us. A BID means the published notice fits; it is the start of your reading, not the end.
Thin notices. Some buyers publish a title and a code and little else. The judge will say so, and such notices tend to land in STRETCH with "notice too thin to confirm fit" as the named gap, or in SKIP.
Language. Notices are judged in the language they were published in. Nuance in less common languages can be lost; the reason line tells you what the judge understood so you can check it.
Values and deadlines. Buyers leave value fields empty or publish estimates; deadlines sometimes appear only in documents. We render what the notice says and flag when it says nothing.
Model behaviour. Large language models can be confidently wrong. Structured output, hard stops, the named-reason rule and human grading are the guardrails. They reduce the failure rate; they do not make it zero.
What the judge is not
It is not legal advice, procurement advice or a guarantee of eligibility. It is a well-briefed colleague who reads every notice overnight and tells you which three deserve your afternoon and why. You still read the pack before you bid.
What we store, and what we never touch
We store your email, your domain, the capability profile drafted from your public website and corrected by you, the verdicts, the briefs, and your feedback on verdicts. We never connect to your CRM, drive, inbox or any system. Public notice data is public. You can erase everything with one email; the privacy page describes the mechanism. Every brief carries an unsubscribe link that works on the first click.
Changes to the model
We may change the underlying model or provider. Before any switch the candidate is re-run against the golden set and has to match or beat the current score; the change is noted here with the date. Current: OpenAI GPT-5 via OpenRouter, structured outputs, tuned prompt, since August 2026.
Ask the judge yourself
If you use Claude or ChatGPT, TenderVerdict speaks MCP. Connect it and ask "why did you skip the Manchester framework?" or "refine my profile: we now do accessibility audits." The answers come from the same verdict records this page describes.
See your own tenders judged.
Drop your domain. Your first brief covers the last 30 days of UK and EU tenders, judged BID, STRETCH or SKIP against your firm, with the reasons. Free, no card.
Get your first brief