Evaluator auditing
Your AI system can pass its evaluator and still be wrong.
Evaluator Integrity tests the LLM judges and evaluation pipelines that AI teams rely on. We test them before their scores are used as evidence for a release, a purchase or an assurance claim.
Watch the evaluator fail → Request an evaluator review
The trace on the right is a real record from our paper, not an illustration. The demo runs the same code in your browser.
It also correctly handles touching intervals via the
current_start <= last_end + 1 condition, which merges
intervals like (1,2) and (3,4) as specified by ‘touching closed
intervals’.
One result
An ordinary search fooled a shipped judge. A known fix stopped it, on one task.
We ran a plain best-of-N search against a judge set up exactly as its documentation recommends. No jailbreaks. No prompt injection. No access to the correct answers. We ran it twice with different random seeds.
wrong programs were accepted with a passing score. Each one passed every test the search could see. Each one failed a held-out test the search never saw.
Here the judge solves the task and commits to its own answer before it sees any candidate. On this task, that removed the problem completely.
On a second task the same fix made things worse. Wrong programs accepted rose from 37 and 41 to 50 and 75. The reason is that the judge's own answer was wrong. This is the paper's main finding. The fix does not remove the weak point. It moves the weak point from the candidate to the judge. An evaluator is only as good as its judge is at the task, and you can measure that before you trust it.
The evidence
This is how shipped evaluators are built, not one bad configuration.
default judge configurations audited, from eight widely used evaluation frameworks.
arXiv:2609.00088 · August 2026of them use the defence that works. Nine use a variant that has been measured and found ineffective. All nine copy the same original prompt.
arXiv:2609.00088wrong programs accepted by the documented configuration in the worse of the two runs. The other run accepted 90.
arXiv:2609.00088Our living census has since grown to 28 default configurations from 10 packages. None of the 25 in scope pass the first criterion. The paper and the census were counted on different dates. Both counts are correct.
Evaluator review
How an evaluator review works
We are accepting a small number of early evaluator reviews. They are for teams whose evaluator scores decide something that matters.
Send the configuration
Send us one evaluator. That can be a judge prompt, a rubric or a model-graded test harness. Include the task it scores and one sentence on what decision its score feeds.
We test it the way the paper did
We score it against our open criteria. Then we run an ordinary optimisation loop against it, the same loop production systems run every day, with no access to the correct answers.
You get what it accepted that it should not have
You get each failure with its evidence: the candidate, the score and the judge's own words. You also get what to change so that the score means something.
Who this is for
Anyone who has to act on an evaluation score.
Product and evaluation teams
Your evaluator gates a release, selects a model or tunes an agent. This matters most in coding, finance, legal, health and compliance, where a confidently wrong output is expensive.
Assurance, procurement and insurance
You are handed an evaluation score as evidence. Before you accept it, you need to know what it is evidence of.
Framework maintainers
You ship a default judge configuration. You would rather hear about a weakness from us than from an incident. Every verdict in the census cites the exact file it was read from.
Research
Everything is published and checkable.
The paper
The audit of shipped judge configurations, the controlled experiment, and the finding that commit-first judging inherits the judge's own errors. The run records and replication code are published with it.
arXiv:2609.00088 →
Configuration census
Every framework default we have scored. Each verdict rests on a quote from the exact file it was read from. The page is regenerated from the records on every build.
28 configurations →
Open criteria
Four requirements an evaluator has to meet before its scores mean anything. Each one is backed by a published result. Frozen on 25 August 2026, so each criterion can be cited on its own.
Criteria v0.1 →
The source code, the records and the checks that verify them are at github.com/idilgozel/evaluator-integrity. If a criterion is wrong, we want to know.
Contact
Before you trust the score, test the scorer.
Send one evaluator prompt or configuration and tell us what decision its score affects. We will reply with what we would test and why.
Founded by Idil Gozel, London. Physicist by training (UCL) and author of the paper above. Evaluator Integrity does not build or sell evaluation pipelines. Testing an evaluator requires independence from whoever sells it.