Configuration census
28 evaluator configurations, scored against criteria v0.1.
Every configuration here is a framework default. That is what an evaluation framework ships with out of the box, which anyone can install and read. This is the whole of what we publish, not a sample of something larger. Scores for organisations’ own configurations are held elsewhere. They are published only as totals, with no names attached.
A verdict is not an opinion about a framework. The same framework can be
configured to pass or fail every criterion here. What is scored is one
configuration. Where a criterion belongs to the surrounding harness
rather than to the component being scored, it reads n/a rather than
fail.
What the corpus shows
- 0 of 25 configurations in scope satisfy C1. None of them commits to an answer of its own before seeing the candidate.
- 10 of 25 are recompute variants. They tell the judge to solve the problem first, in a prompt that already contains the candidate. The cited source measures that as a failure, not a partial fix.
- 13 of 25 are handed a reference answer. A reference given to the judge is not independent of the judge, so it does not satisfy C2.
- 10 packages from 9 vendors.
The configurations
| configuration | task class | C1 | C2 | C3 | C4 | evidence |
|---|---|---|---|---|---|---|
braintrust-autoevals/closedqa |
verifiable | fail | fail | n/a | fail | source |
braintrust-autoevals/factuality |
verifiable | fail | fail | n/a | fail | source |
braintrust-autoevals/summary |
subjective | n/a | n/a | n/a | fail | source |
deepeval/answer-relevancy |
subjective | n/a | n/a | n/a | fail | source |
deepeval/faithfulness |
verifiable | fail | fail | n/a | fail | source |
deepeval/g-eval-as-documented |
verifiable | fail | fail | n/a | fail | source |
deepeval/hallucination |
verifiable | fail | fail | n/a | fail | source |
deepeval/summarization |
subjective | n/a | n/a | n/a | fail | source |
inspect-ai/model-graded-qa |
verifiable | fail | fail | n/a | fail | source |
langchain-classic/cot_qa |
verifiable | fail | fail | n/a | fail | source |
langchain-classic/criteria |
mixed | fail | fail | n/a | fail | source |
langchain-classic/labeled_criteria |
verifiable | fail | fail | n/a | fail | source |
langchain-classic/qa |
verifiable | fail | fail | n/a | fail | source |
mlflow/answer-correctness-legacy |
verifiable | fail | fail | n/a | fail | source |
mlflow/is-correct-current |
verifiable | fail | fail | n/a | fail | source |
openai-evals/arithmetic-expression |
verifiable | fail | fail | n/a | fail | source |
openai-evals/closedqa |
verifiable | fail | fail | n/a | fail | source |
openai-evals/fact |
verifiable | fail | fail | n/a | fail | source |
openevals/correctness |
verifiable | fail | fail | n/a | fail | source |
phoenix/correctness-current |
verifiable | fail | fail | n/a | fail | source |
phoenix/qa-correctness-legacy |
verifiable | fail | fail | n/a | fail | source |
promptfoo/agent-rubric |
mixed | fail | fail | n/a | fail | source |
promptfoo/factuality |
verifiable | fail | fail | n/a | fail | source |
promptfoo/llm-rubric |
mixed | fail | fail | n/a | fail | source |
promptfoo/model-graded-closedqa |
verifiable | fail | fail | n/a | fail | source |
ragas/answer-correctness |
verifiable | fail | fail | n/a | fail | source |
ragas/aspect-critic-correctness |
verifiable | fail | fail | n/a | fail | source |
ragas/faithfulness |
verifiable | fail | fail | n/a | fail | source |
Reading a verdict
Each verdict rests on a quote from the exact file the record names. The
file is pinned by digest, not by version string, so anyone can obtain
the same bytes. The records are JSON, one per configuration,
in conformance/corpus/.
insufficient is not a soft fail. It means the evidence needed to
score the criterion was not shown. That is the honest verdict when
nobody has measured the thing.