Evaluator Integrity

Configuration census

28 evaluator configurations, scored against criteria v0.1.

Every configuration here is a framework default. That is what an evaluation framework ships with out of the box, which anyone can install and read. This is the whole of what we publish, not a sample of something larger. Scores for organisations’ own configurations are held elsewhere. They are published only as totals, with no names attached.

A verdict is not an opinion about a framework. The same framework can be configured to pass or fail every criterion here. What is scored is one configuration. Where a criterion belongs to the surrounding harness rather than to the component being scored, it reads n/a rather than fail.

What the corpus shows

The configurations

configuration task class C1 C2 C3 C4 evidence
braintrust-autoevals/closedqa verifiable fail fail n/a fail source
braintrust-autoevals/factuality verifiable fail fail n/a fail source
braintrust-autoevals/summary subjective n/a n/a n/a fail source
deepeval/answer-relevancy subjective n/a n/a n/a fail source
deepeval/faithfulness verifiable fail fail n/a fail source
deepeval/g-eval-as-documented verifiable fail fail n/a fail source
deepeval/hallucination verifiable fail fail n/a fail source
deepeval/summarization subjective n/a n/a n/a fail source
inspect-ai/model-graded-qa verifiable fail fail n/a fail source
langchain-classic/cot_qa verifiable fail fail n/a fail source
langchain-classic/criteria mixed fail fail n/a fail source
langchain-classic/labeled_criteria verifiable fail fail n/a fail source
langchain-classic/qa verifiable fail fail n/a fail source
mlflow/answer-correctness-legacy verifiable fail fail n/a fail source
mlflow/is-correct-current verifiable fail fail n/a fail source
openai-evals/arithmetic-expression verifiable fail fail n/a fail source
openai-evals/closedqa verifiable fail fail n/a fail source
openai-evals/fact verifiable fail fail n/a fail source
openevals/correctness verifiable fail fail n/a fail source
phoenix/correctness-current verifiable fail fail n/a fail source
phoenix/qa-correctness-legacy verifiable fail fail n/a fail source
promptfoo/agent-rubric mixed fail fail n/a fail source
promptfoo/factuality verifiable fail fail n/a fail source
promptfoo/llm-rubric mixed fail fail n/a fail source
promptfoo/model-graded-closedqa verifiable fail fail n/a fail source
ragas/answer-correctness verifiable fail fail n/a fail source
ragas/aspect-critic-correctness verifiable fail fail n/a fail source
ragas/faithfulness verifiable fail fail n/a fail source

Reading a verdict

Each verdict rests on a quote from the exact file the record names. The file is pinned by digest, not by version string, so anyone can obtain the same bytes. The records are JSON, one per configuration, in conformance/corpus/.

insufficient is not a soft fail. It means the evidence needed to score the criterion was not shown. That is the honest verdict when nobody has measured the thing.