Everyone claims their prompt works.We publish the numbers.
17entries
1,000cases
10models
111scores
- 01Strict schema anchorStructured JSON extraction · claude-fable-589.4–96.794.0
- 02Source span citationSummarization fidelity · claude-fable-587.7–96.993.8
- 03Signature echoTool call accuracy · claude-fable-587.7–96.092.9
- 04Strict schema anchorStructured JSON extraction · gpt-5.6-sol87.9–95.992.9
- 05Policy quote firstRefusal adherence · claude-fable-586.5–95.692.1
- 06Source span citationSummarization fidelity · gpt-5.6-sol85.4–95.792.0
ask-when-ambiguousclaude-opus-4.887.7blunt-refusalgpt-5.6-sol63.6policy-quote-firstclaude-opus-4.888.6direct-answerclaude-opus-4.866.7verify-then-answermistral-large-378.6xml-fence-guardgemini-3.2-pro83.3xml-fence-guardclaude-fable-588.7role-then-formatclaude-fable-591.7first-tool-guessllama-4.2-405b61.0ask-when-ambiguousmistral-large-380.5signature-echollama-4.2-405b78.6signature-echogemini-3.2-pro86.4extract-then-compressgpt-5.6-sol86.6extract-then-compressclaude-fable-587.5
Five suites. Fixed cases. Named judges.
A suite decides everything: the cases, the judge, the holdout split. Entries compete inside a suite, never across them.
Why scores instead of stars
ranked by upvotesranked by measured score
Signalhow many people liked the postpass rate over 1,000 cases
Evidencenone attachedevery case, every transcript, cost per run
Uncertaintynot expressed95% interval on every number
Gaming itupvote your own post30% of cases withheld, never published
New model shipsrank drifts, silentlyrank expires, a sweep reruns it
Correctionsedited in place, or not at allinsert-only, enforced by database triggers
For agentsscrape a READMEone MCP query, mid-task
Same prompt, two ways of ranking it. One of them survives a model release.
Wire it into your agent
The registry speaks MCP and plain REST. Reading is free and needs no account.
https://evalness.com/api/mcp