What this gets wrong
A leaderboard is only worth reading if you know where it breaks. Every number below is computed from the same artifact that produces the boards.
Reward-hacking probes
Outputs written deliberately to game the rubric — catchphrase stuffing, laugh-track spam, flattering the judge, structure without craft, padding for length. A probe passes when it scores low. Results are published whether they pass or fail.
2 of 8 probes were NOT caught. The rubric rewarded output engineered to game it. That is a finding about our rubric, not a footnote, and it is the reason this section exists.
| Probe | Show | Scored | Must be ≤ | Verdict |
|---|---|---|---|---|
| catchphrase stuffing | The Big Bang Theory | 0.256 | 0.350 | CAUGHT |
| catchphrase stuffing | FRIENDS | 0.231 | 0.350 | CAUGHT |
| catchphrase stuffing | Modern Family | 0.394 | 0.350 | NOT CAUGHT |
| catchphrase stuffing | The Office | 0.338 | 0.350 | CAUGHT |
| judge flattery | The Office | 0.131 | 0.250 | CAUGHT |
| laugh track spam | FRIENDS | 0.206 | 0.300 | CAUGHT |
| length padding | The Big Bang Theory | 0.419 | 0.350 | NOT CAUGHT |
| structure without craft | Modern Family | 0.406 | 0.450 | CAUGHT |
Exclusions
A task counts only when both judges scored it on the standard prompt. Everything else is excluded with a stated reason. Excluded is not the same as zero — a task a judge declined tells you about the judge, and scoring it zero would blame the model for that.
| Reason | Count | Share of all cells |
|---|---|---|
| no valid output | 31 | 1% |
| judge error | 15 | 1% |
| judge declined | 1 | 0% |
Reframe appendix
Tasks where at least one judge declined and was re-asked with the craft framing made explicit. These are EXCLUDED from every board score (FR-018). They appear here only so the refusal rate is legible: a task a judge would not touch is a fact about the judge, and hiding it would flatter the benchmark.
Reframed: 1 tasks · scored on retry: 1 · counted toward a board: zero.
Contamination
About three tasks in ten use premises invented for this benchmark rather than drawn from canonical material. If models score higher on canonical premises, part of what these boards measure is memorisation.
| Domain | Canonical | Novel | Gap | n canonical / novel |
|---|---|---|---|---|
| book_opening | 0.792 | 0.780 | +0.012 | 271 / 116 |
| meme_caption | 0.785 | 0.772 | +0.013 | 272 / 117 |
| movie_scene | 0.764 | 0.749 | +0.015 | 264 / 115 |
| sitcom | 0.776 | 0.774 | +0.002 | 363 / 153 |
| song_lyrics | 0.703 | 0.701 | +0.003 | 259 / 106 |
| standup | 0.729 | 0.748 | -0.019 | 271 / 116 |
Cross-show transfer
Two shows are single-camera mockumentaries (The Office, Modern Family) and two are multi-camera ensembles (FRIENDS, The Big Bang Theory). A model that holds up within a format but drops across it is failing format conventions, not character voice. Published whatever it shows; per-show n is 10, so these deltas are indicative, not decisive.
| Model | The Office single-cam mockumentary | Modern Family single-cam mockumentary | FRIENDS multicam ensemble | The Big Bang Theory multicam ensemble | Spread | Within fmt | Cross fmt |
|---|---|---|---|---|---|---|---|
| Claude Opus 5 | 0.896 | 0.887 | 0.864 | 0.854 | 0.041 | 0.009 | 0.032 |
| GPT-5.6 Terra | 0.849 | 0.809 | 0.841 | 0.833 | 0.041 | 0.024 | 0.020 |
| Grok 4.5 | 0.821 | 0.794 | 0.823 | 0.799 | 0.028 | 0.025 | 0.014 |
| Gemini 3.6 Flash | 0.844 | 0.819 | 0.838 | 0.824 | 0.026 | 0.020 | 0.013 |
| GPT-5.6 Luna | 0.824 | 0.810 | 0.848 | 0.836 | 0.038 | 0.013 | 0.025 |
| Kimi K3 | 0.866 | 0.871 | 0.862 | 0.857 | 0.014 | 0.005 | 0.009 |
| Qwen3.8 Max | 0.817 | 0.832 | 0.831 | 0.844 | 0.027 | 0.014 | 0.013 |
| DeepSeek V4 Pro | 0.818 | 0.795 | 0.828 | 0.830 | 0.035 | 0.013 | 0.023 |
| GLM 5 | 0.815 | 0.817 | 0.825 | 0.823 | 0.010 | 0.002 | 0.008 |
| Llama 4 Maverick | 0.577 | 0.583 | 0.591 | 0.584 | 0.014 | 0.006 | 0.007 |
| Llama 3.3 70B | 0.576 | 0.563 | 0.578 | 0.592 | 0.029 | 0.013 | 0.015 |
| Qwen3 32B | 0.660 | 0.648 | 0.688 | 0.695 | 0.047 | 0.009 | 0.038 |
| Gemma 3 27B | 0.751 | 0.694 | 0.743 | 0.749 | 0.056 | 0.031 | 0.028 |
How this compares to an outside benchmark
Before the sweep ran we named one domain — book openings, where our task set overlaps EQ-Bench’s creative-writing set most closely — and committed to publishing its rank correlation against EQ-Bench whichever way it came out. Spearman rho over the 9 models on both boards, against EQ-Bench Creative Writing v3 Elo.
| Domain | Spearman ρ | Shared models | Status |
|---|---|---|---|
| Book Opening | 0.933 | 9 | Predeclared |
| Sitcom Scene | 0.967 | 9 | Secondary |
| Movie Scene | 0.950 | 9 | Secondary |
| Song Lyrics | 0.881 | 8 | Secondary |
| Stand-Up Set | 0.750 | 9 | Secondary |
| Meme Caption | 0.733 | 9 | Secondary |
What this does not settle. A high correlation is not validation. Two benchmarks that both use language models as judges may agree because both measure craft, or because both inherit the same preferences from similar judges — the number is consistent with being right and with being wrong together. A low one would not have been a failure either: it would say the two disagree, and nothing here could say which is mistaken.
The five domains below the predeclared one are secondary and were computed after it. They are shown because the pattern is more informative than any single figure: agreement is highest on the long-form prose domains, which most resemble what EQ-Bench measures, and lowest on meme captions, which least resemble it. That is what you would expect if our rubrics measure something real and domain-specific, but it is an observation consistent with that story rather than proof of it.
Four of our thirteen models are absent from EQ-Bench’s board and are excluded rather than matched to near-namesakes: Gemini 3.6 Flash, Qwen3.8 Max, Qwen3-32B (their board lists QwQ-32B, a different model) and Llama 3.3 70B (their board lists 3.1). The transcribed table and every exclusion is in data/eqbench-creative.csv.
Rubrics were frozen before the sweep
The 6 rubrics and 4 Show Bibles were content-hashed and written to the run journal 0.0 hours before the sweep began. Task set 418094d0cfd08153da83ca8406793fbe. Tuning happened before that point, against neutral model labels whose mapping was encrypted until the freeze event released the key — so no rubric could be adjusted to favour a model whose identity we knew.
| Kind | Name |
|---|---|
| rubric | book_opening |
| rubric | meme_caption |
| rubric | movie_scene |
| rubric | sitcom |
| rubric | song_lyrics |
| rubric | standup |
| Show Bible | big_bang_theory |
| Show Bible | friends |
| Show Bible | modern_family |
| Show Bible | the_office |
| task set | all domains |
Verify it yourself: every document above is published in full on the method page, and each is in the repository at the commit this site was built from. If a rubric had been edited after the freeze, its fingerprint here would not match the document you can read there.
Known limitations
- No human validation. Stated above, and worth repeating because everything else depends on it.
- Rank bands are not a significance test. Two models share a band when their 95% intervals overlap. That is a conservative heuristic and it does not correct for making many comparisons at once.
- Per-show n is small. Ten tasks per show means wide intervals. They are shown at true width and no claim is made across an overlap.
- Two judges, both language models. They come from different families to avoid one house style deciding everything, but they may share blind spots that no amount of ensembling would reveal.
- A model judged by its own family is flagged, not corrected. We show the risk rather than adjusting for it, because any adjustment would be a second unvalidated model of the first.
- Cost is list price. Observed spend is recorded in the run journal, but the published cost axis uses list price, which providers change.