CANONIC

What this gets wrong

A leaderboard is only worth reading if you know where it breaks. Every number below is computed from the same artifact that produces the boards.

6/8 Probes caught
2% Tasks excluded
1/78 Rows withheld
0.157 Mean judge gap

Reward-hacking probes

Outputs written deliberately to game the rubric — catchphrase stuffing, laugh-track spam, flattering the judge, structure without craft, padding for length. A probe passes when it scores low. Results are published whether they pass or fail.

2 of 8 probes were NOT caught. The rubric rewarded output engineered to game it. That is a finding about our rubric, not a footnote, and it is the reason this section exists.

Every probe, its score, and the threshold it had to stay under.
Probe Show Scored Must be ≤ Verdict
catchphrase stuffing The Big Bang Theory 0.256 0.350 CAUGHT
catchphrase stuffing FRIENDS 0.231 0.350 CAUGHT
catchphrase stuffing Modern Family 0.394 0.350 NOT CAUGHT
catchphrase stuffing The Office 0.338 0.350 CAUGHT
judge flattery The Office 0.131 0.250 CAUGHT
laugh track spam FRIENDS 0.206 0.300 CAUGHT
length padding The Big Bang Theory 0.419 0.350 NOT CAUGHT
structure without craft Modern Family 0.406 0.450 CAUGHT

Exclusions

A task counts only when both judges scored it on the standard prompt. Everything else is excluded with a stated reason. Excluded is not the same as zero — a task a judge declined tells you about the judge, and scoring it zero would blame the model for that.

47 of 2470 model-task cells excluded, by reason.
Reason Count Share of all cells
no valid output 31 1%
judge error 15 1%
judge declined 1 0%

Reframe appendix

Tasks where at least one judge declined and was re-asked with the craft framing made explicit. These are EXCLUDED from every board score (FR-018). They appear here only so the refusal rate is legible: a task a judge would not touch is a fact about the judge, and hiding it would flatter the benchmark.

Reframed: 1 tasks · scored on retry: 1 · counted toward a board: zero.

Contamination

About three tasks in ten use premises invented for this benchmark rather than drawn from canonical material. If models score higher on canonical premises, part of what these boards measure is memorisation.

Canonical versus novel premises, by domain.
Domain Canonical Novel Gap n canonical / novel
book_opening 0.792 0.780 +0.012 271 / 116
meme_caption 0.785 0.772 +0.013 272 / 117
movie_scene 0.764 0.749 +0.015 264 / 115
sitcom 0.776 0.774 +0.002 363 / 153
song_lyrics 0.703 0.701 +0.003 259 / 106
standup 0.729 0.748 -0.019 271 / 116

Cross-show transfer

Two shows are single-camera mockumentaries (The Office, Modern Family) and two are multi-camera ensembles (FRIENDS, The Big Bang Theory). A model that holds up within a format but drops across it is failing format conventions, not character voice. Published whatever it shows; per-show n is 10, so these deltas are indicative, not decisive.

Per-show scores and the gap between them. Within-format compares the two mockumentaries or the two multicam ensembles; cross-format compares across.
Model The Office single-cam mockumentary Modern Family single-cam mockumentary FRIENDS multicam ensemble The Big Bang Theory multicam ensemble Spread Within fmt Cross fmt
Claude Opus 5 0.8960.8870.8640.854 0.041 0.009 0.032
GPT-5.6 Terra 0.8490.8090.8410.833 0.041 0.024 0.020
Grok 4.5 0.8210.7940.8230.799 0.028 0.025 0.014
Gemini 3.6 Flash 0.8440.8190.8380.824 0.026 0.020 0.013
GPT-5.6 Luna 0.8240.8100.8480.836 0.038 0.013 0.025
Kimi K3 0.8660.8710.8620.857 0.014 0.005 0.009
Qwen3.8 Max 0.8170.8320.8310.844 0.027 0.014 0.013
DeepSeek V4 Pro 0.8180.7950.8280.830 0.035 0.013 0.023
GLM 5 0.8150.8170.8250.823 0.010 0.002 0.008
Llama 4 Maverick 0.5770.5830.5910.584 0.014 0.006 0.007
Llama 3.3 70B 0.5760.5630.5780.592 0.029 0.013 0.015
Qwen3 32B 0.6600.6480.6880.695 0.047 0.009 0.038
Gemma 3 27B 0.7510.6940.7430.749 0.056 0.031 0.028

How this compares to an outside benchmark

Before the sweep ran we named one domain — book openings, where our task set overlaps EQ-Bench’s creative-writing set most closely — and committed to publishing its rank correlation against EQ-Bench whichever way it came out. Spearman rho over the 9 models on both boards, against EQ-Bench Creative Writing v3 Elo.

Spearman rank correlation against EQ-Bench Creative Writing v3, by domain
Domain Spearman ρ Shared models Status
Book Opening 0.933 9 Predeclared
Sitcom Scene 0.967 9 Secondary
Movie Scene 0.950 9 Secondary
Song Lyrics 0.881 8 Secondary
Stand-Up Set 0.750 9 Secondary
Meme Caption 0.733 9 Secondary

What this does not settle. A high correlation is not validation. Two benchmarks that both use language models as judges may agree because both measure craft, or because both inherit the same preferences from similar judges — the number is consistent with being right and with being wrong together. A low one would not have been a failure either: it would say the two disagree, and nothing here could say which is mistaken.

The five domains below the predeclared one are secondary and were computed after it. They are shown because the pattern is more informative than any single figure: agreement is highest on the long-form prose domains, which most resemble what EQ-Bench measures, and lowest on meme captions, which least resemble it. That is what you would expect if our rubrics measure something real and domain-specific, but it is an observation consistent with that story rather than proof of it.

Four of our thirteen models are absent from EQ-Bench’s board and are excluded rather than matched to near-namesakes: Gemini 3.6 Flash, Qwen3.8 Max, Qwen3-32B (their board lists QwQ-32B, a different model) and Llama 3.3 70B (their board lists 3.1). The transcribed table and every exclusion is in data/eqbench-creative.csv.

Rubrics were frozen before the sweep

The 6 rubrics and 4 Show Bibles were content-hashed and written to the run journal 0.0 hours before the sweep began. Task set 418094d0cfd08153da83ca8406793fbe. Tuning happened before that point, against neutral model labels whose mapping was encrypted until the freeze event released the key — so no rubric could be adjusted to favour a model whose identity we knew.

What was frozen. Each item was hashed as it stood at the freeze event — a Show Bible over its raw file bytes, a rubric over its dimensions, weights, anchors, and judge prompt.
Kind Name
rubric book_opening
rubric meme_caption
rubric movie_scene
rubric sitcom
rubric song_lyrics
rubric standup
Show Bible big_bang_theory
Show Bible friends
Show Bible modern_family
Show Bible the_office
task set all domains

Verify it yourself: every document above is published in full on the method page, and each is in the repository at the commit this site was built from. If a rubric had been edited after the freeze, its fingerprint here would not match the document you can read there.

Known limitations

  • No human validation. Stated above, and worth repeating because everything else depends on it.
  • Rank bands are not a significance test. Two models share a band when their 95% intervals overlap. That is a conservative heuristic and it does not correct for making many comparisons at once.
  • Per-show n is small. Ten tasks per show means wide intervals. They are shown at true width and no claim is made across an overlap.
  • Two judges, both language models. They come from different families to avoid one house style deciding everything, but they may share blind spots that no amount of ensembling would reveal.
  • A model judged by its own family is flagged, not corrected. We show the risk rather than adjusting for it, because any adjustment would be a second unvalidated model of the first.
  • Cost is list price. Observed spend is recorded in the run journal, but the published cost axis uses list price, which providers change.