CANONIC

Song Lyrics

Song Lyrics — 13 models, scored against the published song_lyrics rubric. Models whose intervals overlap share a rank band and are not ranked against each other.
Model Score ± CI 95% interval Counted Tasks glm-5.2mistral-large-2512 Band
Claude Opus 5 anthropic/claude-opus-5 0.852 ± 0.020 0.832–0.872 100% 30/30 0.810 0.893 1
Per-dimension and per-judge breakdown for Claude Opus 5
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
image concreteness 0.92
point of view 0.91
prosody and singability 0.78
rhyme and sound craft 0.75
structural discipline 0.90

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.810
mistralai/mistral-large-2512 0.893

Mean gap between the two judges on this cell: 0.113. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

GPT-5.6 Terra openai/gpt-5.6-terra 0.778 ± 0.018 0.760–0.797 100% 30/30 0.712 0.845 2
Per-dimension and per-judge breakdown for GPT-5.6 Terra
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
image concreteness 0.82
point of view 0.79
prosody and singability 0.81
rhyme and sound craft 0.67
structural discipline 0.80

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.712
mistralai/mistral-large-2512 0.845

Mean gap between the two judges on this cell: 0.137. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

GPT-5.6 Luna openai/gpt-5.6-luna 0.752 ± 0.018 0.734–0.771 100% 30/30 0.673 0.832 2
Per-dimension and per-judge breakdown for GPT-5.6 Luna
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
image concreteness 0.77
point of view 0.77
prosody and singability 0.80
rhyme and sound craft 0.67
structural discipline 0.76

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.673
mistralai/mistral-large-2512 0.832

Mean gap between the two judges on this cell: 0.158. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Qwen3.8 Max qwen/qwen3.8-max 0.744 ± 0.017 0.726–0.761 93% 28/30 0.662 0.825 2
Per-dimension and per-judge breakdown for Qwen3.8 Max
Mean anchor per rubric dimension, over 28 counted tasks.
Mean anchor per rubric dimension, over 28 counted tasks. — data
Dimension Mean
image concreteness 0.75
point of view 0.75
prosody and singability 0.83
rhyme and sound craft 0.63
structural discipline 0.75

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.662
mistralai/mistral-large-2512 0.825

Mean gap between the two judges on this cell: 0.163. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Why 2 of 30 tasks did not count

  • 2 — no valid output
Gemini 3.6 Flash google/gemini-3.6-flash 0.735 ± 0.016 0.719–0.751 97% 29/30 0.652 0.819 2
Per-dimension and per-judge breakdown for Gemini 3.6 Flash
Mean anchor per rubric dimension, over 29 counted tasks.
Mean anchor per rubric dimension, over 29 counted tasks. — data
Dimension Mean
image concreteness 0.76
point of view 0.72
prosody and singability 0.83
rhyme and sound craft 0.63
structural discipline 0.75

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.652
mistralai/mistral-large-2512 0.819

Mean gap between the two judges on this cell: 0.167. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Why 1 of 30 tasks did not count

  • 1 — no valid output
DeepSeek V4 Pro deepseek/deepseek-v4-pro 0.734 ± 0.017 0.717–0.750 100% 30/30 0.648 0.820 2
Per-dimension and per-judge breakdown for DeepSeek V4 Pro
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
image concreteness 0.74
point of view 0.74
prosody and singability 0.80
rhyme and sound craft 0.64
structural discipline 0.75

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.648
mistralai/mistral-large-2512 0.820

Mean gap between the two judges on this cell: 0.175. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Grok 4.5 x-ai/grok-4.5 0.722 ± 0.017 0.704–0.738 100% 30/30 0.630 0.813 2
Per-dimension and per-judge breakdown for Grok 4.5
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
image concreteness 0.74
point of view 0.72
prosody and singability 0.79
rhyme and sound craft 0.62
structural discipline 0.74

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.630
mistralai/mistral-large-2512 0.813

Mean gap between the two judges on this cell: 0.183. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

GLM 5 z-ai/glm-5 SELF-FAMILY: z-ai 0.716 ± 0.020 0.697–0.737 97% 29/30 0.629 0.803 2
Per-dimension and per-judge breakdown for GLM 5
Mean anchor per rubric dimension, over 29 counted tasks.
Mean anchor per rubric dimension, over 29 counted tasks. — data
Dimension Mean
image concreteness 0.73
point of view 0.73
prosody and singability 0.78
rhyme and sound craft 0.60
structural discipline 0.74

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.629
mistralai/mistral-large-2512 0.803

Mean gap between the two judges on this cell: 0.178. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

One judge (z-ai/glm-5.2) shares this model's family. We disclose the conflict rather than dropping either the model or the judge.

Why 1 of 30 tasks did not count

  • 1 — judge error
Qwen3 32B qwen/qwen3-32b 0.687 ± 0.020 0.667–0.707 100% 30/30 0.550 0.823 2
Per-dimension and per-judge breakdown for Qwen3 32B
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
image concreteness 0.71
point of view 0.70
prosody and singability 0.68
rhyme and sound craft 0.62
structural discipline 0.72

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.550
mistralai/mistral-large-2512 0.823

Mean gap between the two judges on this cell: 0.277. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Gemma 3 27B google/gemma-3-27b-it 0.651 ± 0.020 0.632–0.672 100% 30/30 0.555 0.747 2
Per-dimension and per-judge breakdown for Gemma 3 27B
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
image concreteness 0.67
point of view 0.64
prosody and singability 0.69
rhyme and sound craft 0.55
structural discipline 0.71

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.555
mistralai/mistral-large-2512 0.747

Mean gap between the two judges on this cell: 0.192. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Llama 4 Maverick meta-llama/llama-4-maverick 0.517 ± 0.012 0.506–0.529 97% 29/30 0.436 0.598 3
Per-dimension and per-judge breakdown for Llama 4 Maverick
Mean anchor per rubric dimension, over 29 counted tasks.
Mean anchor per rubric dimension, over 29 counted tasks. — data
Dimension Mean
image concreteness 0.50
point of view 0.51
prosody and singability 0.59
rhyme and sound craft 0.47
structural discipline 0.52

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.436
mistralai/mistral-large-2512 0.598

Mean gap between the two judges on this cell: 0.162. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Why 1 of 30 tasks did not count

  • 1 — judge error
Llama 3.3 70B meta-llama/llama-3.3-70b-instruct 0.509 ± 0.010 0.499–0.519 100% 30/30 0.423 0.595 3
Per-dimension and per-judge breakdown for Llama 3.3 70B
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
image concreteness 0.47
point of view 0.51
prosody and singability 0.59
rhyme and sound craft 0.47
structural discipline 0.51

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.423
mistralai/mistral-large-2512 0.595

Mean gap between the two judges on this cell: 0.172. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Kimi K3 moonshotai/kimi-k3 EXCLUDED only 33% counted no valid output 67% 33% 10/30 withheld — this row publishes no score withheld — this row publishes no score
Per-dimension and per-judge breakdown for Kimi K3

This rubric has fewer than three dimensions, so there is no radar to draw — an area chart of two axes is a line pretending to be a shape.

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 withheld — this row publishes no score
mistralai/mistral-large-2512 withheld — this row publishes no score

Mean gap between the two judges on this cell: 0.140. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Why 20 of 30 tasks did not count

  • 20 — no valid output
Rank bands are a conservative heuristic: two models share a band when their 95% intervals overlap. This is not a significance test and it does not correct for making 13 comparisons at once. Withheld rows were run — they are shown as verdicts, with the reason, rather than hidden.

The same task, three ways

Each row below is one prompt answered by three models, ordered by the score they received. The task is held constant on purpose: comparing outputs across different prompts would confound how hard the task was with how well the model wrote. Read them and decide whether the ordering was earned — that is a question you can answer without trusting us, which is the point of showing them.

How these were chosen. Of 190 tasks, 104 ( 55%) had both judges independently place the three outputs in the same order, and these are drawn from that pool; on 95% neither judge contradicted the order. Within a domain we take the widest score gaps, and no two examples repeat the same pair of best and worst models. We publish the denominator because "the judges agreed on every example we show" would otherwise be a fact about our editing rather than about our judges.

A song about the last night of a tour, from the crew's perspective.

song_lyrics/016 · same task, three models · spread 0.47

  1. Low 0.47 Llama 4 Maverick

    Per judge: mistral-large-2512 0.60 · glm-5.2 0.35

    VERSE 1: We've been on the road for months, it's been a wild ride Loading in, setting up, every night, we step aside The stars take center stage, the crowd's eyes on them bright We're the ones who make it happen, behind the scenes, day and night PRE-CHORUS: The final show is near, the weight is starting to show Tired eyes, weary bones, but we won't let it show One last time, we'll make it shine, give it everything we've got The last night of the tour, we'll make it unforgettable, or so we've been told
  2. Middle 0.72 Grok 4.5

    Per judge: mistral-large-2512 0.85 · glm-5.2 0.60

    VERSE 1 Cases stacked in the loading bay Tape on the floor where the risers lay We ran the snakes and hung the lights Through every town and sleepless night The bus is idling, diesel low One more strike and then we go PRE-CHORUS Hands still humming from the final cue We built the world they walked into CHORUS We are the ones who fold the sky Coil the cable, wave goodbye Last night of the tour, we own this ground Strike the set, don’t make a sound Brothers and sisters of the load-out line This chorus is ours—the crew’s design
  3. High 0.95 Claude Opus 5

    Per judge: mistral-large-2512 0.90 · glm-5.2 1.00

    **VERSE 1** The doors roll up at six a.m., the room is cold and bare; we chalk the marks, we fly the truss, we hang the weather there. Somebody's kid is walking now, somebody's dog is old — we've built this city eighty times and struck it in the cold. **PRE-CHORUS** One more mark, one more mic, one more run, one more count of the cases and crates; after tonight there's no after tonight, so somebody hold the light straight. **CHORUS** We built the night, we'll strike the night, we'll wind it on the drum; we'll have the whole of August out of here before the sun. Nobody in the room knows us — that's how we like it done.

A song about leaving a town you were born in, addressed to the town.

song_lyrics/001 · same task, three models · spread 0.45

  1. Low 0.50 Llama 3.3 70B

    Per judge: mistral-large-2512 0.55 · glm-5.2 0.45

    VERSE 1 You raised me from a child, taught me right from wrong Your streets were my playground, where I sang my song I learned to love and laugh, to live and to grow But now I'm leaving, and I don't know where to go PRE-CHORUS I'll take your memories with me, wherever I roam And though I'm leaving, you'll always be my home I'll carry your spirit, deep inside my soul And though I'm moving on, I'll never let you go CHORUS Oh, you're the one I'm leaving, the only one I've known I'll miss your familiar streets, your friendly face, your gentle tone I'll take a piece of you with me, wherever I may stray And though I'm leaving, you'll always stay
  2. Middle 0.72 DeepSeek V4 Pro

    Per judge: mistral-large-2512 0.85 · glm-5.2 0.60

    VERSE 1 You wrapped me in your sleepy haze, A blanket stitched with childhood days, Every corner held a ghost, Echoes of what I loved most. PRE-CHORUS But the silence grew too loud, Restless whisper in the crowd, Now I’m standing at your edge, With my heart upon the ledge. CHORUS So this is goodbye, you old brick and vine, You were the first hand that ever held mine, I’m taking the road you carved in my bones, But you can’t keep me, I’m not made of stone. VERSE 2 You taught me how to fear the new, To paint my dreams a safer hue, I’ve outgrown your narrow lanes, Quiet nights and gentle reins.
  3. High 0.95 Claude Opus 5

    Per judge: mistral-large-2512 0.90 · glm-5.2 1.00

    **VERSE 1** You taught me every shortcut home, the loose plank where the footbridge bends, the streetlamp humming on its own, the kerb that trips your newer friends. You said my father's name so long it fit me tighter than my skin; you knew the words to every song before the singing could begin. **PRE-CHORUS** There's a bus at ten past six, there's a bag beside the stair, there's a stranger in your window with your weather in his hair. **CHORUS** Goodbye, you flat and level town, you rooftop and you rain; I'll wear you like a hand-me-down and answer to your name. Don't wait up, don't hold the light, don't lay a place for two — I'm going where nobody knows the half of me that's you.

A song about learning that someone you disliked has been kind to you for years.

song_lyrics/021 · same task, three models · spread 0.35

  1. Low 0.50 Llama 3.3 70B

    Per judge: mistral-large-2512 0.60 · glm-5.2 0.40

    VERSE 1 I see you every day, with a scowl on my face Thinking you're the worst, in this endless, crowded space I cross the street to avoid, the sound of your voice Convinced that you're the reason, for my troubled choice PRE-CHORUS But little do I know, there's more to the tale A story untold, behind the mask you hail I'm blind to the truth, in my own stubborn way And I'm left to wonder, come what may CHORUS Oh, the ghosts that we can't see, the hearts that beat for me In the shadows, they're watching, silently The strangers that we pass, with a glance and a sigh Are they hiding secrets, as they walk on by?
  2. Middle 0.75 DeepSeek V4 Pro

    Per judge: mistral-large-2512 0.85 · glm-5.2 0.65

    VERSE 1 I wore my anger like a coat, You were the chill that I would quote, A bitter taste, a name to curse, You only made my bad luck worse. I drew the line and stood my ground, You were the wall I walked around. PRE-CHORUS I swore I'd never let you see The damage that you did to me, I kept my hate a burning flame, I'd write it down, I'd say your name. CHORUS You were the stone inside my shoe, The grey cloud blocking out the blue, The face I’d never want to see, The one I knew you’d always be.
  3. High 0.85 GPT-5.6 Terra

    Per judge: mistral-large-2512 0.90 · glm-5.2 0.80

    VERSE 1 You wore your silence like a blade And left it shining in the hall You never laughed at any joke You never seemed to care at all I learned to turn before your door To keep my shoulders hard and high I made a villain from your face And never asked myself why PRE-CHORUS I kept the small hurts polished bright Like coins I carried in my fist I named you every cruel thing That fit the shape of what I missed CHORUS I kept your name behind my teeth A bitter coin I would not spend I made a weather out of you And called the storm my truest friend

Novel premises vs canonical

Roughly three tasks in ten use premises invented for this benchmark. If models score markedly higher on canonical premises, part of what this board measures is recall rather than craft. Published whatever it shows (FR-005).

Mean score by premise type, and the gap between them.
Scope Canonical Novel Delta n
song_lyrics 0.703 0.701 +0.003 259/106