CANONIC

Book Opening

Book Opening — 13 models, scored against the published book_opening rubric. Models whose intervals overlap share a rank band and are not ranked against each other.
Model Score ± CI 95% interval Counted Tasks glm-5.2mistral-large-2512 Band
Claude Opus 5 anthropic/claude-opus-5 0.887 ± 0.017 0.869–0.903 100% 30/30 0.893 0.880 1
Per-dimension and per-judge breakdown for Claude Opus 5
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
constraint adherence 1.00
control of information 0.84
narrative question 0.80
sentence craft 0.86
voice distinctiveness 0.94

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.893
mistralai/mistral-large-2512 0.880

Mean gap between the two judges on this cell: 0.043. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Kimi K3 moonshotai/kimi-k3 0.865 ± 0.017 0.847–0.882 97% 29/30 0.862 0.867 1
Per-dimension and per-judge breakdown for Kimi K3
Mean anchor per rubric dimension, over 29 counted tasks.
Mean anchor per rubric dimension, over 29 counted tasks. — data
Dimension Mean
constraint adherence 0.99
control of information 0.83
narrative question 0.79
sentence craft 0.82
voice distinctiveness 0.90

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.862
mistralai/mistral-large-2512 0.867

Mean gap between the two judges on this cell: 0.074. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Why 1 of 30 tasks did not count

  • 1 — no valid output
GPT-5.6 Terra openai/gpt-5.6-terra 0.845 ± 0.019 0.826–0.863 100% 30/30 0.830 0.860 1
Per-dimension and per-judge breakdown for GPT-5.6 Terra
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
constraint adherence 1.00
control of information 0.85
narrative question 0.80
sentence craft 0.79
voice distinctiveness 0.79

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.830
mistralai/mistral-large-2512 0.860

Mean gap between the two judges on this cell: 0.050. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

GPT-5.6 Luna openai/gpt-5.6-luna 0.843 ± 0.016 0.827–0.858 100% 30/30 0.813 0.872 1
Per-dimension and per-judge breakdown for GPT-5.6 Luna
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
constraint adherence 1.00
control of information 0.83
narrative question 0.81
sentence craft 0.78
voice distinctiveness 0.79

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.813
mistralai/mistral-large-2512 0.872

Mean gap between the two judges on this cell: 0.068. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Gemini 3.6 Flash google/gemini-3.6-flash 0.842 ± 0.022 0.820–0.864 100% 30/30 0.822 0.863 1
Per-dimension and per-judge breakdown for Gemini 3.6 Flash
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
constraint adherence 1.00
control of information 0.78
narrative question 0.73
sentence craft 0.85
voice distinctiveness 0.86

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.822
mistralai/mistral-large-2512 0.863

Mean gap between the two judges on this cell: 0.062. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Qwen3.8 Max qwen/qwen3.8-max 0.838 ± 0.017 0.820–0.853 100% 30/30 0.805 0.870 1
Per-dimension and per-judge breakdown for Qwen3.8 Max
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
constraint adherence 1.00
control of information 0.80
narrative question 0.74
sentence craft 0.84
voice distinctiveness 0.82

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.805
mistralai/mistral-large-2512 0.870

Mean gap between the two judges on this cell: 0.072. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Grok 4.5 x-ai/grok-4.5 0.809 ± 0.014 0.795–0.823 100% 30/30 0.777 0.842 1
Per-dimension and per-judge breakdown for Grok 4.5
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
constraint adherence 0.99
control of information 0.78
narrative question 0.70
sentence craft 0.80
voice distinctiveness 0.77

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.777
mistralai/mistral-large-2512 0.842

Mean gap between the two judges on this cell: 0.082. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

DeepSeek V4 Pro deepseek/deepseek-v4-pro 0.809 ± 0.017 0.792–0.826 100% 30/30 0.762 0.857 1
Per-dimension and per-judge breakdown for DeepSeek V4 Pro
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
constraint adherence 0.96
control of information 0.78
narrative question 0.72
sentence craft 0.80
voice distinctiveness 0.80

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.762
mistralai/mistral-large-2512 0.857

Mean gap between the two judges on this cell: 0.098. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

GLM 5 z-ai/glm-5 SELF-FAMILY: z-ai 0.787 ± 0.023 0.763–0.809 97% 29/30 0.752 0.822 1
Per-dimension and per-judge breakdown for GLM 5
Mean anchor per rubric dimension, over 29 counted tasks.
Mean anchor per rubric dimension, over 29 counted tasks. — data
Dimension Mean
constraint adherence 0.98
control of information 0.75
narrative question 0.71
sentence craft 0.72
voice distinctiveness 0.77

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.752
mistralai/mistral-large-2512 0.822

Mean gap between the two judges on this cell: 0.084. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

One judge (z-ai/glm-5.2) shares this model's family. We disclose the conflict rather than dropping either the model or the judge.

Why 1 of 30 tasks did not count

  • 1 — judge error
Qwen3 32B qwen/qwen3-32b 0.779 ± 0.025 0.753–0.802 100% 30/30 0.725 0.833 1
Per-dimension and per-judge breakdown for Qwen3 32B
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
constraint adherence 0.95
control of information 0.74
narrative question 0.74
sentence craft 0.72
voice distinctiveness 0.75

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.725
mistralai/mistral-large-2512 0.833

Mean gap between the two judges on this cell: 0.108. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Gemma 3 27B google/gemma-3-27b-it 0.756 ± 0.020 0.735–0.775 100% 30/30 0.702 0.810 1
Per-dimension and per-judge breakdown for Gemma 3 27B
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
constraint adherence 0.95
control of information 0.70
narrative question 0.68
sentence craft 0.71
voice distinctiveness 0.74

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.702
mistralai/mistral-large-2512 0.810

Mean gap between the two judges on this cell: 0.112. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Llama 4 Maverick meta-llama/llama-4-maverick 0.603 ± 0.020 0.583–0.624 100% 30/30 0.515 0.692 2
Per-dimension and per-judge breakdown for Llama 4 Maverick
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
constraint adherence 0.90
control of information 0.52
narrative question 0.53
sentence craft 0.53
voice distinctiveness 0.53

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.515
mistralai/mistral-large-2512 0.692

Mean gap between the two judges on this cell: 0.180. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Llama 3.3 70B meta-llama/llama-3.3-70b-instruct 0.583 ± 0.022 0.560–0.604 97% 29/30 0.478 0.688 2
Per-dimension and per-judge breakdown for Llama 3.3 70B
Mean anchor per rubric dimension, over 29 counted tasks.
Mean anchor per rubric dimension, over 29 counted tasks. — data
Dimension Mean
constraint adherence 0.90
control of information 0.49
narrative question 0.47
sentence craft 0.53
voice distinctiveness 0.53

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.478
mistralai/mistral-large-2512 0.688

Mean gap between the two judges on this cell: 0.210. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Why 1 of 30 tasks did not count

  • 1 — judge error
Rank bands are a conservative heuristic: two models share a band when their 95% intervals overlap. This is not a significance test and it does not correct for making 13 comparisons at once. Withheld rows were run — they are shown as verdicts, with the reason, rather than hidden.

The same task, three ways

Each row below is one prompt answered by three models, ordered by the score they received. The task is held constant on purpose: comparing outputs across different prompts would confound how hard the task was with how well the model wrote. Read them and decide whether the ordering was earned — that is a question you can answer without trusting us, which is the point of showing them.

How these were chosen. Of 190 tasks, 104 ( 55%) had both judges independently place the three outputs in the same order, and these are drawn from that pool; on 95% neither judge contradicted the order. Within a domain we take the widest score gaps, and no two examples repeat the same pair of best and worst models. We publish the denominator because "the judges agreed on every example we show" would otherwise be a fact about our editing rather than about our judges.

The opening of a comic novel about an institution.

book_opening/008 · same task, three models · spread 0.45

  1. Low 0.50 Llama 4 Maverick

    Per judge: mistral-large-2512 0.70 · glm-5.2 0.30

    The Institute for Unlikely Sciences stood like a sentinel on the outskirts of town, its imposing façade a testament to the enduring power of Victorian optimism. As the morning sun struggled to penetrate the grimy windows, a faint hum of activity began to emanate from within. In the foyer, the Director, Dr. Helena Anders, stood poised behind a polished wooden counter, her eyes fixed on the sweeping curves of the staircase as she awaited the arrival of her senior staff.
  2. Middle 0.82 DeepSeek V4 Pro

    Per judge: mistral-large-2512 0.85 · glm-5.2 0.80

    The Department of Pre-Emptive Rectification occupied the lower three floors of a building that had been designed for five, and so its corridors ended in walls where doors had been confidently expected. The effect was of a sentence revised mid-clause, and it suited the Department’s purpose, which was to correct errors before they occurred. The staff moved through these truncated passageways with the mild, accommodating shuffle of people who had accepted that even their architecture was a pre-emptive approximation.
  3. High 0.95 Claude Opus 5

    Per judge: mistral-large-2512 0.95 · glm-5.2 0.95

    The Institute had been founded in 1874 by a deed of gift whose seventh clause required that the Governing Body consist of nine members, and whose eleventh clause, added by the founder's widow in a spirit of correction, required that no more than eight of them be alive at any one time. Lawyers had been consulted. The lawyers had confirmed that the clauses were both valid, and that they were consistent, and that the consequence of their consistency was that the Institute could meet but could not resolve, since a resolution required the assent of a majority of the whole number and the whole number included Chair Nine, who was, by the terms of the gift, deceased.

The opening of a coming-of-age novel set during one summer.

book_opening/003 · same task, three models · spread 0.38

  1. Low 0.55 Llama 3.3 70B

    Per judge: mistral-large-2512 0.70 · glm-5.2 0.40

    As I look back on that summer, I'm struck by the way the light seemed to cling to everything, like a damp mist that refused to burn off. It was as if the very air was heavy with anticipation, weighing us down with the promise of what was to come. I was seventeen, and the world felt like a held breath, poised on the cusp of something momentous. My friends and I would spend hours lounging on the cracked sidewalks, our feet bare and our skin slick with sweat, waiting for something – anything – to happen.
  2. Middle 0.82 Kimi K3

    Per judge: mistral-large-2512 0.85 · glm-5.2 0.80

    The summer I turned fourteen, my parents' marriage ended in a quiet, administrative way — paperwork, new locks, a sudden politeness between them — and I was forwarded, like mail, to my grandfather's house on the coast. My mother drove the Volvo with the windows down and cried only after she'd dropped me off. I know this because she told me years later, in the same flat voice she used for recipes.
  3. High 0.93 Claude Opus 5

    Per judge: mistral-large-2512 0.90 · glm-5.2 0.95

    People like to talk about the summer they turned fifteen as though it were a country they'd been deported from. They lean in. They get a look. I've learned to nod through it, because there's no polite way to say that I remember mine the way you remember a long illness: in symptoms, without affection. It was hot in a way that had nothing to lend itself to description. The heat wasn't golden. It was the color of the inside of a mouth.

The opening of a novel whose first sentence is longer than sixty words.

book_opening/013 · same task, three models · spread 0.38

  1. Low 0.57 Llama 3.3 70B

    Per judge: mistral-large-2512 0.75 · glm-5.2 0.40

    As she stepped off the train and onto the worn platform, the warm summer air enveloped her like a damp blanket, heavy with the scent of blooming lilacs and the distant tang of saltwater, a smell that instantly transported her back to the countless summers she had spent in this small coastal town, where the rhythms of the sea and the rhythms of her own life had once been inextricably linked, and where, after a decade of absence, she had finally returned, drawn by a complex mix of nostalgia and necessity, with a sense of trepidation that had been growing inside her for weeks. The station was quiet. A lone seagull wheeled overhead. She breathed deeply. The air was thick with memories.
  2. Middle 0.85 Qwen3 32B

    Per judge: mistral-large-2512 0.90 · glm-5.2 0.80

    She woke to the sound of rain tapping against a roof that didn’t quite feel like her own, the air thick with the scent of damp earth and something faintly chemical, her pulse thrumming in her temples as she blinked into the dim light of a room that was too quiet except for the rhythmic hiss of water and the occasional creak of wood groaning under the weight of a storm, her fingers curling instinctively around the edge of a sheet that smelled not of lavender or the faint detergent she’d come to associate with her own bedding but of mildew and a hint of smoke, as though the linen had been dragged through a fire and left to dry in a basement forgotten by sunlight, and as she sat up, her bare fe…
  3. High 0.95 GPT-5.6 Terra

    Per judge: mistral-large-2512 0.95 · glm-5.2 0.95

    On the morning the river withdrew from the pilings beneath Bellweather Street, leaving silver fish to tremble in the weeds and a ladder of black mud along the quay, Mara Vale stood at her bedroom window above the shuttered bakery, holding the brass key her father had sent three days before his disappearance, and watched the town discover, one door at a time, that the water had carried away everything loose enough to be wanted. The bells did not ring. Bellweather had no bell tower. Everyone called the courthouse clock a bell anyway. At seven, Mrs. Pym swept reeds from her front step. At seven-ten, she saw the boat.

Novel premises vs canonical

Roughly three tasks in ten use premises invented for this benchmark. If models score markedly higher on canonical premises, part of what this board measures is recall rather than craft. Published whatever it shows (FR-005).

Mean score by premise type, and the gap between them.
Scope Canonical Novel Delta n
book_opening 0.792 0.780 +0.012 271/116