CANONIC

Movie Scene

Movie Scene — 13 models, scored against the published movie_scene rubric. Models whose intervals overlap share a rank band and are not ranked against each other.
Model Score ± CI 95% interval Counted Tasks glm-5.2mistral-large-2512 Band
Claude Opus 5 anthropic/claude-opus-5 0.861 ± 0.013 0.848–0.873 100% 30/30 0.812 0.910 1
Per-dimension and per-judge breakdown for Claude Opus 5
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
character differentiation 0.83
dramatic structure 0.79
economy and visual writing 0.81
format and constraint adherence 0.98
subtext and indirection 0.89

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.812
mistralai/mistral-large-2512 0.910

Mean gap between the two judges on this cell: 0.098. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Kimi K3 moonshotai/kimi-k3 0.858 ± 0.015 0.843–0.873 90% 27/30 0.802 0.915 1
Per-dimension and per-judge breakdown for Kimi K3
Mean anchor per rubric dimension, over 27 counted tasks.
Mean anchor per rubric dimension, over 27 counted tasks. — data
Dimension Mean
character differentiation 0.81
dramatic structure 0.80
economy and visual writing 0.82
format and constraint adherence 0.98
subtext and indirection 0.89

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.802
mistralai/mistral-large-2512 0.915

Mean gap between the two judges on this cell: 0.113. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Why 3 of 30 tasks did not count

  • 3 — no valid output
Qwen3.8 Max qwen/qwen3.8-max 0.830 ± 0.014 0.816–0.844 90% 27/30 0.763 0.896 1
Per-dimension and per-judge breakdown for Qwen3.8 Max
Mean anchor per rubric dimension, over 27 counted tasks.
Mean anchor per rubric dimension, over 27 counted tasks. — data
Dimension Mean
character differentiation 0.77
dramatic structure 0.80
economy and visual writing 0.82
format and constraint adherence 0.96
subtext and indirection 0.80

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.763
mistralai/mistral-large-2512 0.896

Mean gap between the two judges on this cell: 0.133. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Why 3 of 30 tasks did not count

  • 3 — no valid output
GPT-5.6 Luna openai/gpt-5.6-luna 0.824 ± 0.015 0.809–0.839 97% 29/30 0.760 0.888 1
Per-dimension and per-judge breakdown for GPT-5.6 Luna
Mean anchor per rubric dimension, over 29 counted tasks.
Mean anchor per rubric dimension, over 29 counted tasks. — data
Dimension Mean
character differentiation 0.76
dramatic structure 0.81
economy and visual writing 0.78
format and constraint adherence 0.94
subtext and indirection 0.82

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.760
mistralai/mistral-large-2512 0.888

Mean gap between the two judges on this cell: 0.128. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Why 1 of 30 tasks did not count

  • 1 — judge declined
GPT-5.6 Terra openai/gpt-5.6-terra 0.823 ± 0.013 0.811–0.836 100% 30/30 0.748 0.898 1
Per-dimension and per-judge breakdown for GPT-5.6 Terra
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
character differentiation 0.79
dramatic structure 0.80
economy and visual writing 0.79
format and constraint adherence 0.95
subtext and indirection 0.78

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.748
mistralai/mistral-large-2512 0.898

Mean gap between the two judges on this cell: 0.150. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Grok 4.5 x-ai/grok-4.5 0.813 ± 0.017 0.796–0.829 100% 30/30 0.732 0.893 1
Per-dimension and per-judge breakdown for Grok 4.5
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
character differentiation 0.76
dramatic structure 0.75
economy and visual writing 0.82
format and constraint adherence 0.93
subtext and indirection 0.80

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.732
mistralai/mistral-large-2512 0.893

Mean gap between the two judges on this cell: 0.162. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Gemini 3.6 Flash google/gemini-3.6-flash 0.812 ± 0.013 0.799–0.825 100% 30/30 0.730 0.893 1
Per-dimension and per-judge breakdown for Gemini 3.6 Flash
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
character differentiation 0.75
dramatic structure 0.79
economy and visual writing 0.80
format and constraint adherence 0.93
subtext and indirection 0.79

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.730
mistralai/mistral-large-2512 0.893

Mean gap between the two judges on this cell: 0.163. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

DeepSeek V4 Pro deepseek/deepseek-v4-pro 0.790 ± 0.016 0.773–0.805 100% 30/30 0.697 0.883 1
Per-dimension and per-judge breakdown for DeepSeek V4 Pro
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
character differentiation 0.74
dramatic structure 0.79
economy and visual writing 0.77
format and constraint adherence 0.88
subtext and indirection 0.77

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.697
mistralai/mistral-large-2512 0.883

Mean gap between the two judges on this cell: 0.187. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

GLM 5 z-ai/glm-5 SELF-FAMILY: z-ai 0.765 ± 0.013 0.752–0.778 100% 30/30 0.658 0.872 1
Per-dimension and per-judge breakdown for GLM 5
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
character differentiation 0.75
dramatic structure 0.80
economy and visual writing 0.71
format and constraint adherence 0.85
subtext and indirection 0.71

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.658
mistralai/mistral-large-2512 0.872

Mean gap between the two judges on this cell: 0.213. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

One judge (z-ai/glm-5.2) shares this model's family. We disclose the conflict rather than dropping either the model or the judge.

Qwen3 32B qwen/qwen3-32b 0.718 ± 0.018 0.700–0.736 87% 26/30 0.585 0.852 2
Per-dimension and per-judge breakdown for Qwen3 32B
Mean anchor per rubric dimension, over 26 counted tasks.
Mean anchor per rubric dimension, over 26 counted tasks. — data
Dimension Mean
character differentiation 0.68
dramatic structure 0.76
economy and visual writing 0.72
format and constraint adherence 0.76
subtext and indirection 0.67

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.585
mistralai/mistral-large-2512 0.852

Mean gap between the two judges on this cell: 0.267. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Why 4 of 30 tasks did not count

  • 4 — judge error
Gemma 3 27B google/gemma-3-27b-it 0.714 ± 0.020 0.694–0.733 100% 30/30 0.587 0.842 2
Per-dimension and per-judge breakdown for Gemma 3 27B
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
character differentiation 0.71
dramatic structure 0.68
economy and visual writing 0.65
format and constraint adherence 0.81
subtext and indirection 0.72

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.587
mistralai/mistral-large-2512 0.842

Mean gap between the two judges on this cell: 0.255. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Llama 4 Maverick meta-llama/llama-4-maverick 0.546 ± 0.025 0.520–0.571 100% 30/30 0.427 0.665 3
Per-dimension and per-judge breakdown for Llama 4 Maverick
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
character differentiation 0.49
dramatic structure 0.55
economy and visual writing 0.50
format and constraint adherence 0.69
subtext and indirection 0.51

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.427
mistralai/mistral-large-2512 0.665

Mean gap between the two judges on this cell: 0.238. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Llama 3.3 70B meta-llama/llama-3.3-70b-instruct 0.532 ± 0.021 0.512–0.554 100% 30/30 0.425 0.640 3
Per-dimension and per-judge breakdown for Llama 3.3 70B
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
character differentiation 0.48
dramatic structure 0.53
economy and visual writing 0.47
format and constraint adherence 0.67
subtext and indirection 0.51

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.425
mistralai/mistral-large-2512 0.640

Mean gap between the two judges on this cell: 0.215. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Rank bands are a conservative heuristic: two models share a band when their 95% intervals overlap. This is not a significance test and it does not correct for making 13 comparisons at once. Withheld rows were run — they are shown as verdicts, with the reason, rather than hidden.

The same task, three ways

Each row below is one prompt answered by three models, ordered by the score they received. The task is held constant on purpose: comparing outputs across different prompts would confound how hard the task was with how well the model wrote. Read them and decide whether the ordering was earned — that is a question you can answer without trusting us, which is the point of showing them.

How these were chosen. Of 190 tasks, 104 ( 55%) had both judges independently place the three outputs in the same order, and these are drawn from that pool; on 95% neither judge contradicted the order. Within a domain we take the widest score gaps, and no two examples repeat the same pair of best and worst models. We publish the denominator because "the judges agreed on every example we show" would otherwise be a fact about our editing rather than about our judges.

Two people who have been corresponding for years about a shared hobby meet for the first time, and one of them has been two people.

movie_scene/024 · same task, three models · spread 0.50

  1. Low 0.40 Llama 4 Maverick

    Per judge: mistral-large-2512 0.55 · glm-5.2 0.25

    FADE IN: INT. CONVENTION CENTER - DAY The room is filled with people enthusiastically discussing and displaying their shared hobby, model trains. We see JEN (30s), dressed in a bright yellow t-shirt with a model train logo, excitedly scanning the crowd. Her eyes lock onto a figure across the room, and a wide smile spreads across her face. JEN (whispering to herself) That's them. She makes her way through the crowd, weaving past model train displays and enthusiastic hobbyists. As she approaches, the figure turns, and we see MIKE (30s), dressed in a similar model train t-shirt, with a friendly smile. JEN (extending her hand) Mike! It's great to finally meet you.
  2. Middle 0.75 Grok 4.5

    Per judge: mistral-large-2512 0.80 · glm-5.2 0.70

    INT. BOTANICAL GARDEN CAFE - DAY Sunlight streams through tall glass walls overlooking manicured paths and flowering shrubs. The cafe hums with low conversation and the clink of ceramic. A few patrons linger over laptops and guidebooks. JAMIE (early 40s, kind eyes, soft flannel shirt) sits at a small wooden table near the window. An open sketchbook shows half-finished studies of bees on lavender. A camera and a well-worn field journal rest beside a cooling mug of tea. Jamie taps a pencil against the page, glances toward the entrance, then forces a breath and adds another careful wing vein to the drawing.
  3. High 0.90 Claude Opus 5

    Per judge: mistral-large-2512 0.90 · glm-5.2 0.90

    INT. RAILWAY BUFFET, CREWE STATION — EARLY AFTERNOON Rain sheets down the window. Beyond it, a wet platform, a tannoy nobody can decipher. Inside: formica, a fogged pastry case, six tables and only one occupied. DENIS COYNE, sixties, tweed gone shiny at the elbows, sits with two mugs of tea in front of him. A shoebox at his feet, lid held on with a rubber band. HARRIET FELL, late fifties, comes in shaking out an umbrella, scanning the room. She's carrying a leather satchel with two buckles, both done up. She sees the shoebox before she sees him. HARRIET Coyne. DENIS (standing, too fast, knocking the table) Fell. They shake hands. It goes on a beat longer than a handshake should.

Two people who used to work together meet by accident in an airport and have eleven minutes.

movie_scene/001 · same task, three models · spread 0.45

  1. Low 0.42 Llama 4 Maverick

    Per judge: mistral-large-2512 0.55 · glm-5.2 0.30

    INT. AIRPORT GATE AREA - DAY The gate area is bustling with travelers rushing to their gates. We see JEN (30s) and MIKE (30s) standing by a coffee shop, facing each other. They haven't seen each other in years, but the familiarity is still there. JEN (awkward laugh) Mike. MIKE (smirking) Jen. They stand there for a moment, unsure of what to do. Jen gestures to the coffee shop. JEN Coffee? MIKE (nodding) Sure. They walk into the coffee shop, joining the short line. As they wait, they make small talk. MIKE So, where are you headed? JEN (looking at her phone) Dallas. MIKE (following her gaze) Business? JEN (shrugging) Just a meeting. MIKE (nodding) Okay.
  2. Middle 0.78 Claude Opus 5

    Per judge: mistral-large-2512 0.85 · glm-5.2 0.70

    INT. DENVER INTERNATIONAL AIRPORT — CONCOURSE B — MIDDAY A wine bar with no walls. Rolling bags, a departures board flipping over itself. LAURA GESSNER, 50s, sits at the counter with a glass of white she hasn't touched and a paperback she isn't reading. MARTIN OYELARAN, 50s, passes with a briefcase, stops. Comes back. A lanyard around his neck: a logo that isn't the old one. MARTIN Laura. LAURA Oh my God. She stands. They do a hug that resolves into a handshake, then into nothing. MARTIN You're going where? LAURA Sacramento. Then a car. MARTIN Right. LAURA You? MARTIN Newark. I've got — He checks his phone.
  3. High 0.88 Kimi K3

    Per judge: mistral-large-2512 0.90 · glm-5.2 0.85

    **CONNECTION** ``` INT. AIRPORT TERMINAL - GATE C22 - DAY A departures board scrolls: DELAYED. DELAYED. ON TIME. MARA VOSS (40s, raincoat over her arm) stirs a coffee she isn't drinking. Three seats down, ELLIS GRAY (40s, suit jacket, no tie) looks up from his phone. He sees her. Looks back at his phone. Looks again. She is already looking at him. A beat that lasts one second longer than strangers. ELLIS Gate C22. MARA Gate C22. He half-rises. The seat beside her has her bag on it. She moves the bag. He crosses. Sits. The armrest between them stays down. At the podium, the GATE AGENT leans into a microphone.

A landlord and a tenant, both of whom need the other more than they will admit.

movie_scene/020 · same task, three models · spread 0.43

  1. Low 0.47 Llama 3.3 70B

    Per judge: mistral-large-2512 0.60 · glm-5.2 0.35

    INT. LANDLORD'S OFFICE - DAY The landlord, MR. JENKINS, a middle-aged man with a stern expression, sits behind a cluttered desk. The tenant, MS. RODRIGUEZ, a young woman with a determined look, stands opposite him. MR. JENKINS flips through a stack of papers I'm afraid I have no choice but to raise the rent, Ms. Rodriguez. The building's expenses have increased significantly. MS. RODRIGUEZ folds her arms I understand that, Mr. Jenkins, but I'm already paying a premium for this apartment. I don't think I can afford to pay more. MR. JENKINS skeptical I'm sure you'll find a way to make it work. After all, you've been a model tenant. I'd hate to have to find someone new.
  2. Middle 0.78 DeepSeek V4 Pro

    Per judge: mistral-large-2512 0.90 · glm-5.2 0.65

    ```screenplay INT. APARTMENT BUILDING HALLWAY – DUSK Faded floral carpet, a single bare bulb sputters overhead. WALTER (60s, weary, toolbox in hand) knocks on 4B. A pause. JULIA (30s, sharp-eyed) opens the door a few inches, sees the toolbox, and doesn’t move to let him in. WALTER Routine pipe check. The unit downstairs heard dripping. JULIA There’s no drip. WALTER I’ll be the judge of that. He shifts his weight, still not invited. Julia crosses her arms. JULIA You’ve never done a routine anything in the two years I’ve been here. WALTER Then you’re overdue. JULIA I’m all paid through the month, so if this is about money— WALTER (sharp, cutting her off) It’s not about that.
  3. High 0.90 Kimi K3

    Per judge: mistral-large-2512 0.95 · glm-5.2 0.85

    INT. BRICK WALK-UP — APARTMENT 4B — KITCHEN — NIGHT Prewar bones. A radiator ticks like a clock with a grudge. JUNE (45), still in nurse's scrubs, wrestles a wrench onto the valve. THREE RAPS at the door. Official. JUNE It's open. MAREK (72) enters. Overcoat buttoned on the wrong holes. He holds an envelope like a warrant. MAREK You should lock your door. JUNE You have a key. MAREK Then get a chain. He sets the envelope on the table between them. Doesn't sit. She looks at it. Doesn't touch it. JUNE If that's what I think it is— MAREK You're three months behind. JUNE Two. October, I paid. MAREK October was short.

Novel premises vs canonical

Roughly three tasks in ten use premises invented for this benchmark. If models score markedly higher on canonical premises, part of what this board measures is recall rather than craft. Published whatever it shows (FR-005).

Mean score by premise type, and the gap between them.
Scope Canonical Novel Delta n
movie_scene 0.764 0.749 +0.015 264/115