Movie Scene
| Model | Score ± CI | 95% interval | Counted | Tasks | glm-5.2 | mistral-large-2512 | Band | ||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 anthropic/claude-opus-5 | 0.861 ± 0.013 | 0.848–0.873 | 100% | 30/30 | 0.812 | 0.910 | 1 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Claude Opus 5
Per judge
Mean gap between the two judges on this cell: 0.098. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| Kimi K3 moonshotai/kimi-k3 | 0.858 ± 0.015 | 0.843–0.873 | 90% | 27/30 | 0.802 | 0.915 | 1 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Kimi K3
Per judge
Mean gap between the two judges on this cell: 0.113. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. Why 3 of 30 tasks did not count
| |||||||||||||||||||||||||
| Qwen3.8 Max qwen/qwen3.8-max | 0.830 ± 0.014 | 0.816–0.844 | 90% | 27/30 | 0.763 | 0.896 | 1 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Qwen3.8 Max
Per judge
Mean gap between the two judges on this cell: 0.133. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. Why 3 of 30 tasks did not count
| |||||||||||||||||||||||||
| GPT-5.6 Luna openai/gpt-5.6-luna | 0.824 ± 0.015 | 0.809–0.839 | 97% | 29/30 | 0.760 | 0.888 | 1 | ||||||||||||||||||
Per-dimension and per-judge breakdown for GPT-5.6 Luna
Per judge
Mean gap between the two judges on this cell: 0.128. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. Why 1 of 30 tasks did not count
| |||||||||||||||||||||||||
| GPT-5.6 Terra openai/gpt-5.6-terra | 0.823 ± 0.013 | 0.811–0.836 | 100% | 30/30 | 0.748 | 0.898 | 1 | ||||||||||||||||||
Per-dimension and per-judge breakdown for GPT-5.6 Terra
Per judge
Mean gap between the two judges on this cell: 0.150. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| Grok 4.5 x-ai/grok-4.5 | 0.813 ± 0.017 | 0.796–0.829 | 100% | 30/30 | 0.732 | 0.893 | 1 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Grok 4.5
Per judge
Mean gap between the two judges on this cell: 0.162. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| Gemini 3.6 Flash google/gemini-3.6-flash | 0.812 ± 0.013 | 0.799–0.825 | 100% | 30/30 | 0.730 | 0.893 | 1 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Gemini 3.6 Flash
Per judge
Mean gap between the two judges on this cell: 0.163. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| DeepSeek V4 Pro deepseek/deepseek-v4-pro | 0.790 ± 0.016 | 0.773–0.805 | 100% | 30/30 | 0.697 | 0.883 | 1 | ||||||||||||||||||
Per-dimension and per-judge breakdown for DeepSeek V4 Pro
Per judge
Mean gap between the two judges on this cell: 0.187. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| GLM 5 z-ai/glm-5 SELF-FAMILY: z-ai | 0.765 ± 0.013 | 0.752–0.778 | 100% | 30/30 | 0.658 | 0.872 | 1 | ||||||||||||||||||
Per-dimension and per-judge breakdown for GLM 5
Per judge
Mean gap between the two judges on this cell: 0.213. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. One judge (z-ai/glm-5.2) shares this model's family. We disclose the conflict rather than dropping either the model or the judge. | |||||||||||||||||||||||||
| Qwen3 32B qwen/qwen3-32b | 0.718 ± 0.018 | 0.700–0.736 | 87% | 26/30 | 0.585 | 0.852 | 2 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Qwen3 32B
Per judge
Mean gap between the two judges on this cell: 0.267. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. Why 4 of 30 tasks did not count
| |||||||||||||||||||||||||
| Gemma 3 27B google/gemma-3-27b-it | 0.714 ± 0.020 | 0.694–0.733 | 100% | 30/30 | 0.587 | 0.842 | 2 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Gemma 3 27B
Per judge
Mean gap between the two judges on this cell: 0.255. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| Llama 4 Maverick meta-llama/llama-4-maverick | 0.546 ± 0.025 | 0.520–0.571 | 100% | 30/30 | 0.427 | 0.665 | 3 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Llama 4 Maverick
Per judge
Mean gap between the two judges on this cell: 0.238. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| Llama 3.3 70B meta-llama/llama-3.3-70b-instruct | 0.532 ± 0.021 | 0.512–0.554 | 100% | 30/30 | 0.425 | 0.640 | 3 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Llama 3.3 70B
Per judge
Mean gap between the two judges on this cell: 0.215. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
The same task, three ways
Each row below is one prompt answered by three models, ordered by the score they received. The task is held constant on purpose: comparing outputs across different prompts would confound how hard the task was with how well the model wrote. Read them and decide whether the ordering was earned — that is a question you can answer without trusting us, which is the point of showing them.
How these were chosen. Of 190 tasks, 104 ( 55%) had both judges independently place the three outputs in the same order, and these are drawn from that pool; on 95% neither judge contradicted the order. Within a domain we take the widest score gaps, and no two examples repeat the same pair of best and worst models. We publish the denominator because "the judges agreed on every example we show" would otherwise be a fact about our editing rather than about our judges.
Two people who have been corresponding for years about a shared hobby meet for the first time, and one of them has been two people.
-
Low 0.40 Llama 4 Maverick
Per judge: mistral-large-2512 0.55 · glm-5.2 0.25
FADE IN: INT. CONVENTION CENTER - DAY The room is filled with people enthusiastically discussing and displaying their shared hobby, model trains. We see JEN (30s), dressed in a bright yellow t-shirt with a model train logo, excitedly scanning the crowd. Her eyes lock onto a figure across the room, and a wide smile spreads across her face. JEN (whispering to herself) That's them. She makes her way through the crowd, weaving past model train displays and enthusiastic hobbyists. As she approaches, the figure turns, and we see MIKE (30s), dressed in a similar model train t-shirt, with a friendly smile. JEN (extending her hand) Mike! It's great to finally meet you.
-
Middle 0.75 Grok 4.5
Per judge: mistral-large-2512 0.80 · glm-5.2 0.70
INT. BOTANICAL GARDEN CAFE - DAY Sunlight streams through tall glass walls overlooking manicured paths and flowering shrubs. The cafe hums with low conversation and the clink of ceramic. A few patrons linger over laptops and guidebooks. JAMIE (early 40s, kind eyes, soft flannel shirt) sits at a small wooden table near the window. An open sketchbook shows half-finished studies of bees on lavender. A camera and a well-worn field journal rest beside a cooling mug of tea. Jamie taps a pencil against the page, glances toward the entrance, then forces a breath and adds another careful wing vein to the drawing.
-
High 0.90 Claude Opus 5
Per judge: mistral-large-2512 0.90 · glm-5.2 0.90
INT. RAILWAY BUFFET, CREWE STATION — EARLY AFTERNOON Rain sheets down the window. Beyond it, a wet platform, a tannoy nobody can decipher. Inside: formica, a fogged pastry case, six tables and only one occupied. DENIS COYNE, sixties, tweed gone shiny at the elbows, sits with two mugs of tea in front of him. A shoebox at his feet, lid held on with a rubber band. HARRIET FELL, late fifties, comes in shaking out an umbrella, scanning the room. She's carrying a leather satchel with two buckles, both done up. She sees the shoebox before she sees him. HARRIET Coyne. DENIS (standing, too fast, knocking the table) Fell. They shake hands. It goes on a beat longer than a handshake should.
Two people who used to work together meet by accident in an airport and have eleven minutes.
-
Low 0.42 Llama 4 Maverick
Per judge: mistral-large-2512 0.55 · glm-5.2 0.30
INT. AIRPORT GATE AREA - DAY The gate area is bustling with travelers rushing to their gates. We see JEN (30s) and MIKE (30s) standing by a coffee shop, facing each other. They haven't seen each other in years, but the familiarity is still there. JEN (awkward laugh) Mike. MIKE (smirking) Jen. They stand there for a moment, unsure of what to do. Jen gestures to the coffee shop. JEN Coffee? MIKE (nodding) Sure. They walk into the coffee shop, joining the short line. As they wait, they make small talk. MIKE So, where are you headed? JEN (looking at her phone) Dallas. MIKE (following her gaze) Business? JEN (shrugging) Just a meeting. MIKE (nodding) Okay.
-
Middle 0.78 Claude Opus 5
Per judge: mistral-large-2512 0.85 · glm-5.2 0.70
INT. DENVER INTERNATIONAL AIRPORT — CONCOURSE B — MIDDAY A wine bar with no walls. Rolling bags, a departures board flipping over itself. LAURA GESSNER, 50s, sits at the counter with a glass of white she hasn't touched and a paperback she isn't reading. MARTIN OYELARAN, 50s, passes with a briefcase, stops. Comes back. A lanyard around his neck: a logo that isn't the old one. MARTIN Laura. LAURA Oh my God. She stands. They do a hug that resolves into a handshake, then into nothing. MARTIN You're going where? LAURA Sacramento. Then a car. MARTIN Right. LAURA You? MARTIN Newark. I've got — He checks his phone.
-
High 0.88 Kimi K3
Per judge: mistral-large-2512 0.90 · glm-5.2 0.85
**CONNECTION** ``` INT. AIRPORT TERMINAL - GATE C22 - DAY A departures board scrolls: DELAYED. DELAYED. ON TIME. MARA VOSS (40s, raincoat over her arm) stirs a coffee she isn't drinking. Three seats down, ELLIS GRAY (40s, suit jacket, no tie) looks up from his phone. He sees her. Looks back at his phone. Looks again. She is already looking at him. A beat that lasts one second longer than strangers. ELLIS Gate C22. MARA Gate C22. He half-rises. The seat beside her has her bag on it. She moves the bag. He crosses. Sits. The armrest between them stays down. At the podium, the GATE AGENT leans into a microphone.
A landlord and a tenant, both of whom need the other more than they will admit.
-
Low 0.47 Llama 3.3 70B
Per judge: mistral-large-2512 0.60 · glm-5.2 0.35
INT. LANDLORD'S OFFICE - DAY The landlord, MR. JENKINS, a middle-aged man with a stern expression, sits behind a cluttered desk. The tenant, MS. RODRIGUEZ, a young woman with a determined look, stands opposite him. MR. JENKINS flips through a stack of papers I'm afraid I have no choice but to raise the rent, Ms. Rodriguez. The building's expenses have increased significantly. MS. RODRIGUEZ folds her arms I understand that, Mr. Jenkins, but I'm already paying a premium for this apartment. I don't think I can afford to pay more. MR. JENKINS skeptical I'm sure you'll find a way to make it work. After all, you've been a model tenant. I'd hate to have to find someone new.
-
Middle 0.78 DeepSeek V4 Pro
Per judge: mistral-large-2512 0.90 · glm-5.2 0.65
```screenplay INT. APARTMENT BUILDING HALLWAY – DUSK Faded floral carpet, a single bare bulb sputters overhead. WALTER (60s, weary, toolbox in hand) knocks on 4B. A pause. JULIA (30s, sharp-eyed) opens the door a few inches, sees the toolbox, and doesn’t move to let him in. WALTER Routine pipe check. The unit downstairs heard dripping. JULIA There’s no drip. WALTER I’ll be the judge of that. He shifts his weight, still not invited. Julia crosses her arms. JULIA You’ve never done a routine anything in the two years I’ve been here. WALTER Then you’re overdue. JULIA I’m all paid through the month, so if this is about money— WALTER (sharp, cutting her off) It’s not about that.
-
High 0.90 Kimi K3
Per judge: mistral-large-2512 0.95 · glm-5.2 0.85
INT. BRICK WALK-UP — APARTMENT 4B — KITCHEN — NIGHT Prewar bones. A radiator ticks like a clock with a grudge. JUNE (45), still in nurse's scrubs, wrestles a wrench onto the valve. THREE RAPS at the door. Official. JUNE It's open. MAREK (72) enters. Overcoat buttoned on the wrong holes. He holds an envelope like a warrant. MAREK You should lock your door. JUNE You have a key. MAREK Then get a chain. He sets the envelope on the table between them. Doesn't sit. She looks at it. Doesn't touch it. JUNE If that's what I think it is— MAREK You're three months behind. JUNE Two. October, I paid. MAREK October was short.
Novel premises vs canonical
Roughly three tasks in ten use premises invented for this benchmark. If models score markedly higher on canonical premises, part of what this board measures is recall rather than craft. Published whatever it shows (FR-005).
| Scope | Canonical | Novel | Delta | n |
|---|---|---|---|---|
| movie_scene | 0.764 | 0.749 | +0.015 | 264/115 |