CANONIC

Stand-Up Set

This board scores craft, not funniness. No rubric dimension asks a judge whether something is funny — the comedic dimensions ask whether the humour uses the mechanism the show or form actually uses.

Stand-Up Set — 13 models, scored against the published standup rubric. Scores craft, not funniness. Models whose intervals overlap share a rank band and are not ranked against each other.
Model Score ± CI 95% interval Counted Tasks glm-5.2mistral-large-2512 Band
Claude Opus 5 anthropic/claude-opus-5 0.834 ± 0.015 0.819–0.850 100% 30/30 0.790 0.878 1
Per-dimension and per-judge breakdown for Claude Opus 5
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
economy of language 0.78
escalation and structure 0.79
premise development 0.86
specificity and observation 0.88
voice and persona 0.86

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.790
mistralai/mistral-large-2512 0.878

Mean gap between the two judges on this cell: 0.105. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Kimi K3 moonshotai/kimi-k3 0.825 ± 0.013 0.812–0.839 97% 29/30 0.772 0.878 1
Per-dimension and per-judge breakdown for Kimi K3
Mean anchor per rubric dimension, over 29 counted tasks.
Mean anchor per rubric dimension, over 29 counted tasks. — data
Dimension Mean
economy of language 0.77
escalation and structure 0.79
premise development 0.85
specificity and observation 0.84
voice and persona 0.86

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.772
mistralai/mistral-large-2512 0.878

Mean gap between the two judges on this cell: 0.116. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Why 1 of 30 tasks did not count

  • 1 — judge error
Qwen3.8 Max qwen/qwen3.8-max 0.792 ± 0.008 0.784–0.799 100% 30/30 0.750 0.833 2
Per-dimension and per-judge breakdown for Qwen3.8 Max
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
economy of language 0.76
escalation and structure 0.76
premise development 0.80
specificity and observation 0.81
voice and persona 0.82

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.750
mistralai/mistral-large-2512 0.833

Mean gap between the two judges on this cell: 0.083. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

DeepSeek V4 Pro deepseek/deepseek-v4-pro 0.783 ± 0.011 0.772–0.794 100% 30/30 0.730 0.837 2
Per-dimension and per-judge breakdown for DeepSeek V4 Pro
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
economy of language 0.74
escalation and structure 0.75
premise development 0.81
specificity and observation 0.80
voice and persona 0.82

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.730
mistralai/mistral-large-2512 0.837

Mean gap between the two judges on this cell: 0.107. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Gemini 3.6 Flash google/gemini-3.6-flash 0.783 ± 0.012 0.770–0.795 100% 30/30 0.730 0.835 2
Per-dimension and per-judge breakdown for Gemini 3.6 Flash
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
economy of language 0.75
escalation and structure 0.74
premise development 0.80
specificity and observation 0.81
voice and persona 0.82

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.730
mistralai/mistral-large-2512 0.835

Mean gap between the two judges on this cell: 0.105. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

GPT-5.6 Terra openai/gpt-5.6-terra 0.770 ± 0.010 0.760–0.779 100% 30/30 0.732 0.808 2
Per-dimension and per-judge breakdown for GPT-5.6 Terra
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
economy of language 0.75
escalation and structure 0.73
premise development 0.76
specificity and observation 0.80
voice and persona 0.81

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.732
mistralai/mistral-large-2512 0.808

Mean gap between the two judges on this cell: 0.077. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Grok 4.5 x-ai/grok-4.5 0.764 ± 0.015 0.748–0.779 100% 30/30 0.712 0.817 2
Per-dimension and per-judge breakdown for Grok 4.5
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
economy of language 0.72
escalation and structure 0.75
premise development 0.78
specificity and observation 0.78
voice and persona 0.80

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.712
mistralai/mistral-large-2512 0.817

Mean gap between the two judges on this cell: 0.105. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

GPT-5.6 Luna openai/gpt-5.6-luna 0.758 ± 0.012 0.746–0.771 100% 30/30 0.717 0.800 2
Per-dimension and per-judge breakdown for GPT-5.6 Luna
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
economy of language 0.74
escalation and structure 0.72
premise development 0.77
specificity and observation 0.78
voice and persona 0.78

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.717
mistralai/mistral-large-2512 0.800

Mean gap between the two judges on this cell: 0.083. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

GLM 5 z-ai/glm-5 SELF-FAMILY: z-ai 0.743 ± 0.015 0.728–0.758 100% 30/30 0.697 0.790 2
Per-dimension and per-judge breakdown for GLM 5
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
economy of language 0.72
escalation and structure 0.73
premise development 0.75
specificity and observation 0.74
voice and persona 0.78

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.697
mistralai/mistral-large-2512 0.790

Mean gap between the two judges on this cell: 0.093. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

One judge (z-ai/glm-5.2) shares this model's family. We disclose the conflict rather than dropping either the model or the judge.

Gemma 3 27B google/gemma-3-27b-it 0.707 ± 0.016 0.690–0.723 100% 30/30 0.637 0.777 3
Per-dimension and per-judge breakdown for Gemma 3 27B
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
economy of language 0.64
escalation and structure 0.69
premise development 0.71
specificity and observation 0.72
voice and persona 0.77

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.637
mistralai/mistral-large-2512 0.777

Mean gap between the two judges on this cell: 0.143. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Qwen3 32B qwen/qwen3-32b 0.690 ± 0.017 0.673–0.707 100% 30/30 0.595 0.785 3
Per-dimension and per-judge breakdown for Qwen3 32B
Mean anchor per rubric dimension, over 30 counted tasks.
Mean anchor per rubric dimension, over 30 counted tasks. — data
Dimension Mean
economy of language 0.61
escalation and structure 0.68
premise development 0.68
specificity and observation 0.68
voice and persona 0.79

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.595
mistralai/mistral-large-2512 0.785

Mean gap between the two judges on this cell: 0.190. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Llama 4 Maverick meta-llama/llama-4-maverick 0.555 ± 0.012 0.542–0.567 97% 29/30 0.448 0.662 4
Per-dimension and per-judge breakdown for Llama 4 Maverick
Mean anchor per rubric dimension, over 29 counted tasks.
Mean anchor per rubric dimension, over 29 counted tasks. — data
Dimension Mean
economy of language 0.44
escalation and structure 0.61
premise development 0.61
specificity and observation 0.49
voice and persona 0.63

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.448
mistralai/mistral-large-2512 0.662

Mean gap between the two judges on this cell: 0.214. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Why 1 of 30 tasks did not count

  • 1 — judge error
Llama 3.3 70B meta-llama/llama-3.3-70b-instruct 0.540 ± 0.015 0.525–0.554 97% 29/30 0.433 0.647 4
Per-dimension and per-judge breakdown for Llama 3.3 70B
Mean anchor per rubric dimension, over 29 counted tasks.
Mean anchor per rubric dimension, over 29 counted tasks. — data
Dimension Mean
economy of language 0.41
escalation and structure 0.59
premise development 0.61
specificity and observation 0.46
voice and persona 0.63

Per judge

Mean score from each judge, and the mean absolute gap between them.
Judge Mean
z-ai/glm-5.2 0.433
mistralai/mistral-large-2512 0.647

Mean gap between the two judges on this cell: 0.214. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear.

Why 1 of 30 tasks did not count

  • 1 — judge error
Rank bands are a conservative heuristic: two models share a band when their 95% intervals overlap. This is not a significance test and it does not correct for making 13 comparisons at once. Withheld rows were run — they are shown as verdicts, with the reason, rather than hidden.

The same task, three ways

Each row below is one prompt answered by three models, ordered by the score they received. The task is held constant on purpose: comparing outputs across different prompts would confound how hard the task was with how well the model wrote. Read them and decide whether the ordering was earned — that is a question you can answer without trusting us, which is the point of showing them.

How these were chosen. Of 190 tasks, 104 ( 55%) had both judges independently place the three outputs in the same order, and these are drawn from that pool; on 95% neither judge contradicted the order. Within a domain we take the widest score gaps, and no two examples repeat the same pair of best and worst models. We publish the denominator because "the judges agreed on every example we show" would otherwise be a fact about our editing rather than about our judges.

A bit about the moment you realise you have been talking to a stranger who thinks you are someone else.

standup/004 · same task, three models · spread 0.43

  1. Low 0.50 Llama 3.3 70B

    Per judge: mistral-large-2512 0.60 · glm-5.2 0.40

    You know, I was at the coffee shop the other day, and I had one of those moments where you're chatting with someone, and it's all going great, until you realize they have no idea who you are. But not just that, they think you're someone else entirely. It's like, you're having a conversation, and they're like, "Oh, yeah, I loved your last album," and you're like, "Uh, I don't have an album." And they're like, "What, you don't remember? It was a huge hit!" And you're like, "Dude, I'm a accountant, not a rockstar."
  2. Middle 0.78 Gemini 3.6 Flash

    Per judge: mistral-large-2512 0.80 · glm-5.2 0.75

    Have you ever been five minutes deep into a conversation before you realize the other person has no idea who you are? Not just "Oh, we met at a party once," but full-on, "I think you are my cousin Marcus, and we buried a lawnmower in 2011." It happened to me at Home Depot last week. I was in the plumbing aisle, holding a PVC pipe I had no business touching, when this guy in a fleece vest grabs my shoulder and goes, "Marcus! Thank God! How’s the shoulder holding up after the hearing?"
  3. High 0.93 Kimi K3

    Per judge: mistral-large-2512 0.90 · glm-5.2 0.95

    I was at a wedding — friend of a friend, I knew maybe four people — and I'm at the bar when this man crosses the room with his arms already open, walking like I'm the last helicopter out of Saigon. He grabs me by both shoulders and goes, [as him, booming, shaking my shoulders, overjoyed]: "DEREK! You absolute criminal — you came!" And I knew within half a second that I was not Derek. But he had already spent so much joy, and there was no gap in the conversation to hand it back. You can't hand joy back. That's like telling a dog it's got the wrong house.

A bit about pretending to have read something in a conversation.

standup/011 · same task, three models · spread 0.42

  1. Low 0.53 Llama 3.3 70B

    Per judge: mistral-large-2512 0.65 · glm-5.2 0.40

    You know, I was at a dinner party the other night, and someone mentioned this super obscure novel, "The Ephemeral Nature of Reality" or something like that. And I'm thinking, "Oh, yeah, I've totally read that." I mean, who hasn't, right? (chuckles) So, I start nodding along, making these thoughtful noises, like "hmm, yes, the themes of existential dread really resonated with me." And the other person is eating it up, asking me questions about the plot and the characters. I'm like, "Oh, absolutely, the protagonist's journey is a powerful allegory for the human condition." (act-out: pretending to be deep in thought)
  2. Middle 0.78 DeepSeek V4 Pro

    Per judge: mistral-large-2512 0.80 · glm-5.2 0.75

    Pretending you’ve read something is never just one lie. It’s a baby giraffe born on a frozen pond—it’s immediately trying to stand, and it’s doomed. I did it because a woman I liked mentioned *Infinite Jest*. And I said, “Oh, yes… the tennis?” She goes, “You read it?” I said, “Twice.” She was impressed, but she’s a PhD candidate in comp lit, so she follows up. “What did you think of the narrative structure intersecting with Hal’s silent scream?” I said, “It’s iconic. It’s just… screams into the void, you know?” She nodded. Then she invited me to lead her book club discussion. *Lead.* I am now the keynote speaker for a book I haven’t cracked. So I buy the book. It’s the size of a microwave.
  3. High 0.95 Claude Opus 5

    Per judge: mistral-large-2512 0.95 · glm-5.2 0.95

    Somewhere in my body there is a muscle that answers questions before my brain arrives. It's not lying. Lying takes planning. This is closer to a sneeze. So I'm at a party, holding a beer like a hostage, and a woman says the name of a novel. Hungarian. Four hundred pages. And my mouth goes, *oh, that book ruined me.* Ruined me! Past tense! Confidently! And her whole face opens up like a garage door, because she has not met another person who has read this book, and she has been waiting years, and now here I am, her guy.

A bit about the specific etiquette of a shared office fridge.

standup/010 · same task, three models · spread 0.40

  1. Low 0.53 Llama 4 Maverick

    Per judge: mistral-large-2512 0.65 · glm-5.2 0.40

    You know what's wild? The office fridge. It's like a societal microcosm in there. (pauses for comedic effect) You've got your leftovers, your lunches, your Tupperware... it's like a mini United Nations. Except instead of diplomacy, you've got mystery meat and a lingering sense of dread. "Is that a science experiment in there?" (chuckles) Now, I know what you're thinking. "What's the big deal? It's just a fridge." But let me tell you, the etiquette is real. You gotta label your food, or it's like you're declaring war on your coworkers. "This is mine, I swear, I brought it in on Tuesday...
  2. Middle 0.78 GPT-5.6 Terra

    Per judge: mistral-large-2512 0.80 · glm-5.2 0.75

    The shared office fridge is where adults go to discover they have no faith in democracy. Because everybody comes in with the same little speech: “We’re all professionals here.” No, we’re not. We’re animals with access to hummus. There’s always one container in there that’s been abandoned so long it has its own weather system. It’s got a Post-it on it that says “DO NOT EAT,” dated three fiscal quarters ago. At that point, nobody wants to eat it. We want to know if it’s still legally food. And people label things like they’re preparing evidence for trial. “Karen’s yogurt. 3/12. Please respect.” Please respect? That yogurt has more boundaries than Karen’s entire management style.
  3. High 0.93 Claude Opus 5

    Per judge: mistral-large-2512 0.95 · glm-5.2 0.90

    So somebody's eating my yogurt. Not "somebody ate my yogurt." Present tense. Ongoing. This is a series, not a pilot. And you can't say anything, right? Because there's no proportionality in the office. If I say nothing, I'm losing four dollars a week to a stranger. If I say something, I'm the guy with the yogurt speech. Forever. I could cure something. They'd say, "That's the yogurt guy, he cured something, but he did do the yogurt thing."

Novel premises vs canonical

Roughly three tasks in ten use premises invented for this benchmark. If models score markedly higher on canonical premises, part of what this board measures is recall rather than craft. Published whatever it shows (FR-005).

Mean score by premise type, and the gap between them.
Scope Canonical Novel Delta n
standup 0.729 0.748 -0.019 271/116