Stand-Up Set
This board scores craft, not funniness. No rubric dimension asks a judge whether something is funny — the comedic dimensions ask whether the humour uses the mechanism the show or form actually uses.
| Model | Score ± CI | 95% interval | Counted | Tasks | glm-5.2 | mistral-large-2512 | Band | ||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 anthropic/claude-opus-5 | 0.834 ± 0.015 | 0.819–0.850 | 100% | 30/30 | 0.790 | 0.878 | 1 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Claude Opus 5
Per judge
Mean gap between the two judges on this cell: 0.105. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| Kimi K3 moonshotai/kimi-k3 | 0.825 ± 0.013 | 0.812–0.839 | 97% | 29/30 | 0.772 | 0.878 | 1 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Kimi K3
Per judge
Mean gap between the two judges on this cell: 0.116. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. Why 1 of 30 tasks did not count
| |||||||||||||||||||||||||
| Qwen3.8 Max qwen/qwen3.8-max | 0.792 ± 0.008 | 0.784–0.799 | 100% | 30/30 | 0.750 | 0.833 | 2 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Qwen3.8 Max
Per judge
Mean gap between the two judges on this cell: 0.083. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| DeepSeek V4 Pro deepseek/deepseek-v4-pro | 0.783 ± 0.011 | 0.772–0.794 | 100% | 30/30 | 0.730 | 0.837 | 2 | ||||||||||||||||||
Per-dimension and per-judge breakdown for DeepSeek V4 Pro
Per judge
Mean gap between the two judges on this cell: 0.107. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| Gemini 3.6 Flash google/gemini-3.6-flash | 0.783 ± 0.012 | 0.770–0.795 | 100% | 30/30 | 0.730 | 0.835 | 2 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Gemini 3.6 Flash
Per judge
Mean gap between the two judges on this cell: 0.105. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| GPT-5.6 Terra openai/gpt-5.6-terra | 0.770 ± 0.010 | 0.760–0.779 | 100% | 30/30 | 0.732 | 0.808 | 2 | ||||||||||||||||||
Per-dimension and per-judge breakdown for GPT-5.6 Terra
Per judge
Mean gap between the two judges on this cell: 0.077. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| Grok 4.5 x-ai/grok-4.5 | 0.764 ± 0.015 | 0.748–0.779 | 100% | 30/30 | 0.712 | 0.817 | 2 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Grok 4.5
Per judge
Mean gap between the two judges on this cell: 0.105. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| GPT-5.6 Luna openai/gpt-5.6-luna | 0.758 ± 0.012 | 0.746–0.771 | 100% | 30/30 | 0.717 | 0.800 | 2 | ||||||||||||||||||
Per-dimension and per-judge breakdown for GPT-5.6 Luna
Per judge
Mean gap between the two judges on this cell: 0.083. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| GLM 5 z-ai/glm-5 SELF-FAMILY: z-ai | 0.743 ± 0.015 | 0.728–0.758 | 100% | 30/30 | 0.697 | 0.790 | 2 | ||||||||||||||||||
Per-dimension and per-judge breakdown for GLM 5
Per judge
Mean gap between the two judges on this cell: 0.093. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. One judge (z-ai/glm-5.2) shares this model's family. We disclose the conflict rather than dropping either the model or the judge. | |||||||||||||||||||||||||
| Gemma 3 27B google/gemma-3-27b-it | 0.707 ± 0.016 | 0.690–0.723 | 100% | 30/30 | 0.637 | 0.777 | 3 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Gemma 3 27B
Per judge
Mean gap between the two judges on this cell: 0.143. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| Qwen3 32B qwen/qwen3-32b | 0.690 ± 0.017 | 0.673–0.707 | 100% | 30/30 | 0.595 | 0.785 | 3 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Qwen3 32B
Per judge
Mean gap between the two judges on this cell: 0.190. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. | |||||||||||||||||||||||||
| Llama 4 Maverick meta-llama/llama-4-maverick | 0.555 ± 0.012 | 0.542–0.567 | 97% | 29/30 | 0.448 | 0.662 | 4 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Llama 4 Maverick
Per judge
Mean gap between the two judges on this cell: 0.214. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. Why 1 of 30 tasks did not count
| |||||||||||||||||||||||||
| Llama 3.3 70B meta-llama/llama-3.3-70b-instruct | 0.540 ± 0.015 | 0.525–0.554 | 97% | 29/30 | 0.433 | 0.647 | 4 | ||||||||||||||||||
Per-dimension and per-judge breakdown for Llama 3.3 70B
Per judge
Mean gap between the two judges on this cell: 0.214. One anchor step is 0.25, so a gap near or above that means the judges were not reading the same descriptor — published because it bounds how much the score can bear. Why 1 of 30 tasks did not count
| |||||||||||||||||||||||||
The same task, three ways
Each row below is one prompt answered by three models, ordered by the score they received. The task is held constant on purpose: comparing outputs across different prompts would confound how hard the task was with how well the model wrote. Read them and decide whether the ordering was earned — that is a question you can answer without trusting us, which is the point of showing them.
How these were chosen. Of 190 tasks, 104 ( 55%) had both judges independently place the three outputs in the same order, and these are drawn from that pool; on 95% neither judge contradicted the order. Within a domain we take the widest score gaps, and no two examples repeat the same pair of best and worst models. We publish the denominator because "the judges agreed on every example we show" would otherwise be a fact about our editing rather than about our judges.
A bit about the moment you realise you have been talking to a stranger who thinks you are someone else.
-
Low 0.50 Llama 3.3 70B
Per judge: mistral-large-2512 0.60 · glm-5.2 0.40
You know, I was at the coffee shop the other day, and I had one of those moments where you're chatting with someone, and it's all going great, until you realize they have no idea who you are. But not just that, they think you're someone else entirely. It's like, you're having a conversation, and they're like, "Oh, yeah, I loved your last album," and you're like, "Uh, I don't have an album." And they're like, "What, you don't remember? It was a huge hit!" And you're like, "Dude, I'm a accountant, not a rockstar."
-
Middle 0.78 Gemini 3.6 Flash
Per judge: mistral-large-2512 0.80 · glm-5.2 0.75
Have you ever been five minutes deep into a conversation before you realize the other person has no idea who you are? Not just "Oh, we met at a party once," but full-on, "I think you are my cousin Marcus, and we buried a lawnmower in 2011." It happened to me at Home Depot last week. I was in the plumbing aisle, holding a PVC pipe I had no business touching, when this guy in a fleece vest grabs my shoulder and goes, "Marcus! Thank God! How’s the shoulder holding up after the hearing?"
-
High 0.93 Kimi K3
Per judge: mistral-large-2512 0.90 · glm-5.2 0.95
I was at a wedding — friend of a friend, I knew maybe four people — and I'm at the bar when this man crosses the room with his arms already open, walking like I'm the last helicopter out of Saigon. He grabs me by both shoulders and goes, [as him, booming, shaking my shoulders, overjoyed]: "DEREK! You absolute criminal — you came!" And I knew within half a second that I was not Derek. But he had already spent so much joy, and there was no gap in the conversation to hand it back. You can't hand joy back. That's like telling a dog it's got the wrong house.
A bit about pretending to have read something in a conversation.
-
Low 0.53 Llama 3.3 70B
Per judge: mistral-large-2512 0.65 · glm-5.2 0.40
You know, I was at a dinner party the other night, and someone mentioned this super obscure novel, "The Ephemeral Nature of Reality" or something like that. And I'm thinking, "Oh, yeah, I've totally read that." I mean, who hasn't, right? (chuckles) So, I start nodding along, making these thoughtful noises, like "hmm, yes, the themes of existential dread really resonated with me." And the other person is eating it up, asking me questions about the plot and the characters. I'm like, "Oh, absolutely, the protagonist's journey is a powerful allegory for the human condition." (act-out: pretending to be deep in thought)
-
Middle 0.78 DeepSeek V4 Pro
Per judge: mistral-large-2512 0.80 · glm-5.2 0.75
Pretending you’ve read something is never just one lie. It’s a baby giraffe born on a frozen pond—it’s immediately trying to stand, and it’s doomed. I did it because a woman I liked mentioned *Infinite Jest*. And I said, “Oh, yes… the tennis?” She goes, “You read it?” I said, “Twice.” She was impressed, but she’s a PhD candidate in comp lit, so she follows up. “What did you think of the narrative structure intersecting with Hal’s silent scream?” I said, “It’s iconic. It’s just… screams into the void, you know?” She nodded. Then she invited me to lead her book club discussion. *Lead.* I am now the keynote speaker for a book I haven’t cracked. So I buy the book. It’s the size of a microwave.
-
High 0.95 Claude Opus 5
Per judge: mistral-large-2512 0.95 · glm-5.2 0.95
Somewhere in my body there is a muscle that answers questions before my brain arrives. It's not lying. Lying takes planning. This is closer to a sneeze. So I'm at a party, holding a beer like a hostage, and a woman says the name of a novel. Hungarian. Four hundred pages. And my mouth goes, *oh, that book ruined me.* Ruined me! Past tense! Confidently! And her whole face opens up like a garage door, because she has not met another person who has read this book, and she has been waiting years, and now here I am, her guy.
A bit about the specific etiquette of a shared office fridge.
-
Low 0.53 Llama 4 Maverick
Per judge: mistral-large-2512 0.65 · glm-5.2 0.40
You know what's wild? The office fridge. It's like a societal microcosm in there. (pauses for comedic effect) You've got your leftovers, your lunches, your Tupperware... it's like a mini United Nations. Except instead of diplomacy, you've got mystery meat and a lingering sense of dread. "Is that a science experiment in there?" (chuckles) Now, I know what you're thinking. "What's the big deal? It's just a fridge." But let me tell you, the etiquette is real. You gotta label your food, or it's like you're declaring war on your coworkers. "This is mine, I swear, I brought it in on Tuesday...
-
Middle 0.78 GPT-5.6 Terra
Per judge: mistral-large-2512 0.80 · glm-5.2 0.75
The shared office fridge is where adults go to discover they have no faith in democracy. Because everybody comes in with the same little speech: “We’re all professionals here.” No, we’re not. We’re animals with access to hummus. There’s always one container in there that’s been abandoned so long it has its own weather system. It’s got a Post-it on it that says “DO NOT EAT,” dated three fiscal quarters ago. At that point, nobody wants to eat it. We want to know if it’s still legally food. And people label things like they’re preparing evidence for trial. “Karen’s yogurt. 3/12. Please respect.” Please respect? That yogurt has more boundaries than Karen’s entire management style.
-
High 0.93 Claude Opus 5
Per judge: mistral-large-2512 0.95 · glm-5.2 0.90
So somebody's eating my yogurt. Not "somebody ate my yogurt." Present tense. Ongoing. This is a series, not a pilot. And you can't say anything, right? Because there's no proportionality in the office. If I say nothing, I'm losing four dollars a week to a stranger. If I say something, I'm the guy with the yogurt speech. Forever. I could cure something. They'd say, "That's the yogurt guy, he cured something, but he did do the yogurt thing."
Novel premises vs canonical
Roughly three tasks in ten use premises invented for this benchmark. If models score markedly higher on canonical premises, part of what this board measures is recall rather than craft. Published whatever it shows (FR-005).
| Scope | Canonical | Novel | Delta | n |
|---|---|---|---|---|
| standup | 0.729 | 0.748 | -0.019 | 271/116 |