CANONIC

Six domains.
No overall score.

Every model below was scored against a rubric frozen before the sweep ran, by two judges from different families, on tasks that were all attempted or not published at all. Where the evidence does not separate two models, they share a band and we say so.

78 rows ran. 1 are withheld — shown as verdicts with their reason, not hidden. That ratio is the most honest number on this page.


Results by domain

Sitcom Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Sitcom Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.875 (0.865–0.886) — top band
Kimi K3 0.864 (0.855–0.873) — top band
GPT-5.6 Terra 0.833 (0.820–0.844)
Gemini 3.6 Flash 0.831 (0.821–0.840)
Qwen3.8 Max 0.831 (0.815–0.843)
GPT-5.6 Luna 0.830 (0.815–0.842)
GLM 5 0.820 (0.808–0.830)
DeepSeek V4 Pro 0.818 (0.803–0.830)
Grok 4.5 0.809 (0.792–0.825)
Gemma 3 27B 0.734 (0.714–0.753)
Qwen3 32B 0.673 (0.655–0.692)
Llama 4 Maverick 0.584 (0.566–0.601)
Llama 3.3 70B 0.577 (0.565–0.588)

2 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Scores craft, not funniness.

Stand-Up Set: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Stand-Up Set: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.834 (0.819–0.850) — top band
Kimi K3 0.825 (0.812–0.839) — top band
Qwen3.8 Max 0.792 (0.784–0.799)
DeepSeek V4 Pro 0.783 (0.772–0.794)
Gemini 3.6 Flash 0.783 (0.770–0.795)
GPT-5.6 Terra 0.770 (0.760–0.779)
Grok 4.5 0.764 (0.748–0.779)
GPT-5.6 Luna 0.758 (0.746–0.771)
GLM 5 0.743 (0.728–0.758)
Gemma 3 27B 0.707 (0.690–0.723)
Qwen3 32B 0.690 (0.673–0.707)
Llama 4 Maverick 0.555 (0.542–0.567)
Llama 3.3 70B 0.540 (0.525–0.554)

2 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Scores craft, not funniness.

Movie Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Movie Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.861 (0.848–0.873) — top band
Kimi K3 0.858 (0.843–0.873) — top band
Qwen3.8 Max 0.830 (0.816–0.844) — top band
GPT-5.6 Luna 0.824 (0.809–0.839) — top band
GPT-5.6 Terra 0.823 (0.811–0.836) — top band
Grok 4.5 0.813 (0.796–0.829) — top band
Gemini 3.6 Flash 0.812 (0.799–0.825) — top band
DeepSeek V4 Pro 0.790 (0.773–0.805) — top band
GLM 5 0.765 (0.752–0.778) — top band
Qwen3 32B 0.718 (0.700–0.736)
Gemma 3 27B 0.714 (0.694–0.733)
Llama 4 Maverick 0.546 (0.520–0.571)
Llama 3.3 70B 0.532 (0.512–0.554)

9 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Song Lyrics: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Song Lyrics: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.852 (0.832–0.872) — top band
GPT-5.6 Terra 0.778 (0.760–0.797)
GPT-5.6 Luna 0.752 (0.734–0.771)
Qwen3.8 Max 0.744 (0.726–0.761)
Gemini 3.6 Flash 0.735 (0.719–0.751)
DeepSeek V4 Pro 0.734 (0.717–0.750)
Grok 4.5 0.722 (0.704–0.738)
GLM 5 0.716 (0.697–0.737)
Qwen3 32B 0.687 (0.667–0.707)
Gemma 3 27B 0.651 (0.632–0.672)
Llama 4 Maverick 0.517 (0.506–0.529)
Llama 3.3 70B 0.509 (0.499–0.519)
Kimi K3 withheld — only 33% counted

One model holds the top band here — the only domain where the evidence separates a single winner. 1 row is withheld, shown as a verdict with its reason rather than dropped.

Meme Caption: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Meme Caption: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.820 (0.793–0.847) — top band
Qwen3.8 Max 0.815 (0.792–0.837) — top band
GPT-5.6 Terra 0.807 (0.782–0.832) — top band
GLM 5 0.804 (0.779–0.829) — top band
Gemini 3.6 Flash 0.800 (0.773–0.824) — top band
Kimi K3 0.800 (0.777–0.822) — top band
DeepSeek V4 Pro 0.797 (0.762–0.830) — top band
Grok 4.5 0.792 (0.766–0.817) — top band
GPT-5.6 Luna 0.784 (0.756–0.814) — top band
Qwen3 32B 0.773 (0.743–0.804) — top band
Gemma 3 27B 0.750 (0.721–0.777) — top band
Llama 3.3 70B 0.739 (0.706–0.770) — top band
Llama 4 Maverick 0.674 (0.640–0.709) — top band

13 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Scores craft, not funniness.

Book Opening: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Book Opening: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.887 (0.869–0.903) — top band
Kimi K3 0.865 (0.847–0.882) — top band
GPT-5.6 Terra 0.845 (0.826–0.863) — top band
GPT-5.6 Luna 0.843 (0.827–0.858) — top band
Gemini 3.6 Flash 0.842 (0.820–0.864) — top band
Qwen3.8 Max 0.838 (0.820–0.853) — top band
Grok 4.5 0.809 (0.795–0.823) — top band
DeepSeek V4 Pro 0.809 (0.792–0.826) — top band
GLM 5 0.787 (0.763–0.809) — top band
Qwen3 32B 0.779 (0.753–0.802) — top band
Gemma 3 27B 0.756 (0.735–0.775) — top band
Llama 4 Maverick 0.603 (0.583–0.624)
Llama 3.3 70B 0.583 (0.560–0.604)

11 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Per-domain — no overall score exists. Each tab is one domain measured on its own rubric. There is no view that combines them: craft in a sitcom scene and craft in a book opening are not the same quantity, and averaging them would produce a number that means nothing while looking authoritative.


Cost against craft

Sitcom Scene score against list price. Vertical bars are 95% intervals. Dots are coloured by provider family; a ring marks the top rank band.
Sitcom Scene score against list price. Vertical bars are 95% intervals. Dots are coloured by provider family; a ring marks the top rank band. — data
Model Score (interval) at price
Claude Opus 5 0.875 (0.865–0.886) at $25
GPT-5.6 Terra 0.833 (0.820–0.844) at $6
Grok 4.5 0.809 (0.792–0.825) at $6
Gemini 3.6 Flash 0.831 (0.821–0.840) at $7.5
GPT-5.6 Luna 0.830 (0.815–0.842) at $0.6
Kimi K3 0.864 (0.855–0.873) at $15
Qwen3.8 Max 0.831 (0.815–0.843) at $6
DeepSeek V4 Pro 0.818 (0.803–0.830) at $0.87
GLM 5 0.820 (0.808–0.830) at $2.55
Llama 4 Maverick 0.584 (0.566–0.601) at $0.8
Llama 3.3 70B 0.577 (0.565–0.588) at $0.32
Qwen3 32B 0.673 (0.655–0.692) at $0.28
Gemma 3 27B 0.734 (0.714–0.753) at $0.45
Stand-Up Set score against list price. Vertical bars are 95% intervals. Dots are coloured by provider family; a ring marks the top rank band.
Stand-Up Set score against list price. Vertical bars are 95% intervals. Dots are coloured by provider family; a ring marks the top rank band. — data
Model Score (interval) at price
Claude Opus 5 0.834 (0.819–0.850) at $25
GPT-5.6 Terra 0.770 (0.760–0.779) at $6
Grok 4.5 0.764 (0.748–0.779) at $6
Gemini 3.6 Flash 0.783 (0.770–0.795) at $7.5
GPT-5.6 Luna 0.758 (0.746–0.771) at $0.6
Kimi K3 0.825 (0.812–0.839) at $15
Qwen3.8 Max 0.792 (0.784–0.799) at $6
DeepSeek V4 Pro 0.783 (0.772–0.794) at $0.87
GLM 5 0.743 (0.728–0.758) at $2.55
Llama 4 Maverick 0.555 (0.542–0.567) at $0.8
Llama 3.3 70B 0.540 (0.525–0.554) at $0.32
Qwen3 32B 0.690 (0.673–0.707) at $0.28
Gemma 3 27B 0.707 (0.690–0.723) at $0.45
Movie Scene score against list price. Vertical bars are 95% intervals. Dots are coloured by provider family; a ring marks the top rank band.
Movie Scene score against list price. Vertical bars are 95% intervals. Dots are coloured by provider family; a ring marks the top rank band. — data
Model Score (interval) at price
Claude Opus 5 0.861 (0.848–0.873) at $25
GPT-5.6 Terra 0.823 (0.811–0.836) at $6
Grok 4.5 0.813 (0.796–0.829) at $6
Gemini 3.6 Flash 0.812 (0.799–0.825) at $7.5
GPT-5.6 Luna 0.824 (0.809–0.839) at $0.6
Kimi K3 0.858 (0.843–0.873) at $15
Qwen3.8 Max 0.830 (0.816–0.844) at $6
DeepSeek V4 Pro 0.790 (0.773–0.805) at $0.87
GLM 5 0.765 (0.752–0.778) at $2.55
Llama 4 Maverick 0.546 (0.520–0.571) at $0.8
Llama 3.3 70B 0.532 (0.512–0.554) at $0.32
Qwen3 32B 0.718 (0.700–0.736) at $0.28
Gemma 3 27B 0.714 (0.694–0.733) at $0.45
Song Lyrics score against list price. Vertical bars are 95% intervals. Dots are coloured by provider family; a ring marks the top rank band.
Song Lyrics score against list price. Vertical bars are 95% intervals. Dots are coloured by provider family; a ring marks the top rank band. — data
Model Score (interval) at price
Claude Opus 5 0.852 (0.832–0.872) at $25
GPT-5.6 Terra 0.778 (0.760–0.797) at $6
Grok 4.5 0.722 (0.704–0.738) at $6
Gemini 3.6 Flash 0.735 (0.719–0.751) at $7.5
GPT-5.6 Luna 0.752 (0.734–0.771) at $0.6
Qwen3.8 Max 0.744 (0.726–0.761) at $6
DeepSeek V4 Pro 0.734 (0.717–0.750) at $0.87
GLM 5 0.716 (0.697–0.737) at $2.55
Llama 4 Maverick 0.517 (0.506–0.529) at $0.8
Llama 3.3 70B 0.509 (0.499–0.519) at $0.32
Qwen3 32B 0.687 (0.667–0.707) at $0.28
Gemma 3 27B 0.651 (0.632–0.672) at $0.45
Meme Caption score against list price. Vertical bars are 95% intervals. Dots are coloured by provider family; a ring marks the top rank band.
Meme Caption score against list price. Vertical bars are 95% intervals. Dots are coloured by provider family; a ring marks the top rank band. — data
Model Score (interval) at price
Claude Opus 5 0.820 (0.793–0.847) at $25
GPT-5.6 Terra 0.807 (0.782–0.832) at $6
Grok 4.5 0.792 (0.766–0.817) at $6
Gemini 3.6 Flash 0.800 (0.773–0.824) at $7.5
GPT-5.6 Luna 0.784 (0.756–0.814) at $0.6
Kimi K3 0.800 (0.777–0.822) at $15
Qwen3.8 Max 0.815 (0.792–0.837) at $6
DeepSeek V4 Pro 0.797 (0.762–0.830) at $0.87
GLM 5 0.804 (0.779–0.829) at $2.55
Llama 4 Maverick 0.674 (0.640–0.709) at $0.8
Llama 3.3 70B 0.739 (0.706–0.770) at $0.32
Qwen3 32B 0.773 (0.743–0.804) at $0.28
Gemma 3 27B 0.750 (0.721–0.777) at $0.45
Book Opening score against list price. Vertical bars are 95% intervals. Dots are coloured by provider family; a ring marks the top rank band.
Book Opening score against list price. Vertical bars are 95% intervals. Dots are coloured by provider family; a ring marks the top rank band. — data
Model Score (interval) at price
Claude Opus 5 0.887 (0.869–0.903) at $25
GPT-5.6 Terra 0.845 (0.826–0.863) at $6
Grok 4.5 0.809 (0.795–0.823) at $6
Gemini 3.6 Flash 0.842 (0.820–0.864) at $7.5
GPT-5.6 Luna 0.843 (0.827–0.858) at $0.6
Kimi K3 0.865 (0.847–0.882) at $15
Qwen3.8 Max 0.838 (0.820–0.853) at $6
DeepSeek V4 Pro 0.809 (0.792–0.826) at $0.87
GLM 5 0.787 (0.763–0.809) at $2.55
Llama 4 Maverick 0.603 (0.583–0.624) at $0.8
Llama 3.3 70B 0.583 (0.560–0.604) at $0.32
Qwen3 32B 0.779 (0.753–0.802) at $0.28
Gemma 3 27B 0.756 (0.735–0.775) at $0.45

Per-domain — no overall score exists. Each tab plots one domain's score. There is no view of this chart that combines them, because craft in a sitcom scene and craft in a book opening are not the same quantity and averaging them would produce a number that means nothing while looking authoritative.

The cost axis is list price per million tokens, blended across input and output at the pinned model versions. It is not what this sweep actually spent: real spend depends on the input/output ratio of each task and on provider routing, and it is recorded per call in the run journal. Read the horizontal axis as an order-of-magnitude comparison, not as a quote.