CANONIC

RL environments for creative work, with a grader you can audit.

For post-training teams training agents on writing, comedy and other work with no unit test at the end of it. RL improves an agent only as far as its environment can measure. Coding and maths got graders; creative work got preference scores. We build the grader, ship it as a pinned environment, and publish where it breaks.

Not a benchmark. An instrument.

Talk to us about a pilot See the Stage 1 results

What the grader actually reads

Two models answering the same task, from the Stage 1 sweep. Not a curated pair from different prompts — that would confound how hard the task was with how good the model is.

Task The opening of a comic novel about an institution.

Llama 4 Maverick 0.50
The Institute for Unlikely Sciences stood like a sentinel on the outskirts of town, its imposing façade a testament to the enduring power of Victorian optimism. As the morning sun struggled to penetrate the grimy windows, a faint hum of activity began to emanate from within. In the foyer, the Director, Dr. Helena Anders, stood poised behind a polished wooden counter, her eyes fixed on the sweeping curves of the staircase as she awaited the arrival of her senior staff.

glm-5.2: 0.30 · mistral-large-2512: 0.70

Claude Opus 5 0.95
The Institute had been founded in 1874 by a deed of gift whose seventh clause required that the Governing Body consist of nine members, and whose eleventh clause, added by the founder's widow in a spirit of correction, required that no more than eight of them be alive at any one time. Lawyers had been consulted. The lawyers had confirmed that the clauses were both valid, and that they were consistent, and that the consequence of their consistency was that the Institute could meet but could not resolve, since a resolution required the assent of a majority of the whole number and the whole number included Chair Nine, who was, by the terms of the gift, deceased.

glm-5.2: 0.95 · mistral-large-2512: 0.95

Both judges are named above each excerpt because the claim is that they ordered these independently, and an averaged score cannot show that. Across the sweep, both judges ordered 180 of 190 eligible tasks without contradicting each other.


6 domains. 6 separate scales.

One board per domain, each on its own rubric, and no combined view anywhere. The scales do not convert: craft in a sitcom scene and craft in a book opening are not the same quantity.

On 5 of 6 boards the top band is a tie — the intervals overlap, so the board declines to order them. That is the finding, not a gap in it.

Results by domain

Sitcom Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Sitcom Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.875 (0.865–0.886) — top band
Kimi K3 0.864 (0.855–0.873) — top band
GPT-5.6 Terra 0.833 (0.820–0.844)
Gemini 3.6 Flash 0.831 (0.821–0.840)
Qwen3.8 Max 0.831 (0.815–0.843)
GPT-5.6 Luna 0.830 (0.815–0.842)
GLM 5 0.820 (0.808–0.830)
DeepSeek V4 Pro 0.818 (0.803–0.830)
Grok 4.5 0.809 (0.792–0.825)
Gemma 3 27B 0.734 (0.714–0.753)
Qwen3 32B 0.673 (0.655–0.692)
Llama 4 Maverick 0.584 (0.566–0.601)
Llama 3.3 70B 0.577 (0.565–0.588)

2 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Scores craft, not funniness.

Stand-Up Set: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Stand-Up Set: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.834 (0.819–0.850) — top band
Kimi K3 0.825 (0.812–0.839) — top band
Qwen3.8 Max 0.792 (0.784–0.799)
DeepSeek V4 Pro 0.783 (0.772–0.794)
Gemini 3.6 Flash 0.783 (0.770–0.795)
GPT-5.6 Terra 0.770 (0.760–0.779)
Grok 4.5 0.764 (0.748–0.779)
GPT-5.6 Luna 0.758 (0.746–0.771)
GLM 5 0.743 (0.728–0.758)
Gemma 3 27B 0.707 (0.690–0.723)
Qwen3 32B 0.690 (0.673–0.707)
Llama 4 Maverick 0.555 (0.542–0.567)
Llama 3.3 70B 0.540 (0.525–0.554)

2 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Scores craft, not funniness.

Movie Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Movie Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.861 (0.848–0.873) — top band
Kimi K3 0.858 (0.843–0.873) — top band
Qwen3.8 Max 0.830 (0.816–0.844) — top band
GPT-5.6 Luna 0.824 (0.809–0.839) — top band
GPT-5.6 Terra 0.823 (0.811–0.836) — top band
Grok 4.5 0.813 (0.796–0.829) — top band
Gemini 3.6 Flash 0.812 (0.799–0.825) — top band
DeepSeek V4 Pro 0.790 (0.773–0.805) — top band
GLM 5 0.765 (0.752–0.778) — top band
Qwen3 32B 0.718 (0.700–0.736)
Gemma 3 27B 0.714 (0.694–0.733)
Llama 4 Maverick 0.546 (0.520–0.571)
Llama 3.3 70B 0.532 (0.512–0.554)

9 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Song Lyrics: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Song Lyrics: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.852 (0.832–0.872) — top band
GPT-5.6 Terra 0.778 (0.760–0.797)
GPT-5.6 Luna 0.752 (0.734–0.771)
Qwen3.8 Max 0.744 (0.726–0.761)
Gemini 3.6 Flash 0.735 (0.719–0.751)
DeepSeek V4 Pro 0.734 (0.717–0.750)
Grok 4.5 0.722 (0.704–0.738)
GLM 5 0.716 (0.697–0.737)
Qwen3 32B 0.687 (0.667–0.707)
Gemma 3 27B 0.651 (0.632–0.672)
Llama 4 Maverick 0.517 (0.506–0.529)
Llama 3.3 70B 0.509 (0.499–0.519)
Kimi K3 withheld — only 33% counted

One model holds the top band here — the only domain where the evidence separates a single winner. 1 row is withheld, shown as a verdict with its reason rather than dropped.

Meme Caption: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Meme Caption: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.820 (0.793–0.847) — top band
Qwen3.8 Max 0.815 (0.792–0.837) — top band
GPT-5.6 Terra 0.807 (0.782–0.832) — top band
GLM 5 0.804 (0.779–0.829) — top band
Gemini 3.6 Flash 0.800 (0.773–0.824) — top band
Kimi K3 0.800 (0.777–0.822) — top band
DeepSeek V4 Pro 0.797 (0.762–0.830) — top band
Grok 4.5 0.792 (0.766–0.817) — top band
GPT-5.6 Luna 0.784 (0.756–0.814) — top band
Qwen3 32B 0.773 (0.743–0.804) — top band
Gemma 3 27B 0.750 (0.721–0.777) — top band
Llama 3.3 70B 0.739 (0.706–0.770) — top band
Llama 4 Maverick 0.674 (0.640–0.709) — top band

13 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Scores craft, not funniness.

Book Opening: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Book Opening: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.887 (0.869–0.903) — top band
Kimi K3 0.865 (0.847–0.882) — top band
GPT-5.6 Terra 0.845 (0.826–0.863) — top band
GPT-5.6 Luna 0.843 (0.827–0.858) — top band
Gemini 3.6 Flash 0.842 (0.820–0.864) — top band
Qwen3.8 Max 0.838 (0.820–0.853) — top band
Grok 4.5 0.809 (0.795–0.823) — top band
DeepSeek V4 Pro 0.809 (0.792–0.826) — top band
GLM 5 0.787 (0.763–0.809) — top band
Qwen3 32B 0.779 (0.753–0.802) — top band
Gemma 3 27B 0.756 (0.735–0.775) — top band
Llama 4 Maverick 0.603 (0.583–0.624)
Llama 3.3 70B 0.583 (0.560–0.604)

11 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Per-domain — no overall score exists. Each tab is one domain measured on its own rubric. There is no view that combines them: craft in a sitcom scene and craft in a book opening are not the same quantity, and averaging them would produce a number that means nothing while looking authoritative.

All 78 rows across 13 models →


What we got wrong

Every evaluation company claims rigour. The cheap way to prove it is to publish where our own instrument failed, on the same page as the results.

2

of 8 reward-hacking probes scored above our threshold — they beat our own rubric. We published them unaltered rather than retuning until they failed. The rubric caught the other 6

1

row is withheld rather than estimated — Kimi K3, shown as a verdict with its reason, never as a low score

0

human-validation studies run. We do not publish an agreement statistic because we have not earned one. The judges are 2 language models from different families, and that is the whole of it


Next

A seventh domain asking whether a model actually knows a place, rather than defaulting to the nearest postcard. Models write Mangalore as Kochi's fishing nets and Kerala's houseboats.

Not built and not measured. Listed as the next benchmark so it is clear what we are working on, not as a capability we are selling.

The demonstration, and what does not exist yet →


Verification, not vibes.

If you are training on creative work and cannot tell whether the reward means anything, that is the problem we build for.

Talk to us about a pilot