Canonic · a verification company for unverifiable domains
RL environments for creative work, with a grader you can audit.
For post-training teams training agents on writing, comedy and other work with no
unit test at the end of it. RL improves an agent only as far as its environment can
measure. Coding and maths got graders; creative work got preference scores. We build
the grader, ship it as a pinned environment, and publish where it breaks.
Two models answering the same task, from the Stage 1 sweep. Not a curated
pair from different prompts — that would confound how hard the task was with how good the
model is.
Task The opening of a comic novel about an institution.
Llama 4 Maverick0.50
The Institute for Unlikely Sciences stood like a sentinel on the outskirts of town, its imposing façade a testament to the enduring power of Victorian optimism. As the morning sun struggled to penetrate the grimy windows, a faint hum of activity began to emanate from within. In the foyer, the Director, Dr. Helena Anders, stood poised behind a polished wooden counter, her eyes fixed on the sweeping curves of the staircase as she awaited the arrival of her senior staff.
glm-5.2: 0.30 · mistral-large-2512: 0.70
Claude Opus 50.95
The Institute had been founded in 1874 by a deed of gift whose seventh clause required that the Governing Body consist of nine members, and whose eleventh clause, added by the founder's widow in a spirit of correction, required that no more than eight of them be alive at any one time. Lawyers had been consulted. The lawyers had confirmed that the clauses were both valid, and that they were consistent, and that the consequence of their consistency was that the Institute could meet but could not resolve, since a resolution required the assent of a majority of the whole number and the whole number included Chair Nine, who was, by the terms of the gift, deceased.
glm-5.2: 0.95 · mistral-large-2512: 0.95
Both judges are named above each excerpt because the claim is that they ordered
these independently, and an averaged score cannot show that. Across the
sweep, both judges ordered
180 of
190 eligible tasks without
contradicting each other.
02
Proof
6 domains. 6 separate scales.
One board per domain, each on its own rubric, and no combined view anywhere. The scales do not convert: craft in a sitcom scene and craft in a book opening are not
the same quantity.
On 5 of 6 boards the top band is a tie — the intervals overlap, so
the board declines to order them. That is the finding, not a gap in it.
Results by domain
Sitcom Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Sitcom Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model
Score (interval)
Claude Opus 5
0.875 (0.865–0.886) — top band
Kimi K3
0.864 (0.855–0.873) — top band
GPT-5.6 Terra
0.833 (0.820–0.844)
Gemini 3.6 Flash
0.831 (0.821–0.840)
Qwen3.8 Max
0.831 (0.815–0.843)
GPT-5.6 Luna
0.830 (0.815–0.842)
GLM 5
0.820 (0.808–0.830)
DeepSeek V4 Pro
0.818 (0.803–0.830)
Grok 4.5
0.809 (0.792–0.825)
Gemma 3 27B
0.734 (0.714–0.753)
Qwen3 32B
0.673 (0.655–0.692)
Llama 4 Maverick
0.584 (0.566–0.601)
Llama 3.3 70B
0.577 (0.565–0.588)
2 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.
Stand-Up Set: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Stand-Up Set: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model
Score (interval)
Claude Opus 5
0.834 (0.819–0.850) — top band
Kimi K3
0.825 (0.812–0.839) — top band
Qwen3.8 Max
0.792 (0.784–0.799)
DeepSeek V4 Pro
0.783 (0.772–0.794)
Gemini 3.6 Flash
0.783 (0.770–0.795)
GPT-5.6 Terra
0.770 (0.760–0.779)
Grok 4.5
0.764 (0.748–0.779)
GPT-5.6 Luna
0.758 (0.746–0.771)
GLM 5
0.743 (0.728–0.758)
Gemma 3 27B
0.707 (0.690–0.723)
Qwen3 32B
0.690 (0.673–0.707)
Llama 4 Maverick
0.555 (0.542–0.567)
Llama 3.3 70B
0.540 (0.525–0.554)
2 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.
Movie Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Movie Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model
Score (interval)
Claude Opus 5
0.861 (0.848–0.873) — top band
Kimi K3
0.858 (0.843–0.873) — top band
Qwen3.8 Max
0.830 (0.816–0.844) — top band
GPT-5.6 Luna
0.824 (0.809–0.839) — top band
GPT-5.6 Terra
0.823 (0.811–0.836) — top band
Grok 4.5
0.813 (0.796–0.829) — top band
Gemini 3.6 Flash
0.812 (0.799–0.825) — top band
DeepSeek V4 Pro
0.790 (0.773–0.805) — top band
GLM 5
0.765 (0.752–0.778) — top band
Qwen3 32B
0.718 (0.700–0.736)
Gemma 3 27B
0.714 (0.694–0.733)
Llama 4 Maverick
0.546 (0.520–0.571)
Llama 3.3 70B
0.532 (0.512–0.554)
9 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.
Song Lyrics: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Song Lyrics: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model
Score (interval)
Claude Opus 5
0.852 (0.832–0.872) — top band
GPT-5.6 Terra
0.778 (0.760–0.797)
GPT-5.6 Luna
0.752 (0.734–0.771)
Qwen3.8 Max
0.744 (0.726–0.761)
Gemini 3.6 Flash
0.735 (0.719–0.751)
DeepSeek V4 Pro
0.734 (0.717–0.750)
Grok 4.5
0.722 (0.704–0.738)
GLM 5
0.716 (0.697–0.737)
Qwen3 32B
0.687 (0.667–0.707)
Gemma 3 27B
0.651 (0.632–0.672)
Llama 4 Maverick
0.517 (0.506–0.529)
Llama 3.3 70B
0.509 (0.499–0.519)
Kimi K3
withheld — only 33% counted
One model holds the top band here — the only domain where the evidence separates a single winner. 1 row is withheld, shown as a verdict with its reason rather than dropped.
Meme Caption: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Meme Caption: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model
Score (interval)
Claude Opus 5
0.820 (0.793–0.847) — top band
Qwen3.8 Max
0.815 (0.792–0.837) — top band
GPT-5.6 Terra
0.807 (0.782–0.832) — top band
GLM 5
0.804 (0.779–0.829) — top band
Gemini 3.6 Flash
0.800 (0.773–0.824) — top band
Kimi K3
0.800 (0.777–0.822) — top band
DeepSeek V4 Pro
0.797 (0.762–0.830) — top band
Grok 4.5
0.792 (0.766–0.817) — top band
GPT-5.6 Luna
0.784 (0.756–0.814) — top band
Qwen3 32B
0.773 (0.743–0.804) — top band
Gemma 3 27B
0.750 (0.721–0.777) — top band
Llama 3.3 70B
0.739 (0.706–0.770) — top band
Llama 4 Maverick
0.674 (0.640–0.709) — top band
13 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.
Book Opening: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Book Opening: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model
Score (interval)
Claude Opus 5
0.887 (0.869–0.903) — top band
Kimi K3
0.865 (0.847–0.882) — top band
GPT-5.6 Terra
0.845 (0.826–0.863) — top band
GPT-5.6 Luna
0.843 (0.827–0.858) — top band
Gemini 3.6 Flash
0.842 (0.820–0.864) — top band
Qwen3.8 Max
0.838 (0.820–0.853) — top band
Grok 4.5
0.809 (0.795–0.823) — top band
DeepSeek V4 Pro
0.809 (0.792–0.826) — top band
GLM 5
0.787 (0.763–0.809) — top band
Qwen3 32B
0.779 (0.753–0.802) — top band
Gemma 3 27B
0.756 (0.735–0.775) — top band
Llama 4 Maverick
0.603 (0.583–0.624)
Llama 3.3 70B
0.583 (0.560–0.604)
11 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.
Per-domain — no overall score exists. Each tab is one domain measured on its own rubric.
There is no view that combines them: craft in a sitcom scene and craft in a book opening are
not the same quantity, and averaging them would produce a number that means nothing while
looking authoritative.
Every evaluation company claims rigour. The cheap way to prove it is to publish where our
own instrument failed, on the same page as the results.
2
of 8 reward-hacking probes scored above our threshold — they beat our own
rubric. We published them unaltered rather than retuning until they failed. The rubric
caught the other 6
1
row is withheld rather than estimated
— Kimi K3, shown as a verdict with its reason, never as a low score
0
human-validation studies run. We do not publish an agreement statistic because we have
not earned one. The judges are 2 language models from different families,
and that is the whole of it
04
Next
Next
A seventh domain asking whether a model actually knows a place,
rather than defaulting to the nearest postcard. Models write Mangalore as Kochi's fishing
nets and Kerala's houseboats.
Not built and not measured. Listed as the next benchmark so it is clear what we are
working on, not as a capability we are selling.