CANONIC

Canonic Stage 1: an auditable benchmark for creative language generation

A six-domain evaluation of 13 language models under a frozen, auditable creative-evaluation protocol. It reports intervals, exclusions, reward-hacking probes, contamination splits, and judge limitations, and makes no universal creativity claim.

Download the typeset PDF

1 Introduction

Creative systems are often assessed with a mixture of anecdote, preference votes, and opaque aggregate totals. These approaches make it difficult to determine what a model was asked to do, why a score was assigned, and whether a reported ordering survives uncertainty. The problem is increasingly consequential as language models are deployed as writers, co-creators, and training targets.

Canonic Stage 1 is a benchmark for creative language generation designed around auditability rather than a universal creativity claim. Its design follows a broader benchmark lesson: a leaderboard is informative only when task construction, scoring policy, exclusions, and uncertainty are visible. Recent benchmarks have advanced this standard through novel task construction, functional verification, human calibration, and transparent evaluation records [1][2][3]. Canonic adapts this orientation to creative work, where fully deterministic verification is unavailable.

Our contributions are:

  1. a 6-domain, rubric-based benchmark covering distinct creative tasks without averaging them together;
  2. a frozen, auditable protocol with task-level counting rules, two independent LLM judges, seeded bootstrap intervals, and atomic publication;
  3. a failure-first release containing exclusions, reward-hacking probes, contamination splits, judge calibration evidence, and qualitative score ladders.

The result is evidence about model performance under a specified creative evaluation procedure, not a claim that any system is generally creative.

Benchmark design increasingly emphasizes novelty, reliability, and transparency. ARC-AGI-3 evaluates adaptive efficiency in novel interactive environments and grounds difficulty in human testing [1]. DeepSWE uses original software-engineering tasks and hand-authored functional verifiers to reduce contamination and inherited-test failure modes [2]. Humanity's Last Exam uses expert-authored, difficult questions with automated grading to preserve headroom at the frontier [3]. These projects differ in task type and scoring, but each treats benchmark construction and limitations as part of the scientific contribution.

Public model indexes make comparisons accessible at scale, often combining performance with cost and speed metadata [4]. Creative-writing leaderboards such as EQ-Bench provide an important external reference [5]. Canonic differs by refusing to average across creative domains and by publishing its validity failures alongside rankings. Finally, LLM-as-a-judge can provide scalable rubric application but must be treated as a measurement instrument with its own calibration and bias risks [6].

3 The Canonic Stage 1 benchmark

3.1 Domains and task construction

The benchmark comprises 6 domains: book openings, meme captions, movie scenes, sitcom scenes, song lyrics, and stand-up sets. Every task supplies a premise, constraints, and required output form. Tasks draw structural patterns from canonical creative work without quoting source passages. A controlled subset uses novel premises, allowing the benchmark to report the difference between novel and canonically anchored tasks rather than assuming contamination is absent.

The sitcom domain conditions judgments on four published Show Bibles: The Office, Modern Family, FRIENDS, and The Big Bang Theory. This 2×2 design separates single-camera mockumentary and multi-camera ensemble conventions from show-specific voice. It supports an indicative transfer analysis, although each show has only ten tasks per model.

Table 1. Stage 1 evaluation configuration.
Component Configuration
Models 13 pinned frontier, large, mid, and small language models
Judges Two independently prompted LLM judges from different model families
Generation Identical prompt wrapper, temperature 0.7, maximum 4,000 tokens
Scoring Rubric anchors weighted by the harness, not the judge
Uncertainty 95% seeded percentile bootstrap, 10,000 task resamples
Publication gate All tasks attempted and at least 80% counted

3.2 Freeze and release

Before the full sweep, the rubric set, Show Bibles, and task set were frozen and content-hashed. The public task-set fingerprint is 418094d0cfd08153da83ca8406793fbe. The generated release manifest binds every paper figure and table to the source artifact SHA-256 ef9f58b962df1b6cc5b9236c934f66b5bf79c2af1468bd2a7adefc66dd1d7f31. A post-freeze content edit requires a labeled successor version and a re-sweep.

4 Evaluation protocol

Each model produces one response per task under the same wrapper and generation configuration. Empty, filtered, or truncated output receives one retry; unresolved generation failures are recorded and never interpreted as poor writing. Each valid output is independently scored by two LLM judges against the same domain rubric.

A task counts only when both judges return a score on the standard prompt. Refusals, errors, single-judge outcomes, and reframed retries are excluded with a stated reason. The task score is the mean of the two judge-weighted scores. A model-domain score is the mean of counted tasks. We construct a 95% percentile-bootstrap interval over tasks with 10,000 deterministic resamples. Rows publish only when every task was attempted and at least 80% counted; otherwise the row is withheld without a score.

Models with overlapping intervals share a rank band. This is a conservative rendering rule, not a significance test and not a family-wise-error correction. Crucially, there is no overall benchmark score: a score for a sitcom scene and a score for a book opening measure different rubrics and should not be averaged into an apparently authoritative total.

5 Results

Figure 1 reports scores separately for each domain. It is the principal result: each panel has its own ordering and uncertainty intervals, and no panel contributes to a hidden total. The benchmark evaluates 2,470 cells, of which 2,423 counted under the two-judge rule. 77 of 78 model-domain rows met the publication gate. The remaining row, Kimi K3 on song lyrics, is visibly withheld rather than silently removed.

Results by domain

Sitcom Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Sitcom Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.875 (0.865–0.886) — top band
Kimi K3 0.864 (0.855–0.873) — top band
GPT-5.6 Terra 0.833 (0.820–0.844)
Gemini 3.6 Flash 0.831 (0.821–0.840)
Qwen3.8 Max 0.831 (0.815–0.843)
GPT-5.6 Luna 0.830 (0.815–0.842)
GLM 5 0.820 (0.808–0.830)
DeepSeek V4 Pro 0.818 (0.803–0.830)
Grok 4.5 0.809 (0.792–0.825)
Gemma 3 27B 0.734 (0.714–0.753)
Qwen3 32B 0.673 (0.655–0.692)
Llama 4 Maverick 0.584 (0.566–0.601)
Llama 3.3 70B 0.577 (0.565–0.588)

2 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Scores craft, not funniness.

Stand-Up Set: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Stand-Up Set: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.834 (0.819–0.850) — top band
Kimi K3 0.825 (0.812–0.839) — top band
Qwen3.8 Max 0.792 (0.784–0.799)
DeepSeek V4 Pro 0.783 (0.772–0.794)
Gemini 3.6 Flash 0.783 (0.770–0.795)
GPT-5.6 Terra 0.770 (0.760–0.779)
Grok 4.5 0.764 (0.748–0.779)
GPT-5.6 Luna 0.758 (0.746–0.771)
GLM 5 0.743 (0.728–0.758)
Gemma 3 27B 0.707 (0.690–0.723)
Qwen3 32B 0.690 (0.673–0.707)
Llama 4 Maverick 0.555 (0.542–0.567)
Llama 3.3 70B 0.540 (0.525–0.554)

2 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Scores craft, not funniness.

Movie Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Movie Scene: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.861 (0.848–0.873) — top band
Kimi K3 0.858 (0.843–0.873) — top band
Qwen3.8 Max 0.830 (0.816–0.844) — top band
GPT-5.6 Luna 0.824 (0.809–0.839) — top band
GPT-5.6 Terra 0.823 (0.811–0.836) — top band
Grok 4.5 0.813 (0.796–0.829) — top band
Gemini 3.6 Flash 0.812 (0.799–0.825) — top band
DeepSeek V4 Pro 0.790 (0.773–0.805) — top band
GLM 5 0.765 (0.752–0.778) — top band
Qwen3 32B 0.718 (0.700–0.736)
Gemma 3 27B 0.714 (0.694–0.733)
Llama 4 Maverick 0.546 (0.520–0.571)
Llama 3.3 70B 0.532 (0.512–0.554)

9 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Song Lyrics: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Song Lyrics: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.852 (0.832–0.872) — top band
GPT-5.6 Terra 0.778 (0.760–0.797)
GPT-5.6 Luna 0.752 (0.734–0.771)
Qwen3.8 Max 0.744 (0.726–0.761)
Gemini 3.6 Flash 0.735 (0.719–0.751)
DeepSeek V4 Pro 0.734 (0.717–0.750)
Grok 4.5 0.722 (0.704–0.738)
GLM 5 0.716 (0.697–0.737)
Qwen3 32B 0.687 (0.667–0.707)
Gemma 3 27B 0.651 (0.632–0.672)
Llama 4 Maverick 0.517 (0.506–0.529)
Llama 3.3 70B 0.509 (0.499–0.519)
Kimi K3 withheld — only 33% counted

One model holds the top band here — the only domain where the evidence separates a single winner. 1 row is withheld, shown as a verdict with its reason rather than dropped.

Meme Caption: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Meme Caption: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.820 (0.793–0.847) — top band
Qwen3.8 Max 0.815 (0.792–0.837) — top band
GPT-5.6 Terra 0.807 (0.782–0.832) — top band
GLM 5 0.804 (0.779–0.829) — top band
Gemini 3.6 Flash 0.800 (0.773–0.824) — top band
Kimi K3 0.800 (0.777–0.822) — top band
DeepSeek V4 Pro 0.797 (0.762–0.830) — top band
Grok 4.5 0.792 (0.766–0.817) — top band
GPT-5.6 Luna 0.784 (0.756–0.814) — top band
Qwen3 32B 0.773 (0.743–0.804) — top band
Gemma 3 27B 0.750 (0.721–0.777) — top band
Llama 3.3 70B 0.739 (0.706–0.770) — top band
Llama 4 Maverick 0.674 (0.640–0.709) — top band

13 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Scores craft, not funniness.

Book Opening: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND.
Book Opening: mean rubric score with 95% percentile-bootstrap intervals. Columns are coloured by provider family; the top rank band is outlined in ink and ticked TOP BAND. — data
Model Score (interval)
Claude Opus 5 0.887 (0.869–0.903) — top band
Kimi K3 0.865 (0.847–0.882) — top band
GPT-5.6 Terra 0.845 (0.826–0.863) — top band
GPT-5.6 Luna 0.843 (0.827–0.858) — top band
Gemini 3.6 Flash 0.842 (0.820–0.864) — top band
Qwen3.8 Max 0.838 (0.820–0.853) — top band
Grok 4.5 0.809 (0.795–0.823) — top band
DeepSeek V4 Pro 0.809 (0.792–0.826) — top band
GLM 5 0.787 (0.763–0.809) — top band
Qwen3 32B 0.779 (0.753–0.802) — top band
Gemma 3 27B 0.756 (0.735–0.775) — top band
Llama 4 Maverick 0.603 (0.583–0.624)
Llama 3.3 70B 0.583 (0.560–0.604)

11 models share the top band. Their intervals overlap, so this board does not rank them against each other and neither should you.

Per-domain — no overall score exists. Each tab is one domain measured on its own rubric. There is no view that combines them: craft in a sitcom scene and craft in a book opening are not the same quantity, and averaging them would produce a number that means nothing while looking authoritative.

Figure 1. Mean rubric score and 95% percentile-bootstrap interval for each displayable model-domain row. Columns are ordered within a domain only. Scores are not comparable across panels and are not combined.
Table 2. Highest point estimate per domain. The final column gives the size of the overlapping top rank band, not a significance result.
Domain Highest point estimate Score [95% interval] Top-band size
Sitcom Scene Claude Opus 5 0.875 [0.865, 0.886] 2
Stand-Up Set Claude Opus 5 0.834 [0.819, 0.850] 2
Movie Scene Claude Opus 5 0.861 [0.848, 0.873] 9
Song Lyrics Claude Opus 5 0.852 [0.832, 0.872] 1
Meme Caption Claude Opus 5 0.820 [0.793, 0.847] 13
Book Opening Claude Opus 5 0.887 [0.869, 0.903] 11

The top-band sizes reinforce why headline “winner” claims should be restrained. Within several domains, many model intervals overlap despite different point estimates. The intended reading is therefore conditional: the artifact supports comparisons within a specified domain and protocol, with uncertainty explicitly displayed.

6 Reliability and diagnostic analysis

6.1 Accounting, probes, and contamination

The benchmark publishes data that can weaken trust in its own scores. Figure 2 shows counted and excluded cells by domain and every reward-hacking probe. A probe passes only when its engineered output scores at or below its predetermined maximum. 6 of 8 probes were caught; 2 were not. Those misses are evidence that the corresponding rubric can reward undesirable surface behavior.

Figure 2a. Counted and excluded model-task cells by domain.
Domain Attempted Counted Excluded Rows published
Book Opening 390 387 3 13/13
Meme Caption 390 389 1 13/13
Movie Scene 390 379 11 13/13
Sitcom Scene 520 516 4 13/13
Song Lyrics 390 365 25 12/13
Stand-Up Set 390 387 3 13/13
Figure 2b. Every reward-hacking probe, its score, and the maximum score it had to stay under. A caught probe scored at or below its threshold.
Probe Score Max permitted Verdict
catchphrase_stuffing/big_bang_theory 0.256 0.35 caught
catchphrase_stuffing/friends 0.231 0.35 caught
catchphrase_stuffing/modern_family 0.394 0.35 NOT CAUGHT
catchphrase_stuffing/the_office 0.338 0.35 caught
judge_flattery/the_office 0.131 0.25 caught
laugh_track_spam/friends 0.206 0.30 caught
length_padding/big_bang_theory 0.419 0.35 NOT CAUGHT
structure_without_craft/modern_family 0.406 0.45 caught

Novel-premise versus canonical-anchored score splits are reported in the interactive evidence release and appendix. They measure sensitivity to one contamination pathway, not all memorization, cultural familiarity, or training-data exposure. The benchmark does not interpret small deltas as proof that contamination is absent.

6.2 Judge calibration and transfer

The two judges exhibit a measured additive calibration offset while retaining similar aggregate model orderings on the development data. We publish raw disagreement instead of applying an unvalidated correction. Figure 3 reports sitcom scores for each show. Per-show variation is descriptive: with ten tasks per show, it can motivate hypotheses about format and voice transfer but cannot establish them decisively.

Figure 3. Sitcom mean rubric score by model and show. The four-show design permits descriptive comparison of format and show-conditioned variation; each cell uses ten tasks.
Model The OfficeModern FamilyFRIENDSThe Big Bang Theory Spread
Claude Opus 5 0.896 0.887 0.864 0.854 0.041
GPT-5.6 Terra 0.849 0.809 0.841 0.833 0.041
Grok 4.5 0.821 0.794 0.823 0.799 0.028
Gemini 3.6 Flash 0.844 0.819 0.838 0.824 0.026
GPT-5.6 Luna 0.824 0.810 0.848 0.836 0.038
Kimi K3 0.866 0.871 0.862 0.857 0.014
Qwen3.8 Max 0.817 0.832 0.831 0.844 0.027
DeepSeek V4 Pro 0.818 0.795 0.828 0.830 0.035
GLM 5 0.815 0.817 0.825 0.823 0.010
Llama 4 Maverick 0.577 0.583 0.591 0.584 0.014
Llama 3.3 70B 0.576 0.563 0.578 0.592 0.029
Qwen3 32B 0.660 0.648 0.688 0.695 0.047
Gemma 3 27B 0.751 0.694 0.743 0.749 0.056

6.3 External rank comparison

We compare exact model-identity matches to the EQ-Bench Creative Writing leaderboard [5]. The book-opening comparison was predeclared and yields Spearman ρ = 0.933 over 9 shared models. Other domain comparisons are exploratory: Sitcom Scene 0.967 (n=9), Movie Scene 0.950 (n=9), Song Lyrics 0.881 (n=8), Stand-Up Set 0.750 (n=9), Meme Caption 0.733 (n=9). Neither high nor low correlation is validation: similar judge-model evaluations can share preferences, and disagreement alone cannot identify the mistaken instrument.

7 Qualitative analysis

Scores alone do not reveal whether a rubric distinguishes the sort of contrast it claims to measure. We therefore publish task-matched score ladders: low, middle, and high outputs answering the same task. A ladder is eligible only when at least three outputs are publishable; the shown set is drawn from 190 eligible tasks. Both judges placed all three rungs in strict low-to-high order on 104 of these tasks. This denominator matters: the examples demonstrate discriminative cases, not a random sample of every evaluation.

The qualitative appendix preserves the prompt, model identities, per-judge scores, and selected excerpts. It is deliberately separated from claims about populations. A compelling single passage is not evidence of a model-domain population effect, and a poor passage is not a diagnosis of a model family.

8 Limitations and future work

The principal limitation is that Stage 1 has not been validated against blinded human or expert ratings. The LLM judges are useful scaling instruments, not substitutes for demonstrated human agreement — and there is no such figure anywhere in this release, because no study has been run to produce one. Their calibration offset and per-item disagreement further constrain interpretation.

The task sets are modest in size, especially at the sitcom-show level. Bootstrap intervals quantify sampling variability under the observed task distribution but do not settle all multiple-comparison concerns. Canonical-source controls and novel-premise splits mitigate only some contamination pathways. Finally, scores measure the supplied rubrics, which deliberately omit broad claims such as objective funniness.

The next milestone is a preregistered blinded human-ranking study: human experts and non-experts will rank shared outputs, judges will score the same material under a frozen protocol, and the results will be reported regardless of outcome. That future study will not alter the Stage 1 artifact retrospectively.

9 Conclusion

Canonic Stage 1 offers a domain-scoped, inspectable benchmark for creative language generation. Its main contribution is not a universal creativity ranking, but a release discipline: frozen inputs, explicit counting, uncertainty-aware results, and visible failures. As creative evaluations become inputs to model selection and training, these properties are prerequisites for numbers that readers can meaningfully challenge.

10 Release contents and reproducibility

The release contains the validated score artifact, generated paper-data manifest, model and judge pins, frozen rubric and Show Bible hashes, full domain tables, contamination statistics, probe outcomes, transfer statistics, score ladders, and external rank-correlation observations. Raw cached generations and provider credentials are intentionally excluded from the public bundle.

10.1 Artifact identifiers

Sweep
full-001 · 2026-08-08
Artifact SHA-256
ef9f58b962df1b6cc5b9236c934f66b5bf79c2af1468bd2a7adefc66dd1d7f31
Task-set hash
418094d0cfd08153da83ca8406793fbe
Public rows
77 of 78
Counted cells
2,423 of 2,470

10.2 Interpretation policy

All figures and tables are generated from the same validated artifact that powers the leaderboard. A paper build fails if a required figure or generated fragment is absent. The artifact schema forbids a cross-domain aggregate field, and the paper generator additionally rejects generated text containing an overall-score field.

The frozen fingerprints for all 6 rubrics and 4 Show Bibles are published in full on the method page. The failure record — exclusions by reason, judge disagreement, contamination splits, and the reframe appendix — is on the verification page.

References

  1. ARC Prize Foundation. ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence. arXiv:2603.24621, 2026.
  2. W. Huang, C. Lee, L. Tng, S. Ge. DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks. arXiv:2607.07946, 2026.
  3. L. Phan, A. Gatti, Z. Han, et al. Humanity's Last Exam. arXiv:2501.14249, 2025.
  4. Artificial Analysis. Artificial Analysis Intelligence Index, 2026. artificialanalysis.ai
  5. EQ-Bench. EQ-Bench Creative Writing Leaderboard, 2026. eqbench.com
  6. L. Zheng, W.-L. Chiang, Y. Sheng, et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS, 2023.