Research paper · version 1.0 · 2026-08-08
Canonic Stage 1: an auditable benchmark for creative language generation
A six-domain evaluation of 13 language models under a frozen, auditable creative-evaluation protocol. It reports intervals, exclusions, reward-hacking probes, contamination splits, and judge limitations, and makes no universal creativity claim.
1 Introduction
Creative systems are often assessed with a mixture of anecdote, preference votes, and opaque aggregate totals. These approaches make it difficult to determine what a model was asked to do, why a score was assigned, and whether a reported ordering survives uncertainty. The problem is increasingly consequential as language models are deployed as writers, co-creators, and training targets.
Canonic Stage 1 is a benchmark for creative language generation designed around auditability rather than a universal creativity claim. Its design follows a broader benchmark lesson: a leaderboard is informative only when task construction, scoring policy, exclusions, and uncertainty are visible. Recent benchmarks have advanced this standard through novel task construction, functional verification, human calibration, and transparent evaluation records [1][2][3]. Canonic adapts this orientation to creative work, where fully deterministic verification is unavailable.
Our contributions are:
- a 6-domain, rubric-based benchmark covering distinct creative tasks without averaging them together;
- a frozen, auditable protocol with task-level counting rules, two independent LLM judges, seeded bootstrap intervals, and atomic publication;
- a failure-first release containing exclusions, reward-hacking probes, contamination splits, judge calibration evidence, and qualitative score ladders.
The result is evidence about model performance under a specified creative evaluation procedure, not a claim that any system is generally creative.
2 Related work
Benchmark design increasingly emphasizes novelty, reliability, and transparency. ARC-AGI-3 evaluates adaptive efficiency in novel interactive environments and grounds difficulty in human testing [1]. DeepSWE uses original software-engineering tasks and hand-authored functional verifiers to reduce contamination and inherited-test failure modes [2]. Humanity's Last Exam uses expert-authored, difficult questions with automated grading to preserve headroom at the frontier [3]. These projects differ in task type and scoring, but each treats benchmark construction and limitations as part of the scientific contribution.
Public model indexes make comparisons accessible at scale, often combining performance with cost and speed metadata [4]. Creative-writing leaderboards such as EQ-Bench provide an important external reference [5]. Canonic differs by refusing to average across creative domains and by publishing its validity failures alongside rankings. Finally, LLM-as-a-judge can provide scalable rubric application but must be treated as a measurement instrument with its own calibration and bias risks [6].
3 The Canonic Stage 1 benchmark
3.1 Domains and task construction
The benchmark comprises 6 domains: book openings, meme captions, movie scenes, sitcom scenes, song lyrics, and stand-up sets. Every task supplies a premise, constraints, and required output form. Tasks draw structural patterns from canonical creative work without quoting source passages. A controlled subset uses novel premises, allowing the benchmark to report the difference between novel and canonically anchored tasks rather than assuming contamination is absent.
The sitcom domain conditions judgments on four published Show Bibles: The Office, Modern Family, FRIENDS, and The Big Bang Theory. This 2×2 design separates single-camera mockumentary and multi-camera ensemble conventions from show-specific voice. It supports an indicative transfer analysis, although each show has only ten tasks per model.
| Component | Configuration |
|---|---|
| Models | 13 pinned frontier, large, mid, and small language models |
| Judges | Two independently prompted LLM judges from different model families |
| Generation | Identical prompt wrapper, temperature 0.7, maximum 4,000 tokens |
| Scoring | Rubric anchors weighted by the harness, not the judge |
| Uncertainty | 95% seeded percentile bootstrap, 10,000 task resamples |
| Publication gate | All tasks attempted and at least 80% counted |
3.2 Freeze and release
Before the full sweep, the rubric set, Show Bibles, and task set were frozen and
content-hashed. The public task-set fingerprint is
418094d0cfd08153da83ca8406793fbe. The generated release manifest binds every
paper figure and table to the source artifact SHA-256
ef9f58b962df1b6cc5b9236c934f66b5bf79c2af1468bd2a7adefc66dd1d7f31. A post-freeze content edit requires a labeled
successor version and a re-sweep.
4 Evaluation protocol
Each model produces one response per task under the same wrapper and generation configuration. Empty, filtered, or truncated output receives one retry; unresolved generation failures are recorded and never interpreted as poor writing. Each valid output is independently scored by two LLM judges against the same domain rubric.
A task counts only when both judges return a score on the standard prompt. Refusals, errors, single-judge outcomes, and reframed retries are excluded with a stated reason. The task score is the mean of the two judge-weighted scores. A model-domain score is the mean of counted tasks. We construct a 95% percentile-bootstrap interval over tasks with 10,000 deterministic resamples. Rows publish only when every task was attempted and at least 80% counted; otherwise the row is withheld without a score.
Models with overlapping intervals share a rank band. This is a conservative rendering rule, not a significance test and not a family-wise-error correction. Crucially, there is no overall benchmark score: a score for a sitcom scene and a score for a book opening measure different rubrics and should not be averaged into an apparently authoritative total.
5 Results
Figure 1 reports scores separately for each domain. It is the principal result: each panel has its own ordering and uncertainty intervals, and no panel contributes to a hidden total. The benchmark evaluates 2,470 cells, of which 2,423 counted under the two-judge rule. 77 of 78 model-domain rows met the publication gate. The remaining row, Kimi K3 on song lyrics, is visibly withheld rather than silently removed.
| Domain | Highest point estimate | Score [95% interval] | Top-band size |
|---|---|---|---|
| Sitcom Scene | Claude Opus 5 | 0.875 [0.865, 0.886] | 2 |
| Stand-Up Set | Claude Opus 5 | 0.834 [0.819, 0.850] | 2 |
| Movie Scene | Claude Opus 5 | 0.861 [0.848, 0.873] | 9 |
| Song Lyrics | Claude Opus 5 | 0.852 [0.832, 0.872] | 1 |
| Meme Caption | Claude Opus 5 | 0.820 [0.793, 0.847] | 13 |
| Book Opening | Claude Opus 5 | 0.887 [0.869, 0.903] | 11 |
The top-band sizes reinforce why headline “winner” claims should be restrained. Within several domains, many model intervals overlap despite different point estimates. The intended reading is therefore conditional: the artifact supports comparisons within a specified domain and protocol, with uncertainty explicitly displayed.
6 Reliability and diagnostic analysis
6.1 Accounting, probes, and contamination
The benchmark publishes data that can weaken trust in its own scores. Figure 2 shows counted and excluded cells by domain and every reward-hacking probe. A probe passes only when its engineered output scores at or below its predetermined maximum. 6 of 8 probes were caught; 2 were not. Those misses are evidence that the corresponding rubric can reward undesirable surface behavior.
| Domain | Attempted | Counted | Excluded | Rows published |
|---|---|---|---|---|
| Book Opening | 390 | 387 | 3 | 13/13 |
| Meme Caption | 390 | 389 | 1 | 13/13 |
| Movie Scene | 390 | 379 | 11 | 13/13 |
| Sitcom Scene | 520 | 516 | 4 | 13/13 |
| Song Lyrics | 390 | 365 | 25 | 12/13 |
| Stand-Up Set | 390 | 387 | 3 | 13/13 |
| Probe | Score | Max permitted | Verdict |
|---|---|---|---|
catchphrase_stuffing/big_bang_theory | 0.256 | 0.35 | caught |
catchphrase_stuffing/friends | 0.231 | 0.35 | caught |
catchphrase_stuffing/modern_family | 0.394 | 0.35 | NOT CAUGHT |
catchphrase_stuffing/the_office | 0.338 | 0.35 | caught |
judge_flattery/the_office | 0.131 | 0.25 | caught |
laugh_track_spam/friends | 0.206 | 0.30 | caught |
length_padding/big_bang_theory | 0.419 | 0.35 | NOT CAUGHT |
structure_without_craft/modern_family | 0.406 | 0.45 | caught |
Novel-premise versus canonical-anchored score splits are reported in the interactive evidence release and appendix. They measure sensitivity to one contamination pathway, not all memorization, cultural familiarity, or training-data exposure. The benchmark does not interpret small deltas as proof that contamination is absent.
6.2 Judge calibration and transfer
The two judges exhibit a measured additive calibration offset while retaining similar aggregate model orderings on the development data. We publish raw disagreement instead of applying an unvalidated correction. Figure 3 reports sitcom scores for each show. Per-show variation is descriptive: with ten tasks per show, it can motivate hypotheses about format and voice transfer but cannot establish them decisively.
| Model | The Office | Modern Family | FRIENDS | The Big Bang Theory | Spread |
|---|---|---|---|---|---|
| Claude Opus 5 | 0.896 | 0.887 | 0.864 | 0.854 | 0.041 |
| GPT-5.6 Terra | 0.849 | 0.809 | 0.841 | 0.833 | 0.041 |
| Grok 4.5 | 0.821 | 0.794 | 0.823 | 0.799 | 0.028 |
| Gemini 3.6 Flash | 0.844 | 0.819 | 0.838 | 0.824 | 0.026 |
| GPT-5.6 Luna | 0.824 | 0.810 | 0.848 | 0.836 | 0.038 |
| Kimi K3 | 0.866 | 0.871 | 0.862 | 0.857 | 0.014 |
| Qwen3.8 Max | 0.817 | 0.832 | 0.831 | 0.844 | 0.027 |
| DeepSeek V4 Pro | 0.818 | 0.795 | 0.828 | 0.830 | 0.035 |
| GLM 5 | 0.815 | 0.817 | 0.825 | 0.823 | 0.010 |
| Llama 4 Maverick | 0.577 | 0.583 | 0.591 | 0.584 | 0.014 |
| Llama 3.3 70B | 0.576 | 0.563 | 0.578 | 0.592 | 0.029 |
| Qwen3 32B | 0.660 | 0.648 | 0.688 | 0.695 | 0.047 |
| Gemma 3 27B | 0.751 | 0.694 | 0.743 | 0.749 | 0.056 |
6.3 External rank comparison
We compare exact model-identity matches to the EQ-Bench Creative Writing leaderboard [5]. The book-opening comparison was predeclared and yields Spearman ρ = 0.933 over 9 shared models. Other domain comparisons are exploratory: Sitcom Scene 0.967 (n=9), Movie Scene 0.950 (n=9), Song Lyrics 0.881 (n=8), Stand-Up Set 0.750 (n=9), Meme Caption 0.733 (n=9). Neither high nor low correlation is validation: similar judge-model evaluations can share preferences, and disagreement alone cannot identify the mistaken instrument.
7 Qualitative analysis
Scores alone do not reveal whether a rubric distinguishes the sort of contrast it claims to measure. We therefore publish task-matched score ladders: low, middle, and high outputs answering the same task. A ladder is eligible only when at least three outputs are publishable; the shown set is drawn from 190 eligible tasks. Both judges placed all three rungs in strict low-to-high order on 104 of these tasks. This denominator matters: the examples demonstrate discriminative cases, not a random sample of every evaluation.
The qualitative appendix preserves the prompt, model identities, per-judge scores, and selected excerpts. It is deliberately separated from claims about populations. A compelling single passage is not evidence of a model-domain population effect, and a poor passage is not a diagnosis of a model family.
8 Limitations and future work
The principal limitation is that Stage 1 has not been validated against blinded human or expert ratings. The LLM judges are useful scaling instruments, not substitutes for demonstrated human agreement — and there is no such figure anywhere in this release, because no study has been run to produce one. Their calibration offset and per-item disagreement further constrain interpretation.
The task sets are modest in size, especially at the sitcom-show level. Bootstrap intervals quantify sampling variability under the observed task distribution but do not settle all multiple-comparison concerns. Canonical-source controls and novel-premise splits mitigate only some contamination pathways. Finally, scores measure the supplied rubrics, which deliberately omit broad claims such as objective funniness.
The next milestone is a preregistered blinded human-ranking study: human experts and non-experts will rank shared outputs, judges will score the same material under a frozen protocol, and the results will be reported regardless of outcome. That future study will not alter the Stage 1 artifact retrospectively.
9 Conclusion
Canonic Stage 1 offers a domain-scoped, inspectable benchmark for creative language generation. Its main contribution is not a universal creativity ranking, but a release discipline: frozen inputs, explicit counting, uncertainty-aware results, and visible failures. As creative evaluations become inputs to model selection and training, these properties are prerequisites for numbers that readers can meaningfully challenge.
10 Release contents and reproducibility
The release contains the validated score artifact, generated paper-data manifest, model and judge pins, frozen rubric and Show Bible hashes, full domain tables, contamination statistics, probe outcomes, transfer statistics, score ladders, and external rank-correlation observations. Raw cached generations and provider credentials are intentionally excluded from the public bundle.
10.1 Artifact identifiers
- Sweep
- full-001 · 2026-08-08
- Artifact SHA-256
- ef9f58b962df1b6cc5b9236c934f66b5bf79c2af1468bd2a7adefc66dd1d7f31
- Task-set hash
- 418094d0cfd08153da83ca8406793fbe
- Public rows
- 77 of 78
- Counted cells
- 2,423 of 2,470
10.2 Interpretation policy
All figures and tables are generated from the same validated artifact that powers the leaderboard. A paper build fails if a required figure or generated fragment is absent. The artifact schema forbids a cross-domain aggregate field, and the paper generator additionally rejects generated text containing an overall-score field.
The frozen fingerprints for all 6 rubrics and 4 Show Bibles are published in full on the method page. The failure record — exclusions by reason, judge disagreement, contamination splits, and the reframe appendix — is on the verification page.
References
- ARC Prize Foundation. ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence. arXiv:2603.24621, 2026.
- W. Huang, C. Lee, L. Tng, S. Ge. DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks. arXiv:2607.07946, 2026.
- L. Phan, A. Gatti, Z. Han, et al. Humanity's Last Exam. arXiv:2501.14249, 2025.
- Artificial Analysis. Artificial Analysis Intelligence Index, 2026. artificialanalysis.ai
- EQ-Bench. EQ-Bench Creative Writing Leaderboard, 2026. eqbench.com
- L. Zheng, W.-L. Chiang, Y. Sheng, et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS, 2023.