CANONIC

Models write the nearest postcard

Ask a model for a place it has read little about and it does not give you vagueness. It gives you the nearest famous neighbour, rendered confidently and in detail. This page shows one instance of that failure and argues it is worth measuring. It does not measure it. The difference between those two verbs is the whole of what we do, so it is marked on every claim below.

Planned Nothing on this page is a measurement. No sweep has run on this, no score exists, and the benchmark described in What we would measure has not been built.

The substitution

A model that has read little about a place has not learned nothing about it. It has learned the region, and the region's most photographed member stands in. The error is not vagueness — vagueness would be honest. It is a confident, fluent, specific description of somewhere else.

The specific substitutions this benchmark would test for, and where each of those things actually is. Every row is a real place; none of them is this one. This is the hypothesis, drawn from the author's own knowledge of the city — not a measured rate, because no rate has been measured.
What comes back Where it belongs
Chinese fishing nets Kochi, 350 km south
Houseboats on still backwaters Alappuzha, Kerala
A cable-stayed landmark bridge nowhere on this estuary
White pleasure craft at anchor the Mediterranean
Manicured resort landscaping Bali

This matters commercially because the people who notice are the people who live there. A model that gets a region confidently wrong is unusable for the market that region represents, and nobody building for that market can currently tell you by how much.


A demonstration, not a result

Two renders from the same video model. The only difference is the prompt. Press play on each — nothing moves until you do.

n = 1 per side. This is a single pair of renders. It illustrates what the failure looks like; it establishes no rate, no ranking, and no score, and it is not comparable to anything on the results boards, every number of which comes from a 2,470-cell sweep with published intervals. If this pair is all you take from the page, take it as a hypothesis.

Generic prompt A tropical coastal scene assembled from the region's most familiar images. Legible, attractive, and not this city.
Grounded prompt Red laterite soil, Mangalore-tiled roofs, a plain girder highway bridge, working trawlers moored three-deep. The particulars, not the region.

Both prompts, verbatim

Reproduced exactly. The second is hand-written, not generated by anything — see What does not exist.

Generic prompt 1 line
Drone view of Mangalore city with a bridge, fishing boats and coconut trees, golden hour.
Grounded prompt ~350 words
Aerial drone photograph of Mangalore (Mangaluru), coastal Karnataka, India — Konkan/Canara
coast, NOT the Mediterranean, NOT Southeast Asia. Golden hour, humid tropical haze softening
the horizon over the Arabian Sea to the west.

GEOGRAPHY AND WATER: The wide, silty Netravati river meeting the sea near Ullal; a long
modern four-lane concrete road bridge with plain girder spans crossing the river — a
functional Indian highway bridge (NH-66), not a suspension bridge, not a cable-stayed
landmark, no arches. Brown-green river water with sandbars visible at low tide. Estuary
mouth with a dark stone breakwater.

BOATS: Wooden deep-sea fishing trawlers and purse-seiners moored three-deep at the Old
Bunder fishing harbour — high flared bows, hulls painted in bold colour bands (blue, red,
green, white), registration numbers stencilled on the bow, small flags and faded tarpaulin
canopies. Crowded, working boats — no yachts, no white pleasure craft, no gondolas, no
Mediterranean sailing boats.

BUILT ENVIRONMENT: Dense low-rise cityscape climbing gentle laterite hills — red clay
Mangalore-tiled roofs (flat interlocking terracotta tiles) on older houses and shops,
mixed with plain concrete mid-rise apartment blocks in white/cream/pastel; compound walls
of exposed red-brown laterite blocks; church spires and temple gopurams and mosque minarets
sharing the same skyline. Narrow roads with auto-rickshaws and buses. Shop signboards in
Kannada script alongside English.

VEGETATION AND SOIL: Coconut palms in dense groves along the water and between houses —
tall, slightly leaning, interspersed with areca nut palms, jackfruit and mango trees.
Exposed red laterite soil on cuttings and open ground. Lush deep-green monsoon vegetation,
not manicured resort landscaping.

LIGHT AND MOOD: Warm low sun over the sea, long shadows across tiled roofs, haze and
woodsmoke in the air, everyday working-city energy — a lived-in Indian port city, not a
tourist postcard.

NEGATIVE: no European architecture, no Bali-style resorts or rice terraces, no Chinese
fishing nets (that's Kochi), no houseboats (that's Kerala backwaters), no skyscraper
skyline, no suspension bridge.

Why this is a measurement problem

The obvious reading of the demonstration above is "write longer prompts". That reading is why this is worth building rather than blogging about.

Somebody had to know that Mangalore has laterite soil and interlocking terracotta tiles and a plain girder bridge, and had to know to say not Kochi's fishing nets, not Kerala's houseboats. The grounded prompt is not a prompting technique. It is a compressed piece of local knowledge, hand-written by someone who has the knowledge, for one place, once.

Three questions follow, and none of them is answerable by looking at two videos: which places fail, how badly, and whether supplying the knowledge reliably fixes it. Those are measurements. The instrument for taking them is a benchmark, which is the thing this company already knows how to build.


What we would measure Planned

Specified, not built. A seventh text domain on the harness that already runs the other six, so it inherits the two-judge counting rule, the bootstrap intervals, the freeze discipline and the atomic publish gate without inventing any new statistics.

Reference Bibles
A frozen, content-hashed document per subject, injected verbatim into the judge's prompt. The judge scores fidelity to that document, never to its own knowledge. This is not a new idea here: it is exactly how the sitcom domain already scores fidelity to a Show Bible, and it exists because the judges' own knowledge of an under-documented subject is precisely the thing under test.
A declared salience axis
Every subject is classified well-known or not, against a rule written down before any subject is chosen, and published with its justification and the date it was checked. The finding is the score gap between the two groups — with its own confidence interval, and if that interval crosses zero we publish that there is no established gap.
Subjects chosen to be hard
Roughly fifteen, deliberately niche and deliberately spread — geography, food, festivals, rituals, crafts, and regional speech, across regions — with a small well-known control group so a low score is attributable to salience rather than to the documents being unusually demanding.
A probe that can kill it
One adversarial fixture writes the nearest famous neighbour, fluently and confidently — Kathakali in place of Theyyam. If the rubric scores it well, the benchmark does not detect the one error it exists to detect, and that blocks publication rather than shipping as a caveat.

One thing that rules itself out: no dimension asks whether writing feels authentic. A rubric rewarding that would reward exoticism — a model writing colourful prose about a place it knows nothing about would beat one writing plainly and accurately, which is the exact inversion of the finding. Every dimension scores traceability to a published document instead.


What does not exist

This page argues from two videos and a plan. Here is everything it would be reasonable to assume from that, and is not true today.

  1. No sweep has run, and there is no score.

    Not one model has been evaluated on any of this. Every number on this site comes from the six-domain sweep and none of it touches this work.

  2. The corpus does not exist.

    No Reference Bible has been written, no task authored, no rubric built. The benchmark's own corpus is currently 100% US and Western in origin.

  3. There is no context engine, and no profile format.

    No module, no schema, no retrieval code, no API, no tests. The grounded prompt above was typed by a person. Calling it "what a Reference Profile delivers" describes an intention, not a system, and no such system is being sold.

  4. We measure text, not video.

    Every domain, both judges and the entire harness are text. The renders here show the failure because it is visible in a video; checking a video automatically would be a second harness that has not been designed. Nothing on this page should be read as a claim to evaluate images or video.

  5. Nobody has verified that supplying context helps, at scale.

    The comparison that would show it — running every task with and without the reference document and publishing the gap — is deliberately scoped as separate work, to be done after the baseline exists and is frozen. Until then the pair above is a hypothesis with n = 1.

If your market is one models get wrong

The useful conversation is about which places or domains your users would notice being wrong, and what evidence would convince you either way. What we can offer today is the instrument, and a public record of where it fails.