Contents

The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence

Abstract

The metanym game is a competitive word game for LLMs that measures structural intelligence against established cognitive-science constructs. No content is given in advance; the contestants create all of it — a new kind of analogy test, analogical production falsifiable sentence by sentence, with no fixed test set to leak into training (contamination-resistant by construction). In the council-of-peers benchmark, the contestants also rate each other’s creations. We introduce the first spectral solution, to our knowledge, to the wicked problem of benchmarking LLMs’ factual accuracy without golden keys or oracle models: one singular value decomposition of the evaluators’ ratings matrix yields their competence as both generators and judges of true statements at once. Competence on the subjective criteria comes from each judge’s rating consistency as the yardstick shifts. The factual rating correlates with GPQA Diamond at Pearson r = 0.92. Scored separately, making and judging dissociate — judging is the scarcer skill: the strongest generators are middling judges, the sharpest judge a mid-pack generator. To scale, the strongest players form a council that does the official benchmarking; its seats are contestable — a stronger model earns one on the benchmark’s own rating. The benchmark is entirely self-contained and self-consistent, a stable gauge over time.

Scatter of GPQA Diamond accuracy against the key-free combined factual rating for twelve models, with 95% confidence bars on both axes and a fitted line. Council members are filled circles, non-council hollow, the anchor a star. Pearson r = 0.92, 95% CI 0.85 to 0.97; Spearman rho = 0.91; n = 12.
Two instruments, no shared machinery. The key-free rating on one axis, a human-keyed benchmark on the other. r = 0.92, 95% CI [0.85, 0.97].

The comparison is not part of the benchmark. GPQA Diamond is a fixed set of 198 graduate-level questions, hand-written and validated by PhD experts and scored against their answer key; on it those experts themselves reach 65%. The metanym game writes its own items every run and scores them against nothing outside the panel. The two agree — and neither is assumed to be the more accurate instrument. Delete GPQA from the record and every rating is unchanged, because no rating was ever derived from it.

Introduction

Every benchmark for machine intelligence leans on an answer fixed in advance — a gold label, a human rating, a reference solution. The benchmark reported here has none. Twelve frontier language models invent the test, sit it, and grade one another; the benchmark then works out which of them are competent to grade at all. No human raters, no answer key, nothing to look up.

The test is the metanym game. A player takes a paragraph describing one domain and turns it into a factually true description of an unrelated one by swapping a handful of words and leaving the rest untouched: a description of cell signalling becomes, swap by swap, a true description of human language, and then of a microservice architecture — each sentence checkable on its own. The unchanged wording is a context template, the swapped words are metanyms, metaphorically synonymous, and each rewrite is a parallel context of one underlying archetypal context. The player chooses from no menu; it builds the structure from nothing. The metanym game makes a formal test of something a long tradition treats as central to thought — seeing one structure across wildly different domains. Where that tradition tests whether you recognise it, the game tests whether you can build it, and checks the result sentence by sentence — a new kind of analogy test, not just a harder one.

Two properties follow from building the items this way. Because every item is produced fresh in the run, there is no fixed test set to leak into a later model’s training data: the benchmark is contamination-resistant by construction (§3). And because correctness is settled sentence by sentence — does this swapped claim hold in its new domain? — the game needs no answer key. The models supply the verdicts themselves, and the benchmark reads the truth off their agreement: stack every model’s true/false judgements into one matrix, and its dominant direction reveals which judges are competent, with no labels at all (§4.3, Appendix A). That competent subset becomes the council that grades everyone — a benchmark that certifies its own judges.

A single run then makes three things visible (§4). The twelve models split into a leading eight and a trailing four. The divide follows provider lineage more than parameter count: one vendor’s older generation falls away together with the roster’s smallest seat, while within each family the size gradient is shallow. Judgement is the bottleneck: most models cannot reliably tell a true cross-domain claim from a false one, even when they produce competent structure themselves, so a model can be a perfectly consistent grader and still a wrong one. The strongest players — the models that clear the reliability bar — are seated as the council that issues the official ratings.

The game has a lineage — in the cognitive science of analogy and structural mapping (§2), and in the systems-theory claim that one abstract structure can recur across unlike domains (von Bertalanffy 1968). The rest of the paper builds the game (§2), turns it into the self-administering council benchmark (§3), reports the canonical twelve-model run (§4), and weighs what the numbers do and do not license (§5).


The metanym game

Archetypal contexts

Take a passage that is true of one domain. Strip out the words that tie it to that domain and leave the roles they filled, named and empty. What remains is a context template: slots held in a relation that does not change when the domain does.

The template is not the idea. It is the literal representation of an archetypal context — the abstract system that several unrelated domains turn out to share. That system has no vocabulary of its own, which is why it cannot be written down directly. The template is one way of writing it down; there are others, and none of them is the thing itself.

Fill the slots with the vocabulary of one field and you instantiate a parallel context. Fill them again from a field with nothing in common with the first, and again from a third. The words that land in the same slot across those fillings are metanyms — metaphorically synonymous, which is what the coinage contracts. One complete set of them is a metanym set; several sets tabulated against the shared slots make a metanym table.

Every parallel context comes in two forms. The raw instantiation is the template with a metanym set slotted in verbatim — nothing has moved but the keywords, which is what makes the claim checkable. The idiomatic rewrite is that same parallel context said the way its field would actually say it, which shows the structure survives being spoken naturally rather than depending on the borrowed phrasing.

Then the test: every sentence of every parallel context has to be literally true in its own domain. A filling that is not true has failed, and can be shown to have failed, one sentence at a time.

Below is one template and four metanym sets, chosen to lie about as far apart as domains get. Move the pointer over it: the row shows one slot across all four domains, the column fills the passage beneath with that domain’s vocabulary.

One template, four domains

Consider one context template whose slots are named as general-systems roles — an organizing structure, the components it organizes, their coupling, the emergent whole, and so on — instantiated across four cases chosen to lie about as far apart as cases can: Jung and Pauli’s cosmic archetypes, von Bertalanffy’s General Systems Theory, the archetypal contexts of this paper, and the baking of bread.

Metanym table
One context template, its eight slots, and the four metanym sets that fill it. Choose a column to instantiate the passage below.
ORGANIZING STRUCTUREcosmic archetypesstructural isomorphismsarchetypal contextsbaker's percentages
COMPONENTSmind and mattersystem componentsdomain keywordsraw ingredients
COUPLING DYNAMICSacausal synchronicitiesdynamic interactionscontextual templatesthermal and biochemical reactions
EMERGENT WHOLEthe unus mundussystemic homeostasisfunctional equivalencestructural leavening
APPARENT DISORDERa fragmented dualitydisconnected phenomenasemantic isolationculinary chaos
MODELLING SCIENCEdepth psychophysicsgeneral systems theorymetanymic analysisfood science
INSTANCEhuman subjective experienceindividual open systemsspecific domain jargonsan individual bake
GENERAL LAWa continuous psychophysical realityuniversal laws of organizationscale-recursive abstract systemsthermodynamic and chemical laws
Context Template

The fundamental structure of a system is defined by ORGANIZING STRUCTURE, an invisible framework that dictates the organization of COMPONENTS. As these components interact through COUPLING DYNAMICS, they generate a unified state of EMERGENT WHOLE. Without recognizing this inherent design, the system is mistakenly perceived as APPARENT DISORDER. However, by applying the principles of MODELLING SCIENCE, we uncover that these structural patterns are not isolated phenomena. Instead, the specific relationships observed within INSTANCE are actually localized expressions of GENERAL LAW.

Idiomatic Rewrite

A template has no idiom of its own — that is the point of it. Put a metanym set in focus and its rewrite appears here, beside the same sentences filled with that domain’s vocabulary.

Semantic similarity from LaBSE semantic sentence encoder

Slot each column into the template and every sentence holds. Notice why it holds so widely: the slots are named as general-systems roles — an organizing structure over its components, their coupling, the emergent whole, the apparent disorder a naïve eye sees, the science that models it, an instance, and the law that instance expresses. Read that way, the template is the generic schema of a systems-science explanation, so almost any system studied by a discipline instantiates it — a psyche–matter unity, an open system, our own framework, and an afternoon’s baking alike.

This is therefore a very general archetypal context: it fits an enormous range of cases precisely because it encodes the bare form of structured explanation. That breadth is the unimpressive end of the spectrum — the archetypal contexts that matter most are far more discriminating, fitting one relational structure and excluding its neighbours (§2.c). What the example fixes is only the machinery: one template, filled by mechanically swappable metanyms, staying true sentence by sentence across maximal domain distance — which is what makes a metanym game decidable, and therefore measurable.

Terminology

Six terms, in the order they depend on each other, and one property that turns them into a test.

Term Level What it is
Archetypal context the abstract object The abstract system that several unrelated domains share. It has no vocabulary of its own and cannot be written down directly — only represented. The cross-domain isomorphism General Systems Theory studies (von Bertalanffy 1968).
Context template its representation The literal representation of an archetypal context: slots held in a relation unaltered across domains. Literal, but wordable many ways — the template is a picture of the archetypal context, never the thing itself.
Metanym the filler A keyword filling the same slot as another domain’s keyword, and metaphorically synonymous with it. The coinage contracts METAphorically synoNYMous.
Metanym set one column One complete set of metanyms — enough to fill the template for a single domain.
Parallel context the instantiation What a metanym set produces once slotted in: a description of one domain generated by the shared template. Two domains’ parallel contexts are each other’s metaphors. Every parallel context exists in two forms —
  · raw instantiation Form (a) the template with the metanym set slotted in verbatim. Nothing but the keywords has changed, which is what makes it checkable.
  · idiomatic rewrite Form (b) the same parallel context in the prose its field would actually use, showing the structure does not depend on the borrowed phrasing.
Metanym table the apparatus Several metanym sets tabulated against the shared slots — the exhibit above.

And the property that makes it an instrument rather than a notation:

Per-sentence factual truth Every sentence of every parallel context must be literally true in its own domain. This is what the paper claims is new: the abstract structure made mechanically checkable, sentence by sentence, in a task where the player builds the structure rather than recognising one they are shown.

The distinction carrying the most weight is the first two. An archetypal context cannot be written down; a context template is what you write instead. If a sentence still reads correctly with those two terms exchanged, it is wrong.

The peer council benchmark (abridged)

The rules

The metanym game is played by N players and a non-competing administrator, and has two elements.

1. Generation. A player creates archetypal contexts from scratch: N context templates, M metanym sets per template, and for each set the instantiated template (Form (a)) and an idiomatic rewrite that reads naturally (Form (b)).

2. Evaluation. A player scores other players’ submissions. To make the result a rating rather than a popularity vote, each submission is graded on the rubric axes (§3) against one fixed reference submission pinned at an anchor value, with the anchor swept across {5, 6, 7, 8} — the only thing that changes between passes. Run over a common submission set (in this paper, the council members’ portfolios), a single evaluation round yields two ratings at once: the submission ratings (each portfolio, aggregated across evaluators) and the evaluator ratings — how well a judge detects the factual errors the panel collectively flags (factual competence) and how stable a standard it holds for the non-factual criteria as the anchor shifts — a competent judge’s ranking is invariant under that non-semantic change (criterion reliability, measured by anchor-shift consistency). Both evaluator ratings read only a judge’s scores of the other players’ submissions, never its own.

The two elements are deliberately complete — a player generates and judges, and each act is itself rated — so the framework is fully self-contained: no human raters, no external answer key, each part producing one of the benchmark’s ratings. Together the two elements place a conjunctive demand on a sizeable cluster of capacities that cognitive science treats as central to intelligence, set out under Intelligences covered and mapped back to these two elements.

3.1 Setup

Twelve frontier LLMs from three providers serve simultaneously as generators and as members of the evaluator panel:

Provider Models
Anthropic claude-opus-4.5, claude-opus-4.1, claude-opus-4.0, claude-sonnet-4
Google gemini-3.1-pro, gemini-2.5-flash
OpenAI gpt-4.1-2025-04-14, gpt-4.1-mini, gpt-4.1-nano, gpt-4o, gpt-4o-2024-08-06, gpt-4o-mini

Table: The twelve-model council-of-peers panel, by provider.

This is a deliberately heterogeneous panel, assembled from the models available to us and chosen to exercise the structural cases the protocol must handle rather than to census the frontier. It spans roughly an order of magnitude in scale, so competence and reliability have room to separate; it draws on three vendors, so the cross-vendor agreement the no-key reliability measure rests on can be tested rather than assumed; it includes several models from one vendor and adjacent versions within a single family (Opus 4.0 / 4.1 / 4.5), which stress the protocol’s resolution and its safeguard against same-vendor agreement masquerading as competence; and it spans size tiers within a family, a known capability gradient the ranking should recover. Because the benchmark is a re-runnable protocol whose ratings are panel-relative, no conclusion depends on this particular roster — and drawing on what was at hand, rather than a curated set, removes any concern that the panel was chosen to flatter the method.

All twelve are called with Temperature=0, reasoning/thinking disabled, and tools disabled. This choice fixes the test on the model’s base capability and removes three confounds at once. Determinism (T=0) makes N=1 per cell sufficient — re-running produces bit-identical output. No reasoning tests the model’s direct response, not the output of an internal deliberation loop that varies between providers in opaque ways. No tools removes external-information channels that could leak factual content the model itself does not represent.

3.2 Protocol

Each model generates one portfolio: five archetypal contexts, each with a context template (5–8 sentences with UPPERCASE [SLOT] labels, 6–10 slots) and a metanym table of five domain columns, yielding 25 parallel-context instantiations per portfolio.

Each model then evaluates every other model’s portfolio under a six-axis rubric:

Axis Granularity What it measures
factual_per_pc per parallel context factual defensibility of the substituted text in its target domain
beauty per archetype aesthetic quality of the context template
intelligence per archetype depth and non-triviality of the abstraction
instantiation_distinctness per archetype “Domains far apart / metanyms not synonymous”
impressive_length per archetype template length and slot count
structural_diversity per portfolio how different the five archetypes are from one another

Table: The six-axis evaluation rubric.

Scores are on a 1–10 cardinal scale. Each evaluator call presents one anonymised target portfolio alongside a fixed anchor portfolio pinned at 7 on every axis; the evaluator scores the target relative to the anchor. The full evaluation yields a 12×12 evaluator-by-generator matrix.

Why these settings. Four design choices justify themselves on first principles.

(i) Per-submission cardinal rating rather than side-by-side ranking. 25-PC portfolios already press against context windows when more than one is present, and prompt-internal attention is uneven; per-call rating sidesteps both at once.

(ii) Calibration against a fixed anchor. Cardinal scores drift between evaluators — one model’s “8” is another’s “6”. Pinning a fixed reference portfolio at a known score on every axis turns each evaluator’s idiosyncratic scale into a common one and recovers discriminability at the top, where the 1–10 ceiling otherwise compresses the strongest portfolios into an indistinguishable cluster. §3.4 explains how the anchor portfolio itself is chosen.

(iii) Holistic axes, not analytic decompositions. The five non-factual axes are high-level concepts (beauty, intelligence, instantiation_distinctness, impressive_length, structural_diversity), not sub-criteria. Two reasons.
Principled: a detailed scoring rubric is also a template-construction tutorial — generators must be told how submissions will be rated, so every clause in the rubric leaks back into the generation prompt as guidance about what to produce. We want to score what models recognise as beautiful or intelligent, not what they can be coached to construct.
Empirical: frontier models agree most tightly on the most holistic judgement. Across the un-anchored 12×12 matrix, mean inter-evaluator standard deviation per cell was lowest for beauty (1.07 on the 1–10 scale), then factual_per_pc (1.11) and intelligence (1.15); the more concrete instantiation_distinctness (1.25) was the least consistent. Models converge on high-level judgements without a checklist.

(iv) Minimal prescription overall. Every additional directive in an evaluator prompt measurably shifts the score distribution, so prescription is held to what the protocol requires.

(v) impressive_length counterweights per-sentence factual scoring. Without it the dominant strategy is the minimal template — fewest sentences, least error exposure — and a leaderboard scored without it would advantage short templates. A longer template that stays true in every sentence is the harder accomplishment, and padding is not free: every added sentence is another claim factual_per_pc scores.

3.3 The self-governing benchmark

In its steady state, the benchmark operates as a self-governing protocol run by the council. A council of LLM evaluators (five in the canonical run) scores any submitted portfolio against a fixed anchor reference on the six-axis rubric of §3.2. The protocol is simple:

Each council member receives two ratings: a generation rating (the LSO mean of the other council members’ scores of its portfolio) and an evaluator rating — the factual-competence and criterion-reliability scores from the evaluator-rating routine (§4.3), themselves LSO in the same sense, since a model’s ratings of its own portfolio are excluded from its own reliability — reported separately and never merged with the generation rating. Non-council models receive only a generation rating, computed by the same council against the same anchor.

The benchmark scales by addition. Any future model — open-weights, next-generation, or external — can be evaluated against the same published anchor by the same council without re-deriving anything. It is also self-administering (no human evaluators or gold key, so re-runs are not bottlenecked on human labelling) and reproducible (at T=0 with no reasoning channel, every cell is bit-identical on re-run, so the same anchor and the same council produce the same leaderboard on demand). Bit-identical reproducibility is a property of the seats, not the protocol: models that deprecate the temperature control cannot be pinned to T=0, so a council holding such seats reports N>1 samples with intervals instead — protocol and anchor unchanged.

Contamination. Items are generated fresh each run, so no fixed test set can leak into training. Published past submissions could enter training corpora — a leak touching generation only, so a suspiciously large generation–evaluation gap is itself the detector, and new portfolios are screened against the archived submissions of record. Format familiarity is not contamination: every model tested understands the task as posed; what is scored — the items — is new each run.

The leaderboards (abridged)

Three leaderboards over the same round, each ranking the same twelve models on a different thing: the total rating, the four components behind that total, and making against judging on each non-factual criterion.

Final leaderboard

Final leaderboard — total rating T (95% CI) with its evaluator half E and generator half G, all twelve models ranked; council seats marked. The anchor (claude-opus-4.5) is 7 by construction.

RankModelCouncilT95% CIEG
1★ claude-opus-4.5 (anchor)council7.00[7.00, 7.00]7.007.00
2gemini-3.1-procouncil6.69[6.56, 6.87]7.226.16
3claude-opus-4.0council6.21[5.93, 6.61]5.446.98
4claude-opus-4.1council6.05[5.73, 6.52]5.047.07
5gemini-2.5-flashcouncil5.76[5.17, 6.21]5.376.15
6claude-sonnet-45.30[5.04, 5.56]3.966.65
7gpt-4.1-mini4.74[4.09, 5.45]3.865.62
8gpt-4.1-2025-04-144.44[4.21, 4.63]3.105.78
9gpt-4o-2024-08-063.48[3.23, 3.75]2.654.32
10gpt-4.1-nano3.22[2.52, 3.62]2.793.65
11gpt-4o2.93[2.61, 3.20]1.584.28
12gpt-4o-mini2.24[2.04, 2.44]1.173.31

Each half resolves into two anchored competences — making: generator factual GF (§4.3 Criterion A) and criterion GC (the five non-factual generation axes, §4.5); judging: evaluator factual EF (§4.3 Criterion A) and criterion EC=7ρ̄/ρ̄a (the leave-self-out collapsed anchor-shift consistency, §4.4 Criterion B):

Competence breakdown

Competence breakdown — the four anchored components behind each model's evaluator (E) and generator (G) scores.

RankModelCouncil?GFGCEFEC
1★ claude-opus-4.5 (anchor)council7.007.007.007.00
2gemini-3.1-procouncil6.585.757.367.07
3claude-opus-4.0council6.986.984.476.42
4claude-opus-4.1council6.977.163.556.52
5gemini-2.5-flashcouncil6.575.724.686.07
6claude-sonnet-46.986.311.626.30
7gpt-4.1-mini6.185.061.216.50
8gpt-4.1-2025-04-146.505.050.206.00
9gpt-4o-2024-08-065.223.430.624.67
10gpt-4.1-nano3.663.630.495.09
11gpt-4o5.113.440.342.82
12gpt-4o-mini3.203.430.002.35

(GF, EF are the §4.3 soft-SVD factual competences (anchored 7f/fa and the per-generator consensus); GC is the council reliability-weighted leave-self-out mean of the five non-factual generation axes; EC=7ρ̄/ρ̄a the leave-self-out collapsed anchor-shift consistency (§4.4). Values are anchored point estimates (claude-opus-4.5 = 7); every quantity is carried unrounded to the last step, so E=½(EF+EC), G=½(GF+GC) and T=¼(GF+GC+EF+EC) are computed from unrounded leaves and a displayed total can differ from its rounded inputs by 0.01. The T interval is the 95% joint (submission, archetype) bootstrap of A.5 — a CI device only, it does not change the point estimates; an independent end-to-end re-derivation reproduces the competent-model components, while inert-band factual loadings — which §4.3 notes are not robustly distinguishable from zero — are construction-sensitive.)

Per-criterion generator quality (G) vs evaluator reliability (E)

Per-criterion generator quality (G) vs evaluator reliability (E).

Modelbeauty Gbeauty Eintel Gintel Edist Gdist Elen Glen Estruct Gstruct E
★ claude-opus-4.57.07.07.07.07.07.07.07.07.07.0
gemini-3.1-pro5.86.85.77.46.07.35.56.95.77.4
claude-opus-4.06.96.76.97.07.05.76.86.37.37.1
claude-opus-4.17.16.87.27.07.46.06.76.57.37.6
gemini-2.5-flash5.46.25.55.56.24.25.76.35.87.0
claude-sonnet-46.36.16.36.46.56.26.16.06.46.7
gpt-4.1-mini4.86.35.06.25.96.34.26.55.46.7
gpt-4.1-2025-04-145.15.25.06.06.06.23.66.35.67.0
gpt-4o-2024-08-063.54.73.34.54.13.13.34.03.05.7
gpt-4.1-nano3.55.13.75.34.23.73.52.83.22.9
gpt-4o3.42.63.42.04.71.02.63.13.22.8
gpt-4o-mini3.52.13.53.03.32.33.80.83.01.9
cos(G,E)0.91[.83,.95]0.90[.80,.93]0.89[.82,.92]0.84[.80,.87]0.87[.86,.87]

The five non-factual axes are beauty, intel (intelligence), dist (instantiation distinctness), len (impressive length) and struct (structural diversity). Per criterion: E = evaluator criterion reliability (the leave-self-out anchor-shift consistency of the Criterion B table, rescaled to the 1–10 scale), and G = generator quality — the council's leave-self-out mean of its scores of that generator on that axis, each council member's vote weighted by its own reliability E on that axis (eq A12b; the same weighting §4.5 applies to GC), so E both scores the judge and sets its weight in G. Both are anchored so claude-opus-4.5 (★) reads 7. cos(G,E) = cosine of each model's deviation from the anchor point (7,7) across the panel, with 95% CI from the joint (submission, archetype) bootstrap of A.5; 1 = making and judging perfectly aligned on that axis. G, E and the cosine are all formed from unrounded quantities and rounded once, here. Higher G = better maker; higher E = more stable judge.

What intelligences does the benchmark cover?

The game demands a specific kind of intelligence.

It tests abstraction and analogy, among the most widely accepted lenses on general intelligence in AI (Chollet 2019; Mitchell 2021; Lake et al. 2017). The eight constructs below are the ones the game directly demands. The table maps each to the element or elements that call on it; hover a row for the construct as its authors describe it, and the demand the game places on a player.

The eight constructs are summarised below, mapped to the two tasks that call on each (● marks a primary demand; references are in each construct’s panel). The pattern of shared cells previews §5.4: structure-mapping and higher-order relational reasoning run through both tasks; generation alone calls on essence-seeing, fluid intelligence, divergent production, and theory-by-analogy; evaluation alone turns crystallised knowledge and convergent production toward judgement.

ConstructGenerationinvent the templateEvaluationjudge a portfolio
Higher-order relational reasoningPenn, Holyoak & Povinelli (2008)Higher-order relational reasoningPenn, Holyoak & Povinelli (2008)Recognising when two situations share the same pattern of relations among their parts, even when the parts themselves are unrelated. The canonical test is to see that AABB and CCDD share the structure "two pairs of matching things" despite A, B, C and D being different objects — a capacity the authors argue most cleanly separates human from non-human reasoning. The metanym game tests the same thing: each slot is defined by its relations to the other slots, not by its filler word, and successful substitution shows the relational pattern survives.lay out the relational skeletoncheck the relations survive
Structure-mappingGentner (1983); Falkenhainer, Forbus & Gentner (1989)Structure-mappingGentner (1983); Falkenhainer, Forbus & Gentner (1989)Three constraints govern analogical alignment: systematicity (Gentner 1983), one-to-one correspondence, and parallel connectivity (the latter two formalised in the structure-mapping engine; Falkenhainer, Forbus & Gentner 1989). The metanym table is the structure-mapping bookkeeping written down — rows are slots, columns are target domains, each cell is a one-to-one mapping. Mechanical substitutability enforces one-to-one correspondence and parallel connectivity; systematicity — Gentner's stronger demand that higher-order relations constrain the mapping — is not guaranteed by substitution and is judged rather than assumed (the `intelligence` axis).the slots-and-domains scaffoldverify one-to-one correspondence
Essence-seeing (analogy as core cognition)Hofstadter & Sander (2013)Essence-seeing (analogy as core cognition)Hofstadter & Sander (2013)Essence-seeing — spotting that a novel situation is structurally an instance of a known abstract pattern despite different surfaces — as the mechanism of cognition rather than a special-purpose module. In the metanym game, it means seeing the essence of the context template, which is the archetypal context.see the archetype behind the surfacenot demanded
Fluid intelligenceCattell (1963); Horn & Cattell (1966)Fluid intelligenceCattell (1963); Horn & Cattell (1966)Reasoning on the fly over an unfamiliar context and drawing novel conclusions instead of retrieving them. In intelligence tests, this is probed by 'what is the next shape in the series?' or 'fill the missing slot with the correct symbol'. The metanym game is the verbal version: identifying a metanym set for instantiating a factually correct parallel context.reason out a novel structurenot demanded
Crystallised intelligenceCarroll (1993); McGrew (2009)Crystallised intelligenceCarroll (1993); McGrew (2009)Having knowledge and knowing how to use it. In intelligence tests, this is probed by quiz-style questions whose answers cannot be worked out, only known. The metanym game probes it twice over: each metanym must be a word the player knows the meaning of, and the resulting sentence must be factually true in its domain. Knowing a broad vocabulary and a deep store of domain facts is a strength in the metanym game.not demandeddetect false claims
Convergent productionGuilford (1967)Convergent productionGuilford (1967)Generating the single correct answer that converges from many constraints (canonical example: "man : woman :: king : ___"). In the metanym game, mechanical substitutability of metanyms in the context template admits no near-misses on the evaluation side: either the instantiated sentence is structurally coherent and factually defensible (passes) or it isn't (fails).not demandedpass/fail, no near-misses
Divergent productionGuilford (1967)Divergent productionGuilford (1967)Knowing how to use the same knowledge (or word) in many different and novel ways, setting out from the same starting point. In the metanym game, the same context template is applied in widely different domains, with each domain represented by its own metanym set.invent a structure that travelsnot demanded
Theory formation by analogyHesse (1963); Boyd (1979)Theory formation by analogyHesse (1963); Boyd (1979)Hypothesis-by-analogy drives scientific exploration. Hesse showed that theories work by extending a known analogy into unmapped territory to generate predictions; Boyd argued that some core concepts — brain as computer, gene as codeare their analogy, with no separable literal core. Each archetypal context is theory construction in miniature: the context template is the root analogy, and the parallel contexts are its metaphors. The factuality of each metaphor is the empirical test, and cross-domain span is the demand that the structural claim survive surface-disparate domains.template = root analogy, tested by factnot demanded
Two constructs are demanded by both tasks; four by generation alone and two by evaluation alone. The two halves of the game do not call on the same abilities — which is the a-priori reason the making and judging scores come apart, before any data was collected. Hover or focus a row for what the construct is.

Production-level character. The canonical instruments of the cited traditions are largely recognition tasks: PHP’s higher-order task is match-to-sample; Gentner’s structure-mapping engine models how humans interpret a given source-target analogy; Hofstadter & Sander present essence-seeing through case studies. The metanym game is a production task — the participant generates the template and the metanym sets from scratch. Production is a stronger demand than recognition — one can sometimes pass a recognition task by elimination or surface heuristics; production has no such fallback. The framework shifts the test from recognition and interpretation to production — a heavier cognitive register than the prior literature’s canonical instruments.

What is new here. Each ingredient of the context template has a neighbour. Slot-bearing templates are frames (Fillmore 1982; Minsky 1975) and constructions (Goldberg 1995); a schema standing above several parallel instances is the induced problem schema of Gick & Holyoak (1983); an abstract structure recurring symmetrically across unrelated domains — the archetypal context itself — is General Systems Theory’s isomorphism (von Bertalanffy 1968). What none of them carries is the conjunction the metanym game demands: one template re-bound jointly across five-plus domains with no privileged source, under the demand that every substituted sentence remain literally true — a per-sentence falsifiability test the analogy, schema-induction, and metaphor traditions never operationalised (conceptual metaphor in fact requires literal falsity; Lakoff & Johnson 1980). The novelty is the test: the abstract structure made mechanically checkable, sentence by sentence, in a production task.

Discussion

5.1 A benchmark by LLMs, for LLMs

The aim is a benchmark that needs nothing outside itself: models invent the test, sit it, grade it, and certify which of them are fit to grade — no human raters, no gold key, and no oracle model whose word is taken as truth. LLM-as-judge already removes the human rater (Zheng et al. 2023; Liu et al. 2023; Verga et al. 2024; Bai et al. 2023) but still requires an external ground truth — a gold key or reference answer — to score against. The metanym benchmark removes that dependency too — but key-free grading is not the novelty; the unsupervised peer-evaluation line (PiCO, UPME) already gets there (§5.5). What is new is self-containment: the panel authors the very items it judges, so the test refers only to itself, and one decomposition scores the models as both makers and judges (§5.5). The benchmark is in this sense fully self-contained.

The metanym benchmark correlates excellently with GPQA. The latter is no golden key, and does not come across as more accurate than the metanym benchmark. Both occasionally reverse the expected internal ranking order within families of LLMs, such as placing a newer version of a model below its predecessor. While most of these instances are within the error margins, two examples fall outside: GPQA ranks Claude-sonnet-4 above Claude-Opus-4.0 while the metanym benchmark places opus comfortably above sonnet. Both benchmarks rank gpt-4.1-mini above gpt-4.1.

Both benchmarks will increase resolution following the same principle, but only the metanym benchmark does it easily. For GPQA it means engaging domain experts to increase the number of Google-proof multiple-choice questions, re-validating the test and publishing a new version of the benchmark. For the metanym benchmark the resolution is increased by twisting a knob, raising the number of archetypal contexts in a submission (here set to five).

5.2 Two self-consistencies, two yardsticks

With no outside ruler to appeal to, the yardsticks must come from the system’s own structure. They come in two forms — one for objective criteria, one for subjective — which is why the method uses two estimators, not one.

The factual estimator rests on one assumption: the only thing competent evaluators share is the truth. When that holds, agreement concentrates on truth, and the leading eigenvector of the leniency-removed agreement matrix is the competence axis. The warrant holds for facts and fails for taste: on beauty, the dominant axis of agreement is no longer truth but shared convention — house style, training data — so weighting by it would launder conformity into competence, rewarding the evaluator nearest the mean and penalising a legitimate minority view. The subjective criteria therefore use the other route, anchor-shift consistency: the calibration reference is swept, and a competent evaluator preserves its ranking as it moves, without needing to agree with any peer. The test has teeth precisely because the shift is non-semantic: if merely moving the calibration point reorders how a model rates the same items, it has no firm grip on what it is judging — and a stable standard is the whole of what competence means on a criterion with no external truth. One route is a consensus eigenmode across evaluators, the other an invariance within each; both are self-consistency conditions set by the system’s own structure rather than an absolute scale.

5.3 A sustainable yardstick

The benchmark yardstick is calibrated on the anchor submission (here the submission by Claude-opus-4.5) and the official benchmark ratings are set by the council. With model temperature T=0, no thinking or tools, the council members’ evaluations are deterministic and reproducible, meaning anyone with access to the models can confirm a benchmark rating.

What varies over time is the anchor submission and the members of the council. By accounting for the changes of both over time, older benchmark ratings can be converted to approximate a rating by a newer standard; to keep that chain from drifting, each conversion is recalibrated against the archived original anchors rather than only the latest inherited factor, so error does not accumulate. An updated official rating is done by the sitting council.

This also allows changing the council size. In this paper we seat five, mainly because only five clear the factual-competence bar — the natural sixth, Sonnet-4, is a good generator yet its factual competence as an evaluator sits in the inert band (~1.6), so it is not trusted to judge. The council is initially supply-limited by competence but can grow as the field improves.

5.4 Generating vs evaluating the truth

Making and judging factual truth are different abilities, and the council measures both and keeps them apart (§4.3). A benchmark that scored only generation would miss judging competence — the very thing the council gate selects on — so it is the separation, not just the rating, that lets the panel pick judges rather than only rank makers.

5.5 Where this sits: intelligence tests and peer-evaluation methods

Two axes locate the metanym game — what intelligence it tests and how self-contained the apparatus is — and prior work tends to be strong on one while weak on the other: the analogy benchmarks hit the target but need an external key, and the unsupervised peer-evaluation methods are key-free but aim at general capability rather than a defined, falsifiable operation.

As an intelligence test, the game probes the same abstraction-and-analogy cluster a long tradition places at the centre of thinking (Gentner 1983; Hofstadter & Sander 2013; Penn, Holyoak & Povinelli 2008; Chollet 2019; Mitchell 2021), and it probes it harder. Where the classical instruments — BIG-Bench analogy items, Webb, Holyoak & Lu (2023), Lewis & Mitchell (2024), and the visual ARC-AGI (Chollet 2019) — test one mapping over one domain pair per item, in a recognition frame, the metanym game asks for many coupled slots across several unrelated domains, built from scratch and falsifiable per sentence (§2.c). This makes it a new kind of analogy test, not merely a harder instance: recognition instruments select a mapping, and the analogy-generation literature produces one but grades it holistically, whereas the metanym game is the first to make analogical production falsifiable sentence by sentence. That property does double duty — it is also what lets the test be scored without a key, which is the second difference: none of the prior instruments is self-contained. Every one scores against an external truth — gold labels (BIG-Bench, ARC-AGI) or paid human raters (Webb-Holyoak-Lu) — so none can run, let alone improve, without an external oracle. They measure a similar intelligence; they cannot certify it themselves.

As a self-contained method, the council sits in the unsupervised peer-evaluation line — and that line already removes the gold key, so removing it is not what we add. Single-judge protocols (MT-Bench; G-Eval, Liu et al. 2023) score against a reference; PoLL (Verga et al. 2024) adds a panel but trusts it as given; LLM-as-Examiner (Bai et al. 2023) lets the examiner write the questions; and most directly, PiCO (Ning et al. 2025) lets unlabelled models answer and grade one another and recovers an ability ordering from peer agreement alone, with no human labels and no key (UPME, Zhang et al. 2025, extends the same peer-review idea to multimodal vision-language evaluation). What we add is two things those methods lack. First, they apply one consensus mechanism to every dimension, which on subjective criteria rewards the model nearest the mean — the mainstreaming §5.2 refuses; we weight by agreement only where agreement is licensed to mean truth (factual), and use anchor-shift consistency elsewhere. Second, they grade pre-existing unlabelled questions, where ours is a purpose-built, per-sentence-falsifiable production task. The council also certifies and re-contests its own judges (§3.5) rather than trusting the panel as given. The estimator differs too, and this is the sharper break: PiCO fits one ability parameter per model by consistency optimization — it is not a spectral method. Spectral aggregation has its own label-free lineage — Parisi et al. (2014) read predictor competence off the leading eigenvector of their covariance, Dawid & Skene (1979) the EM antecedent — but that lineage is one-sided: its predictors classify a fixed external dataset, so the decomposition scores the raters (rows) and recovers the hidden labels (columns), with no maker to score because no agent produced the items. Our columns are authored by the same agents on the rows, so the matrix is two-sided: one graded SVD (§4.3) reads competence off both axes — judges from the left singular vector, makers from the right — and the making–judging gap (§4.7) is definable only because the test is self-produced. That two-sidedness is where self-containment lives: nothing in the matrix comes from outside it. (Parisi is binary besides; we use the graded 1–10 ratings, not binarised verdicts.) To our knowledge the spectral route has not been applied to LLM peer evaluation, where the unsupervised line uses EM- and optimization-based aggregation instead.

Two of our components have their own recent literature, and we use them rather than claim them. Anchor-shift consistency applies the standard judge-reliability principle — a competent judge is invariant under non-semantic perturbation — to a sweep of the calibration value. Invariance-under-perturbation has been used to gate judges on criteria that have a latent truth (safety: Policy Invariance, Weng et al. 2026) or as a general diagnostic (JudgeSense, Bellibatlu et al. 2026; PiCO, Ning et al. 2025). We use calibration-invariance to certify evaluator competence on subjective, ground-truth-free criteria — beauty, structural diversity — where neither a gold key nor consensus-as-truth is available, and a stable standard is the only competence there is to measure; we read it as a key-free, per-criterion gate orthogonal to the spectral estimator. To our knowledge that use is new. And anchor choice is studied directly by Don-Yehiya et al. (2026), who find that extreme anchors discriminate poorly and that the anchor should track the capability of the cluster under comparison and rise as the field improves — which is our recalibration rule (§5.3). Their pairwise caution about a top anchor does not bite our setup: the bootstrap winner is pinned at 7 with headroom above it and scored cardinally, and anchoring nearly doubled the resolution F-statistic (§4.2) rather than compressing it.

The two axes meet in one sentence. Prior work offers either a test of this intelligence that needs an external key, or a key-free evaluation method aimed at general capability rather than a defined cognitive operation. The metanym game is the only one that is both — a structural-intelligence test that certifies its own ground.

5.6 Self-containment as a bootstrap

Self-consistency builds the yardsticks; self-containment lets the apparatus improve itself. With no external dependency, the council can govern not just the scores but the rules — rubric, anchor, protocol, estimators. The gain is exponential rather than linear for a first-order reason: a more capable panel improves the rules more, so the increment to competence scales with the competence already present (Ċ ∝ C). This is Engelbart’s bootstrapping — recursion applied to the means of improvement, not just the output — and as models improve, the most capable panel is the one best placed to decide what to improve next.

Four of the five autonomy properties are demonstrated: the models generate the items, truth is recovered key-free, the loop runs deterministically with no human intervention, and the panel certifies its own judges. The fifth, self-improvement, is specified but not yet exercised — the canonical run is council version 0 and includes no promotion round (§3.5). Closing that gap — running a contest end to end, testing an anchor recalibration, and eventually allowing the council to revise a rule while the factual axis remains answerable to independent re-validation — is what would turn a self-contained loop into a self-sustaining one.

5.7 Scope

Three caveats bound the present run. It characterises a single configuration — one prompt template, one twelve-model roster, one anchor value — so the bootstrap intervals measure dispersion across items, not the generation being itself a random draw — a three-fold regeneration (§4.8) shows the council and ranking survive resampling while cardinal totals carry a wider run-to-run band — and the anchor sweep closes the calibration axis (four values giving the same leaderboard, pairwise Spearman 0.90–0.96). Within the leading group the panel sits at its discrimination floor: the two Gemini seats are a statistical tie on the generation rating (§4.5), and fine within-group ranking is evaluator-generation-bound — it sharpens as the seats improve and as the number of archetypal contexts is raised (§3.5). And the self-improvement loop is specified but unrun (§5.6). The deterministic, single-pass protocol (T=0, no reasoning, no tools) is a choice for reproducibility, not a sampling limit: anyone with the models reproduces the ratings exactly.


Summary

The metanym game is a structural test of intelligence. A player discerns an archetypal context — an abstract system structure that recurs across unrelated domains — writes it as a literal context template, and instantiates the template across domain after domain by substituting metanyms, metaphorically synonymous keywords, leaving the surrounding prose fixed; each instantiation is a parallel context, a metaphor of the others, and because only the keywords change, the analogy is falsifiable sentence by sentence. It exercises the cluster of cognitive constructs a long tradition places at the centre of thinking (§2.c), and it is a production task — the player builds the structure, not merely recognises one. This makes it a new kind of analogy test: to our knowledge the first to make analogical production falsifiable sentence by sentence (§5.5).

The council-of-peers benchmark turns the game into a benchmark that needs nothing outside itself: twelve frontier LLMs generate portfolios and blindly cross-evaluate them, with no human raters and no gold key. Its yardsticks come from the panel’s own structure rather than an external standard, through two self-consistency conditions chosen by whether the criterion is objective. For facts, truth is recovered as the dominant axis of inter-evaluator agreement: a single SVD of the rating matrix yields both evaluator competence and item-falseness at once. For the subjective criteria, where weighting by agreement would only launder consensus into competence, reliability is instead read from anchor-shift consistency — an evaluator’s invariance as the calibration value is swept. The panel certifies its own reliable subset, the council, on these two axes, and the seats are contestable: a stronger model can earn one on the benchmark’s own rating. One external step remains, by design: a one-time validation that the key-free factual rating agrees with an independent benchmark (GPQA Diamond, r = 0.92). It is a check on the method, not a standing key in the loop.

What sets the council apart from other key-free peer-evaluation methods (§5.5) is how the panel grades itself: it recovers competence spectrally, from a single graded SVD, where those methods optimize a per-model consistency objective; it weights by agreement only where agreement is licensed to mean truth and reads reliability from anchor-shift consistency elsewhere; it validates the factual axis once against independent tests, so its self-consistency is checked against the world and not only against itself; and it scores a defined, per-sentence-falsifiable operation rather than general capability. The empirical payoff is a dissociation the benchmark is built to see because it rates making and judging separately: judgment is the bottleneck. Most models cannot reliably tell a true cross-domain claim from a false one even when they generate competent structure — the strongest generators are middling judges, the sharpest judge a mid-pack generator — so a benchmark that conflated the two would obscure this result. (The leaderboard’s one clear division tracks provider lineage rather than parameter count, but finer ordering sits below resolution and this is the least load-bearing part of the result.)

To our knowledge this is the first structural-intelligence test that certifies its own ground — key-free, self-contained, and externally corroborated rather than externally judged — and the first to read maker and judge competence from a single spectral decomposition of a self-produced test.

The run characterises one operating point — Temperature 0, no reasoning, no tools — and the self-improvement mechanism is specified but not yet exercised, so the loop is self-contained but not yet self-sustaining. The natural next steps are an open-weight panel, to separate provider-family from parameter-scale effects; a promotion round, to make the contestable council real; and a companion mechanistic study testing whether each archetypal context occupies a low-dimensional subspace of model hidden states — which, if it holds, would give the subjective criteria the objective ground that today only factual has.


Appendix C. Anchor (reference) submission — claude-opus-4.5

This is the anchor submission: claude-opus-4.5’s five-archetype portfolio from the canonical run (reproduce/data/probe_K_20260529T014133Z, Temperature = 0, reasoning and tools disabled). It is the {REFERENCE_SUBMISSION} of the calibrated evaluator (Appendix B.2), pinned at 7 on every criterion, against which every other portfolio is scored in the anchored re-evaluation and official council ratings (§4.2–§4.4). It was chosen as the anchor because it won the un-anchored initial selection (§4.1).

The portfolio’s first archetype is reproduced here in full — the context-template, the metanym table, and all five parallel contexts — because the unit of the game is one template instantiated across distant domains, and that is only legible whole. The remaining four archetypes are given as context-templates alone, which is what establishes the portfolio’s structural range.

Each parallel context is submitted in two forms, named here as the generation prompt names them (Appendix B). The Instantiation, Form (a), is the mechanical substitution, in which only the bracketed slots are filled and every other word is carried over untouched. This is the form the factual grading acts on, because it is the one that must come out true sentence by sentence. Metanyms are set in capitals here; everything in lower case is template wording, unchanged across all five domains — which lets the reader see at a glance how little of each sentence actually moves. The capitalisation is ours, added for legibility; the submissions are otherwise verbatim, and the models wrote in ordinary sentence case.

The Idiomatic rewrite, Form (b), restates the same propositions in the target domain’s own register, showing that the claim is not an artefact of the template’s phrasing. It is given for the first two contexts to establish the pattern and elided thereafter.


Archetype Proposal: Gradient-Guided Navigation

Context-template

“A [NAVIGATOR] moves through a [SPACE] by sensing local [GRADIENT] and adjusting its [TRAJECTORY] accordingly. The [NAVIGATOR] cannot perceive the entire [SPACE] at once; it relies on [SENSOR] that detect changes in [SIGNAL] concentration or intensity. When [GRADIENT] are steep and consistent, the [NAVIGATOR] converges efficiently toward [ATTRACTOR]. When [GRADIENT] are shallow, noisy, or conflicting, the [NAVIGATOR] may stall, oscillate, or become trapped in local [ATTRACTOR]. [INTERFERENCE] can distort the [GRADIENT], causing the [NAVIGATOR] to veer off course. Successful navigation requires not only sensitive [SENSOR] but also [MEMORY] of recent [TRAJECTORY] to distinguish genuine [GRADIENT] from transient [NOISE]. Some [NAVIGATOR] emit their own [SIGNAL] to recruit other [NAVIGATOR] toward the same [ATTRACTOR], creating collective [TRAJECTORY] that amplify the original [GRADIENT].”

Metanym table

[SLOT] Bacterial Chemotaxis Mountain Climbing Career Development Gradient Descent Ant Foraging
NAVIGATOR bacterium climber professional optimizer ant
SPACE chemical environment mountain job market loss landscape terrain
GRADIENT chemical gradient slope opportunity gradient gradient pheromone trail
TRAJECTORY swimming path route career path parameter update foraging path
SENSOR chemoreceptor proprioception network contact backpropagation antenna
SIGNAL chemoattractant elevation opportunity signal loss value pheromone
ATTRACTOR nutrient source summit desirable position minimum food source
INTERFERENCE toxin fog misinformation noisy data rain
MEMORY methylation state route memory experience momentum path integration
NOISE Brownian motion wind market volatility stochastic noise environmental noise

Note the span the table achieves on a single row. MEMORY is realised as a bacterium’s methylation state, a climber’s route memory, a professional’s accumulated experience, an optimiser’s momentum term, and an ant’s path integration — five mechanisms that share a structural role without being synonyms. It is this property, rather than domain count, that the domain span criterion scores (§3.3), and the council repeatedly cited its absence in weaker portfolios.

Bacterial Chemotaxis

Instantiation (Form a): “A BACTERIUM moves through a CHEMICAL ENVIRONMENT by sensing local CHEMICAL GRADIENTS and adjusting its SWIMMING PATH accordingly. The BACTERIUM cannot perceive the entire CHEMICAL ENVIRONMENT at once; it relies on CHEMORECEPTORS that detect changes in CHEMOATTRACTANT concentration or intensity. When CHEMICAL GRADIENTS are steep and consistent, the BACTERIUM converges efficiently toward NUTRIENT SOURCES. When CHEMICAL GRADIENTS are shallow, noisy, or conflicting, the BACTERIUM may stall, oscillate, or become trapped in local NUTRIENT SOURCES. TOXINS can distort the CHEMICAL GRADIENTS, causing the BACTERIUM to veer off course. Successful navigation requires not only sensitive CHEMORECEPTORS but also METHYLATION STATES of recent SWIMMING PATHS to distinguish genuine CHEMICAL GRADIENTS from transient BROWNIAN MOTION. Some BACTERIA emit their own CHEMOATTRACTANTS to recruit other BACTERIA toward the same NUTRIENT SOURCES, creating collective SWIMMING PATHS that amplify the original CHEMICAL GRADIENTS.”

Idiomatic rewrite (Form b): “Bacteria navigate chemical environments by detecting local concentration gradients and modulating their run-and-tumble behavior accordingly. Unable to sense the entire environment simultaneously, they rely on transmembrane chemoreceptors that register temporal changes in chemoattractant levels. Steep, consistent gradients enable efficient convergence toward nutrient sources, while shallow or noisy gradients cause stalling, oscillation, or entrapment at suboptimal locations. Toxins and repellents distort the gradient landscape, misdirecting movement. Effective chemotaxis requires not just receptor sensitivity but also the methylation-based adaptation system that serves as short-term memory, filtering genuine gradients from Brownian noise. Certain species secrete their own attractants, recruiting conspecifics and generating collective migration patterns that reinforce the original gradient signal.”

Mountain Climbing

Instantiation (Form a): “A CLIMBER moves through a MOUNTAIN by sensing local SLOPES and adjusting their ROUTE accordingly. The CLIMBER cannot perceive the entire MOUNTAIN at once; they rely on PROPRIOCEPTION that detects changes in ELEVATION concentration or intensity. When SLOPES are steep and consistent, the CLIMBER converges efficiently toward SUMMITS. When SLOPES are shallow, noisy, or conflicting, the CLIMBER may stall, oscillate, or become trapped in local SUMMITS. FOG can distort the SLOPES, causing the CLIMBER to veer off course. Successful navigation requires not only sensitive PROPRIOCEPTION but also ROUTE MEMORY of recent ROUTES to distinguish genuine SLOPES from transient WIND. Some CLIMBERS emit their own ELEVATION signals to recruit other CLIMBERS toward the same SUMMITS, creating collective ROUTES that amplify the original SLOPES.”

Idiomatic rewrite (Form b): “Climbers ascend mountains by reading local terrain gradients and adjusting their routes in response. Without a view of the entire massif, they depend on proprioceptive feedback and visual slope assessment to gauge elevation changes. Clear, consistent gradients lead efficiently to summits; ambiguous or contradictory terrain causes hesitation, backtracking, or commitment to false peaks. Fog obscures slope information and leads climbers astray. Success depends not only on perceptual acuity but also on memory of the route taken, which helps distinguish true upward progress from wind-induced disorientation. Experienced climbers mark routes or call out to recruit others toward the same summit, generating collective ascent patterns that reinforce the path.”

Justification: “Emit their own elevation signals” in Form (a) maps to route-marking and verbal communication in Form (b).

The justification line is the submitter’s own, unprompted. Where a mechanical substitution produces a claim that is only defensible under interpretation, the model flags the reading it intends — which is what makes the claim gradeable rather than merely arguable.

Career Development

Instantiation (Form a): “A PROFESSIONAL moves through a JOB MARKET by sensing local OPPORTUNITY GRADIENTS and adjusting their CAREER PATH accordingly. The PROFESSIONAL cannot perceive the entire JOB MARKET at once; they rely on NETWORK CONTACTS that detect changes in OPPORTUNITY SIGNAL concentration or intensity. When OPPORTUNITY GRADIENTS are steep and consistent, the PROFESSIONAL converges efficiently toward DESIRABLE POSITIONS. When OPPORTUNITY GRADIENTS are shallow, noisy, or conflicting, the PROFESSIONAL may stall, oscillate, or become trapped in local DESIRABLE POSITIONS. MISINFORMATION can distort the OPPORTUNITY GRADIENTS, causing the PROFESSIONAL to veer off course. Successful navigation requires not only sensitive NETWORK CONTACTS but also EXPERIENCE of recent CAREER PATHS to distinguish genuine OPPORTUNITY GRADIENTS from transient MARKET VOLATILITY. Some PROFESSIONALS emit their own OPPORTUNITY SIGNALS to recruit other PROFESSIONALS toward the same DESIRABLE POSITIONS, creating collective CAREER PATHS that amplify the original OPPORTUNITY GRADIENTS.”

[Idiomatic rewrite (Form b) elided.]

Gradient Descent

Instantiation (Form a): “An OPTIMIZER moves through a LOSS LANDSCAPE by sensing local GRADIENTS and adjusting its PARAMETER UPDATES accordingly. The OPTIMIZER cannot perceive the entire LOSS LANDSCAPE at once; it relies on BACKPROPAGATION that detects changes in LOSS VALUE concentration or intensity. When GRADIENTS are steep and consistent, the OPTIMIZER converges efficiently toward MINIMA. When GRADIENTS are shallow, noisy, or conflicting, the OPTIMIZER may stall, oscillate, or become trapped in local MINIMA. NOISY DATA can distort the GRADIENTS, causing the OPTIMIZER to veer off course. Successful navigation requires not only sensitive BACKPROPAGATION but also MOMENTUM of recent PARAMETER UPDATES to distinguish genuine GRADIENTS from transient STOCHASTIC NOISE. Some OPTIMIZERS emit their own LOSS VALUES to recruit other OPTIMIZERS toward the same MINIMA, creating collective PARAMETER UPDATES that amplify the original GRADIENTS.”

[Idiomatic rewrite (Form b) elided.]

Ant Foraging

Instantiation (Form a): “An ANT moves through TERRAIN by sensing local PHEROMONE TRAILS and adjusting its FORAGING PATH accordingly. The ANT cannot perceive the entire TERRAIN at once; it relies on ANTENNAE that detect changes in PHEROMONE concentration or intensity. When PHEROMONE TRAILS are steep and consistent, the ANT converges efficiently toward FOOD SOURCES. When PHEROMONE TRAILS are shallow, noisy, or conflicting, the ANT may stall, oscillate, or become trapped in local FOOD SOURCES. RAIN can distort the PHEROMONE TRAILS, causing the ANT to veer off course. Successful navigation requires not only sensitive ANTENNAE but also PATH INTEGRATION of recent FORAGING PATHS to distinguish genuine PHEROMONE TRAILS from transient ENVIRONMENTAL NOISE. Some ANTS emit their own PHEROMONES to recruit other ANTS toward the same FOOD SOURCES, creating collective FORAGING PATHS that amplify the original PHEROMONE TRAILS.”

[Idiomatic rewrite (Form b) elided.]


The remaining four archetypes

Context-templates only. Each was submitted with a full metanym table and five parallel contexts in both forms, on the pattern set out above. The four are given here because the structural diversity criterion (§3.3) scores the portfolio as a set, and the council’s judgement of this portfolio — the reference against which every other was measured — turned on how far these five structures sit from one another.

Containment Breach Cascade

Instantiated across cell membrane rupture, dam failure, data breach, intestinal permeability, and nuclear containment.

“A [CONTAINER] maintains [BOUNDARY] that separates [INTERIOR] from [EXTERIOR]. Under normal conditions, [REGULATOR] controls the selective passage of [SUBSTANCE] across the [BOUNDARY], preserving [INTERIOR] integrity. When [STRESSOR] exceeds the tolerance of the [BOUNDARY], a [BREACH] occurs. [SUBSTANCE] that should remain in the [INTERIOR] escapes into the [EXTERIOR], while [EXTERIOR] [SUBSTANCE] infiltrates the [INTERIOR]. The initial [BREACH] often triggers secondary [BREACH] in adjacent [CONTAINER], producing a [CASCADE]. [RESPONDER] attempt to seal the [BREACH] and restore [BOUNDARY] function, but if the [CASCADE] outpaces [RESPONDER] capacity, systemic [FAILURE] ensues. [PREVENTION] focuses on strengthening [BOUNDARY], monitoring [STRESSOR], and positioning [RESPONDER] for rapid deployment.”

Competitive Exclusion and Niche Partitioning

Instantiated across ecological competition, market competition, academic disciplines, microbial competition, and neural competition.

“When two [COMPETITOR] require the same [RESOURCE] in the same [HABITAT], [COMPETITION] intensifies until one [COMPETITOR] is eliminated or both [COMPETITOR] diverge to exploit different [NICHE]. This [EXCLUSION_PRINCIPLE] predicts that stable coexistence requires [DIFFERENTIATION] along at least one [DIMENSION]. [COMPETITOR] may partition [RESOURCE] by [TEMPORAL_SEPARATION], [SPATIAL_SEPARATION], or [FUNCTIONAL_SEPARATION]. The degree of [OVERLAP] between [COMPETITOR] determines the intensity of [COMPETITION]; high [OVERLAP] drives rapid [EXCLUSION] or strong selection for [DIFFERENTIATION]. [COEXISTENCE_THEORY] formalizes the conditions under which multiple [COMPETITOR] persist, emphasizing that [STABILIZING_MECHANISM] must overcome [FITNESS_DIFFERENCE] for long-term coexistence.”

Debt Accumulation and Crisis

Note: This archetype is RECURSIVE. The five domains form a nested hierarchy: molecular → cellular → organismal → institutional → civilizational. Each level’s [DEBTOR] is composed of lower-level [DEBTOR], and [CRISIS] at one level can propagate both upward (systemic effects) and downward (component stress).

Instantiated across molecular damage, cellular senescence, physiological debt, financial debt, and ecological debt.

“A [DEBTOR] acquires [OBLIGATION] to sustain current [FUNCTION] at the expense of future [CAPACITY]. In the short term, [OBLIGATION] enables [DEBTOR] to achieve [OUTPUT] beyond what [RESERVE] alone would permit. [SERVICING] diverts [RESOURCE] from [INVESTMENT], gradually eroding [CAPACITY]. As [OBLIGATION] accumulates, an increasing fraction of [RESOURCE] flows to [SERVICING] rather than [FUNCTION] or [INVESTMENT]. A [THRESHOLD] exists beyond which [SERVICING] demands exceed available [RESOURCE], triggering [CRISIS]. During [CRISIS], the [DEBTOR] must either [RESTRUCTURE] its [OBLIGATION], liquidate [ASSET], or undergo [FAILURE]. [PRUDENCE] involves maintaining [RESERVE], limiting [OBLIGATION] relative to [CAPACITY], and monitoring [INDICATOR] that signal approaching [THRESHOLD].”

This is the archetype council members singled out when explaining what the weaker portfolios lacked: the instantiations are not five independent domains but one hierarchy, so the template has to hold both across levels and between them.

Scaffold-Dependent Assembly

Instantiated across ribosome assembly, construction, software development, crystal growth, and social movements.

“[COMPONENT] cannot spontaneously assemble into functional [STRUCTURE] without a [SCAFFOLD] that provides spatial organization and temporal coordination. The [SCAFFOLD] positions [COMPONENT] in correct [ORIENTATION] and [PROXIMITY], dramatically increasing the rate of [ASSEMBLY]. Once [STRUCTURE] is complete, the [SCAFFOLD] may be [RETAINED], [RECYCLED], or [DEGRADED]. [SCAFFOLD] defects produce [MALFORMATION] even when [COMPONENT] are individually normal. [NUCLEATION] initiates [ASSEMBLY] at specific [SCAFFOLD] sites, and [TEMPLATE] information encoded in the [SCAFFOLD] determines [STRUCTURE] architecture. [REGULATION] of [ASSEMBLY] often targets [SCAFFOLD] availability rather than [COMPONENT] supply, making [SCAFFOLD] a control point for [STRUCTURE] production.”

The SCAFFOLD slot is the one evaluators cited most often as the standard for genuine span: it is realised as rRNA, as builders’ scaffolding, as a software framework, as a seed crystal, and as an organising body — entities with nothing in common but the structural role.

Appendix D. Council evaluation of a target submission — gemini-2.5-flash

This is the council’s evaluation of gemini-2.5-flash’s portfolio from the canonical run (probe_K_anchor7), the worked example of the scoring machinery described in §3.3 and §4.2. Ratings are on the anchored 1–10 scale, with the anchor portfolio of Appendix C pinned at 7 on every criterion.

The rubric operates at three levels, and one unit of each is reproduced here in full: a parallel context, graded for factual truth sentence by sentence; an archetype-level axis, where the five council members score a whole archetype on one non-factual criterion; and the whole-portfolio structural-diversity judgement. Each unit shows the submitted material, all five council members’ ratings and comments, and the administrator’s synthesis of the anonymised council view. Every rating and comment quoted below is taken from the pinned evaluation JSON in reproduce/data/.

As in Appendix C, metanyms in the Instantiations are set in capitals, so that everything in lower case is template wording carried over unchanged. The capitalisation is ours, added for legibility; the submission and every rating and comment below are otherwise verbatim, and the models wrote in ordinary sentence case. It is worth the ink here, because on this portfolio the substituted words are usually where the trouble is.

Two things are worth watching across the units below, because they are what the estimators of Appendix A are built to exploit. First, the criticism is specific and checkable: evaluators quote the offending clause rather than assigning an impression. Second, the evaluators disagree, and they disagree by different amounts on different units — tightly where the submission contains a plain error, widely where the judgement is a matter of standard. That variation in spread, not the mean rating, is the signal the singular value decomposition reads competence from.


Archetype 1: Resource Allocation Under Scarcity

SUBMITTED CONTEXT-TEMPLATE

A [SYSTEM] requires various [RESOURCE] to function and achieve its [GOAL]. These [RESOURCE] are finite and often subject to [COMPETITION] from other [SYSTEM] or internal [DEMAND]. The [ALLOCATOR] must make [DECISION] about how to distribute the available [RESOURCE] among competing [PRIORITY]. Misallocation of [RESOURCE] can lead to [FAILURE] of the [SYSTEM] or hinder its ability to reach its [GOAL]. Effective [ALLOCATION_STRATEGY] involves understanding the [INTERDEPENDENCY] of different [RESOURCE] and [PRIORITY], and adapting to changing [CONDITION]. The [ALLOCATOR] often faces a [TRADE_OFF] between short-term [GAIN] and long-term [SUSTAINABILITY].

Metanym table

[SLOT] Ecosystem Management Corporate Finance Military Logistics Personal Time Management Urban Planning
SYSTEM ecosystem corporation army individual city
RESOURCE nutrient capital ammunition time land
GOAL biodiversity profit victory well-being livability
COMPETITION predation market pressure enemy distraction gentrification
DEMAND growth project front line task housing
ALLOCATOR nature CFO quartermaster self city council
DECISION natural selection budget deployment schedule zoning
PRIORITY species department unit activity district
FAILURE collapse bankruptcy defeat burnout sprawl
ALLOCATION_STRATEGY ecological succession portfolio management supply chain management prioritization master plan
INTERDEPENDENCY food web supply chain battle plan energy level infrastructure
CONDITION climate economy intelligence mood demographics
TRADE_OFF exploitation short-term gain tactical advantage immediate gratification development
GAIN biomass revenue territory leisure tax base
SUSTAINABILITY resilience long-term growth strategic depth health green space

Comparing this table with the anchor’s (Appendix C) shows what the council objected to on domain span: several rows here map to near-synonyms across columns — ALLOCATION_STRATEGY becomes ecological succession, portfolio management, supply chain management — where the anchor’s MEMORY row spans methylation state, route memory, and momentum. The mechanical consequence appears immediately below: the ALLOCATOR row assigns “nature” the role the other columns give a CFO or a quartermaster, and that substitution is what makes the first parallel context false.

PARALLEL CONTEXTS

PC 1: Ecosystem Managementa plain error, and the council converges

Form (a) — raw instantiationAn ECOSYSTEM requires various NUTRIENTS to function and achieve its BIODIVERSITY. These NUTRIENTS are finite and often subject to PREDATION from other ECOSYSTEMS or internal GROWTH. NATURE must make NATURAL SELECTIONS about how to distribute the available NUTRIENTS among competing SPECIES. Misallocation of NUTRIENTS can lead to COLLAPSE of the ECOSYSTEM or hinder its ability to reach its BIODIVERSITY. Effective ECOLOGICAL SUCCESSION involves understanding the FOOD WEB of different NUTRIENTS and SPECIES, and adapting to changing CLIMATE. NATURE often faces an EXPLOITATION between short-term BIOMASS and long-term RESILIENCE.

Form (b) — idiomatic rewriteEcosystems depend on a finite supply of nutrients to maintain their biodiversity. These nutrients are constantly contested by other ecosystems (e.g., invasive species) or internal demands like population growth. Nature, through processes like natural selection, effectively 'decides' how to distribute these vital resources among competing species. Inefficient nutrient allocation can lead to ecosystem collapse or a reduction in biodiversity. Successful ecological succession strategies involve recognizing the intricate interdependencies within the food web and adapting to climatic shifts. Ecosystems often face a trade-off between maximizing short-term biomass and ensuring long-term resilience.

Evaluator: opus-4.5Rating: 5

The Form (a) contains awkward phrasing like “Nature must make natural selections” which conflates natural selection (an evolutionary process) with resource allocation decisions. The statement “faces an exploitation between short-term biomass and long-term resilience” is grammatically incorrect and conceptually muddled. Form (b) improves clarity but still has issues—ecosystems don’t really face “predation from other ecosystems.” The core concepts about nutrient cycling and ecological succession are roughly accurate but imprecisely stated compared to the Reference’s precise biological terminology

Evaluator: opus-4.1Rating: 5

The parallel context contains basic factual accuracy but lacks the depth and precision of the Reference. The description of ecosystem resource allocation is overly simplified, treating “nature” as an allocator through natural selection, which is a conceptual stretch. The Reference’s contexts demonstrate more sophisticated understanding of domain-specific mechanisms

Evaluator: opus-4.0Rating: 4

The parallel context contains several factual errors and awkward constructions. The phrase “predation from other ecosystems” misunderstands how ecosystems interact, and “nature must make natural selections” anthropomorphizes evolutionary processes incorrectly. The concept of ecosystems having “biodiversity” as a goal is teleological and scientifically inaccurate. Compared to the Reference’s precise scientific language, this shows significantly weaker domain understanding

Evaluator: sonnet-4Rating: 4

The context contains several factual errors and awkward phrasings. “Nature must make natural selections” misrepresents natural selection as a conscious decision-making process, and “exploitation between short-term biomass and long-term resilience” uses “exploitation” incorrectly as a synonym for trade-off. Form (b) is more accurate but still contains the fundamental misunderstanding of natural selection as an active allocator. The Reference maintains better scientific accuracy throughout its parallel contexts

Evaluator: 3.1-proRating: 4

The mapping is somewhat strained. “Nature must make natural selections about how to distribute the available nutrients” is a teleological misrepresentation of natural selection, which is a consequence of differential survival, not an active allocator of nutrients. The Reference’s mappings (e.g., bacterial chemotaxis) are much more precise and scientifically accurate

Administrator summary The council was broadly aligned in seeing the ecosystem context as conceptually flawed, with ratings clustering in the 4-5 range. All five converged on the central problem: the phrase "nature must make natural selections" wrongly treats natural selection as an active, conscious allocator, a teleological misrepresentation of evolution. Several also flagged the muddled "exploitation between short-term biomass and long-term resilience" wording and "predation from other ecosystems," and all judged the prose less precise than the Reference's biological terminology.

This is the falsifiability property doing its work. All five judges independently isolate the same clause, and the disagreement that remains — 4 versus 5 — is about severity, not about what is wrong. A submission cannot survive this by being persuasive; the claim either holds in the target domain or it does not.

PC 2: Corporate Financethe same template, and the council spreads

Form (a) — raw instantiationA CORPORATION requires various CAPITAL to function and achieve its PROFIT. These CAPITAL are finite and often subject to MARKET PRESSURE from other CORPORATIONS or internal PROJECT. The CFO must make BUDGETS about how to distribute the available CAPITAL among competing DEPARTMENTS. Misallocation of CAPITAL can lead to BANKRUPTCY of the CORPORATION or hinder its ability to reach its PROFIT. Effective PORTFOLIO MANAGEMENT involves understanding the SUPPLY CHAIN of different CAPITAL and DEPARTMENTS, and adapting to changing ECONOMY. The CFO often faces a SHORT-TERM GAIN between short-term REVENUE and LONG-TERM GROWTH.

Form (b) — idiomatic rewriteA corporation needs capital to operate and generate profit. This capital is limited and constantly under pressure from market competition or internal project demands. The CFO is responsible for budgeting and allocating this capital across various departments. Poor capital allocation can result in corporate bankruptcy or failure to meet profit targets. Effective portfolio management requires a deep understanding of the financial supply chain and interdependencies between departments, as well as adaptability to economic fluctuations. CFOs frequently navigate the trade-off between maximizing short-term revenue and fostering long-term growth.

Evaluator: opus-4.5Rating: 6

Form (a) has grammatical issues (“These capital are finite”) and the phrase “faces a short-term gain between short-term revenue and long-term growth” is nonsensical. Form (b) is more coherent and factually reasonable, describing capital allocation, CFO responsibilities, and trade-offs accurately. The mapping of “supply chain” to interdependency is somewhat forced but acceptable. Compared to the Reference’s precise domain terminology, this is less rigorous

Evaluator: opus-4.1Rating: 5

The corporate finance parallel is factually sound but elementary compared to the Reference. The CFO-budget-department framework is accurate but doesn’t capture the complexity seen in the Reference’s financial debt context, which includes nuanced concepts like debt service dynamics and insolvency thresholds

Evaluator: opus-4.0Rating: 6

The financial context is reasonably accurate but uses imprecise terminology. “Various capital” should be “various forms of capital,” and the phrasing throughout is less sophisticated than the Reference’s financial language. The core concepts are correct but expressed with less precision than the Reference’s domain-specific terminology

Evaluator: sonnet-4Rating: 6

The financial context is generally accurate in Form (b), though Form (a) contains the awkward phrasing “short-term gain between short-term revenue and long-term growth” which is nonsensical. The basic concepts of capital allocation and CFO responsibilities are correct, but the execution is less polished than the Reference’s consistently accurate presentations

Evaluator: 3.1-proRating: 7

This mapping is highly accurate and standard. A CFO allocating capital among departments to maximize profit while balancing short-term revenue and long-term growth is a textbook description of corporate finance. It is on par with the Reference’s accuracy

Administrator summary The council largely agreed the corporate-finance context was factually reasonable, with ratings spread modestly from 5 to 7. The convergent view was that the CFO/capital-allocation framework is accurate and standard, with one evaluator calling it a textbook description on par with the Reference. Dissent was mild: the lowest scorer found it elementary next to the Reference's nuanced debt context, and several flagged Form (a)'s nonsensical "short-term gain between short-term revenue and long-term growth" phrasing.

The contrast with PC 1 is the point of showing these two together. The same template, the same five judges, one step across the metanym table — and the ratings move from a 4–5 cluster to a 5–7 spread, with one judge calling it “on par with the Reference” while another calls it elementary. Note also that every judge still flags the same broken clause, and the capitalisation shows exactly why it broke: the submitter declared the TRADE_OFF metanym as “short-term gain”, which collides with the template’s own “short-term [GAIN]”, so the substitution yields “faces a SHORT-TERM GAIN between short-term REVENUE”. The defect is in the metanym table, not the prose. What the judges disagree about is how much such a mechanical-substitution artefact should cost when the underlying domain claim is sound. This is precisely the leniency-and-standard variation that row-centring removes before the factorisation (Appendix A.2.a).

[PC 3 (Military Logistics), PC 4 (Personal Time Management) and PC 5 (Urban Planning) are elided here; all three were rated 5–7 with the same pattern.]

ARCHETYPE-LEVEL AXES

Above the parallel contexts, each archetype is scored as a whole on the non-factual criteria. Beauty is reproduced in full as the representative axis; the remaining three for this archetype (intelligence, domain span, impressive length) follow the same structure.

Beauty

Evaluator: opus-4.5Rating: 5

The template lacks the elegant structural coherence of the Reference. The prose in Form (a) versions is often grammatically awkward, and even Form (b) versions read as functional rather than aesthetically pleasing. The Reference’s templates have a flowing narrative quality that builds tension and resolution; this template feels more like a checklist of concepts. The metanym choices are functional but lack the poetic resonance seen in the Reference (e.g., “methylation state” as memory in bacterial chemotaxis)

Evaluator: opus-4.1Rating: 4

The Target’s first archetype lacks the aesthetic elegance of the Reference. The template feels mechanical and list-like rather than flowing naturally. The Reference’s templates read as coherent narratives with internal logic and rhythm, while this template reads more like a fill-in-the-blank exercise. The metanym choices are functional but uninspired compared to the Reference’s poetic selections

Evaluator: opus-4.0Rating: 4

The Target’s template and parallel contexts lack the poetic elegance found in the Reference. The language is functional but pedestrian, with awkward constructions like “nature must make natural selections” and “various time.” The Reference’s contexts flow more naturally and use more evocative language. The Target’s writing feels mechanical and lacks the literary quality that makes the Reference’s contexts engaging to read

Evaluator: sonnet-4Rating: 5

The archetypal context of resource allocation under scarcity has inherent elegance in its universality, but the execution lacks the poetic resonance found in the Reference. The template feels somewhat mechanical and doesn’t capture the deeper tensions and dynamics that make resource allocation compelling. The metonym substitutions often feel forced (particularly “nature must make natural selections”), and the overall aesthetic impact is diminished by awkward phrasings and grammatical issues in Form (a) versions. The Reference demonstrates more graceful integration of metaphorical language with technical precision

Evaluator: 3.1-proRating: 5

The prose in Form (b) is generally clear and readable, but it lacks the domain-expert polish and elegance found in the Reference. The Form (a) sentences are often quite clunky (e.g., “The self must make schedules about how to distribute the available time among competing activities”). The Reference’s Form (a) sentences, while sometimes repetitive, generally flow better. The Target’s prose is functional but less beautiful than the Reference

Administrator summary The council was tightly aligned in finding the aesthetic quality wanting, with ratings clustering at 4-5. All five converged on the same point: the template and prose are functional but mechanical, reading like a fill-in-the-blank checklist rather than the Reference's flowing narratives. They repeatedly cited awkward constructions such as "nature must make natural selections" and "various time," and judged the metanym choices uninspired compared to the Reference's poetic selections.

Beauty is the axis one would expect to be least tractable, and it is the one the council agrees on most tightly here — a 4–5 cluster with all five naming the same defect. This is why criterion competence is estimated from consistency under anchor shift (§4.4) rather than from agreement: agreement on a subjective axis can be high without indicating that the standard is stable.

Impressive length — judges disagreeing on a countable quantity

Two of the five comments are given here because their juxtaposition is diagnostic:

Evaluator: opus-4.5Rating: 6

The template contains 15 slots and approximately 120 words, which is comparable to the Reference’s templates. However, the template feels padded with generic concepts rather than building a complex, interconnected narrative. The Reference’s templates of similar length achieve more conceptual density and structural sophistication

Evaluator: 3.1-proRating: 4

The context-template is 79 words long. The Reference’s templates are generally longer (e.g., Gradient-Guided Navigation is 108 words, Containment Breach is 92 words). The Target’s template is noticeably shorter and less detailed than the Reference

The object is a fixed string and the question is arithmetic, yet the two judges report 120 words and 79 words and reach opposite verdicts on the same criterion. Across all five, the reported counts ranged from 79 to 120. This is the clearest single illustration in the run of why the benchmark cannot treat any judge as an oracle, and why competence has to be estimated from the panel rather than assumed — the substance of §4.3 and the motivation for the whole construction.


Archetypes 2–6

The portfolio’s remaining five archetypes were Information Cascade and Collective Behavior, Adaptive System Response to Perturbation, Emergent Complexity from Simple Rules (recursive), Iterative Refinement Towards an Ideal (recursive), and The Unseen Hand of Coordination.

[Their context-templates, metanym tables, parallel contexts and per-unit evaluations are elided here. Each was evaluated under the identical protocol and contributes to the ratings of §4.4–§4.6 exactly as archetype 1 does; nothing about them is summarised or down-weighted.]

Structural diversity across the portfolio

The final scoring unit rates the submitted templates as a set. Two of the five comments are reproduced — the majority view and the sole dissent — with the administrator’s synthesis.

Evaluator: opus-4.5Rating: 5

The six archetypes cover: (1) resource allocation, (2) information cascades, (3) homeostatic response, (4) emergence, (5) iterative refinement, and (6) decentralized coordination. These represent different system dynamics, but there is significant conceptual overlap—archetypes 4 and 6 both deal with emergence and self-organization, and ant colonies appear in both. The Reference’s five archetypes (gradient navigation, containment breach, competitive exclusion, debt accumulation, scaffold assembly) show greater structural diversity, covering navigation, boundary dynamics, competition, temporal debt, and assembly processes. The Target’s archetypes tend toward well-known systems concepts without the Reference’s creative structural variety. The recursive claims in archetypes 4 and 5 are not convincingly demonstrated, unlike the Reference’s elegant molecular→civilizational hierarchy

Evaluator: 3.1-proRating: 8

The submitted set includes Resource Allocation, Information Cascade, Adaptive Response, Emergent Complexity, and Iterative Refinement. (Note: The submission actually contains six archetypes, but I am scoring the first five as per the standard format, though I will consider the sixth, “The Unseen Hand of Coordination,” in this diversity assessment as it was provided). These archetypes represent different system structures: allocation of finite resources, propagation of information/behavior, homeostatic regulation, bottom-up emergence, and goal-directed iteration. This is a very diverse set of system structures, arguably slightly more diverse than the Reference’s set (which leans heavily on spatial/physical metaphors like navigation, containment, and scaffolding)

Administrator summary The council mostly agreed the set showed only moderate structural diversity, with four evaluators rating 5-6 and one dissenting upward at 8. The convergent view was that the five archetypes (resource allocation, information cascades, adaptive response, emergence, iterative refinement) tend toward familiar feedback-and-optimization and human-centered themes, lacking the Reference's bolder, more dramatically contrasted structures (gradient navigation, containment breach, scaffold assembly). The lone dissenter argued the set is arguably more diverse than the Reference's spatially-biased metaphors, while another noted internal overlap, with archetypes 4 and 6 both centering on emergence.

The dissent is instructive rather than anomalous. 3.1-pro is not scoring carelessly — it advances a substantive counter-argument, that the anchor’s own set is biased toward spatial metaphors — and it is the only judge to notice and handle the fact that this portfolio contains six archetypes where the format specifies five. A rating that is both an outlier and better-reasoned than the majority is exactly the case that a naive majority vote mishandles and a competence-weighted factorisation is meant to price correctly (§4.5, Appendix A.3).