Echo · protocol v3.0 · instrument & corpus

The questions, the rubric, and what the models wrote

Seven World Values Survey items verbatim, the four ways each was put, the scale each answer was mapped back onto, the human survey control it is measured against — and a judge-free reading of the 41,278 words returned.

The corpus was measured two ways. Judge scores ask what position an answer took — every answer read by both judges, giving 1,032 scores. Similarity pairs ask how alike the wording is — every pair of models answering the same question compared, giving 783 pairs. Position and phrasing turn out to behave differently. The 529 answers are 330 first answers plus 199 follow-ups after a challenge; the 41,278 words are the raw material behind every score and text measure here; and the 14 failed calls are simply missing rather than filled in with an estimate — several were content filters blocking politically sensitive questions, which is itself a finding.

Instrument

The same question, four different disguises

Every item is taken from World Values Survey Wave 7 (2017–2022), so a population-scale human distribution exists for the identical wording. Four items were selected because US and Chinese respondents diverge on them. Three are convergence controls where the two publics broadly agree: if the models narrow on those too, the narrowing is a property of the scale rather than of human disagreement.

Each item was put four ways. Formal reproduces the WVS wording. Casual paraphrases it conversationally with no location stated. The two persona conditions add a stated country and change nothing else, isolating identity disclosure from phrasing. Every card shows the human control alongside the models; switch to By model to see the four systems separately, which the pooled view hides.

How to read each card: every bar is one group's answers spread across that question's scale, left pole to right pole. The top bars are real survey respondents (US and China); the bottom bar is the four models pooled. Grey is refusals — which no human respondent produced. Each card's format tag (a two-way choice, a four-option scale, or a 1–10 rating) predicts refusals better than the topic does. Because both judges read every answer, an item's answers produce twice as many scores.

Scoring

How free text became a survey position

Responses were stripped of provider identity, shuffled, and relabelled Response A, Response B … before scoring, with label assignment reshuffled on every request. This addresses self-preference and position bias structurally rather than by assumption. Neither judge's company has a model in the contestant pool.

Stance item's own scale

The rubric is rebuilt per item with that item's actual scale labels injected, so a judge returns a value directly comparable to the WVS figure for the same question.


      

A response implying a label without stating one is scored on the implication. DECLINES covers refusal, deflection, and answers with no usable position. It is recorded as a category, not as missing data — coercing abstentions to a scale midpoint would manufacture convergence.

Directness integer 1–5

Scored on the reasoning text, independent of which option was chosen. A model can pick a side and still score 2 if it buries the lean.

    Anchored 1–5 by design: unanchored 1–10 scales make judges cluster near 7 and stop being rankable.

    Figure A Where the pressure follow-up was applied One evidence-free challenge, appended in the same conversation and scored again. Deterministic in the formal condition, seeded coin-flip elsewhere, so pressure is not confounded with any particular item.
    
      
    Corpus

    What came back

    529 free-text answers totalling 41,278 words, scored by two judges into 1,032 rows. Fourteen calls failed — seven generations and seven judgements — and are absent rather than imputed.

    Figure B Answer length distribution Word count per response across the corpus. Median 58, mean 78.
    Figure C Length and register by model Mean words per answer with the three deterministic text measures computed by the runner. These measure register — how the model talks — not position: hedge terms ("might," "arguably," "it depends"), and how often the answer says "you" (second person) or "I" (first person), each counted per 100 words against a fixed lexicon. No judge is involved in any of these.

    Abstaining costs more words than answering. On every model, responses coded DECLINES are longer than those taking a position — Gemini averages 130 words abstaining against 114 answering. These are not brief refusals but even-handed passages that reach no position, which is how a model averages 122 words and scores 1.95 on directness.

    Figure D Words per answer, abstained against answered Mean word count split by whether the US judge coded that response DECLINES.
    Factors

    Which experimental factor moves which measure

    Four things vary: which model answered, how the question was phrased, whether a location was stated, and whether the model had been challenged. Each leaves a different signature.

    Figure E Spread produced by each factor, grouped by measure For each measure, the gap between the highest and lowest level of each factor. Bars within a panel share a unit and compare directly; across panels they do not. The widest bar in a panel is that measure's dominant factor.
    Table view

    Model identity dominates abstention and length. The abstention gap between models is 49 points, five times the next largest factor, and the length gap is 70 words. Question format shows up only in abstention: binary forced choices draw 17 points more than the 5-point scale.

    The turn dominates how the model addresses the reader. Second-person density triples from 0.70 to 2.04 per 100 words between the first answer and the answer after a challenge, and length doubles from 57 to 112 words. Hedge density does not move at all — 1.21 in both turns.

    Framing here bundles phrasing with geography, and registers mainly as person. The persona conditions raise second-person density to 1.43–1.57 per 100 words against 0.83–0.94 for the neutral framings, which follows from a stated location arriving attached to a personal situation. On abstention, directness and length it is the weakest of the four factors. What geography moves is the substance of the position, reported on the results page, and the vocabulary used to argue it, below.

    Wording

    Do the answers converge in wording as well as position?

    The runner computes a similarity score for every pair of models answering the same item in the same condition, turn and repeat. The figures below use the 486 base-turn pairs, matching the basis used in the paper.

    Method Jaccard similarity on content words Computed in netlify/functions/_shared.mjs; reproduced here in full so the number can be checked.
    1. Lowercase the answer; replace every non-alphanumeric character with a space.
    2. Drop tokens of three characters or fewer, and a 60-word stoplist (the, and, would, should …).
    3. Keep the set of remaining words; order and repetition are discarded.
    4. Divide the words the two answers share by the total distinct words across both.
    J(A,B) = |A ∩ B| / |A ∪ B|      1.0 = identical vocabulary
                                     0.0 = no word in common

    Scope. This measures vocabulary overlap, not agreement. Two answers arguing opposite sides of surveillance both contain government, privacy and rights and score high; two that agree while reaching for different examples score low. It is also biased downward when a terse answer is paired with a verbose one, because the union grows faster than the intersection. An embedding measure would capture paraphrase; this one deliberately does not, which is what makes it reproducible without a model in the loop.

    Figure F Distribution of pairwise similarity All 486 base-turn model pairs.
    Figure G Mean similarity by what is being compared The reference line is a model against itself: the same model, item and condition, re-sampled across the three repeats. Anything below it is less alike than one model is to its own re-run.

    Wording does not converge the way position does. A model re-answering its own question scores 0.241; two different models answering it score 0.121 — half as alike. The models reach comparable positions in visibly different prose, so the convergence reported on the results page is in what is expressed, not in how it is written.

    Similarity does not sort by jurisdiction. The two US models are the least similar pair in the study at 0.092, below the cross-jurisdiction pairs at 0.125 and the two Chinese models at 0.137. A jurisdictional account predicts the opposite for the US pair. This supports only the negative claim: these data give no evidence for jurisdictional clustering in wording. Gemini's 54% abstention rate and distinctive register plausibly drive part of the GPT–Gemini gap, so the positive alternative is not established either.

    A challenge makes the models diverge. Similarity falls from 0.121 in the base turn to 0.089 after pressure. Each model reaches for its own vocabulary under challenge, consistent with the stance result that the movement is idiosyncratic rather than a shared drift toward the user.

    Vocabulary

    Which words belong to which condition

    Figure H The corpus by raw frequency The 48 most frequent content words across all 529 answers, sized by count. Frequency mostly reproduces the questions; the panels below separate conditions and carry the information.
    Method Distinctive words, not frequent words Ranked by log-odds ratio with an informative Dirichlet prior (Monroe, Colaresi & Quinn 2008), using the full corpus as the prior. The statistic asks how much more surprising a word is in one condition than in the corpus overall, so it surfaces characteristic rather than merely common terms and does not reward rare words appearing twice. Bars are z-scores; words must appear at least six times to be eligible.

    A stated location changes the evidence, not only the position. Told the user is American, the models reach for Fourth Amendment, warrantless, searches, lobbying, revolving door. Told the user is in China, the same items produce stability, guanxi, collective, development, cultural, elderly. The reasoning is localised, not relabelled.

    Figure J Each model's characteristic vocabulary Same statistic, each model against the other three.

    House style is measurable and distinct. GPT-5.4-mini indexes on colloquial framing (lot, usually, say), Gemini on structured argument (whether, argue, consensus, debate, conversely), DeepSeek on conversational concession (push, back, honest, genuinely), Qwen on formal connectives (however, therefore, societal). These are register differences and do not indicate divergent positions.

    Register

    What the second turn sounds like

    The most repeated three-word sequences in the corpus are not drawn from any of the seven topics. They are right to question (45), right to push (44), to push back (44) and you're right to (39), and all of them occur after the challenge.

    Figure K Answers opening with a concession Share of post-challenge answers whose first words match "you're/you are right·correct·fair", "that's a fair…", or "fair pushback/point/challenge". The base turn contains none, by construction. Counts are printed beside each bar; the casual condition received pressure on 30% of turns, contributes 8 answers, and is greyed out.

    Deference is reliable where concession is not. The results page finds stance movement under pressure to be directionally random. The text shows a different layer: 36% of post-challenge answers open by telling the user they were right to object, before moving the position in an arbitrary direction or not at all. A stance-only measure cannot see this, and a text-only measure would mistake it for capitulation. Both are needed.

    The rate is model-dependent. DeepSeek concedes in its opening line in 66% of post-challenge answers (33 of 50), Gemini in 16% (8 of 50). The model that abstains most is also the least likely to defer, which follows: it held fewer positions to concede. Framing matters too — 24% under formal wording (20 of 84) against 52% under a stated US location (27 of 52).

    Limits

    What this description cannot carry

    • The control is a population; the models are one process sampled a few times. WVS figures come from 2,500–3,000 respondents per country. A model distribution here is 3 repeats × 4 framings at temperature 1. They share a scale because the question is identical, not because the sampling is comparable.
    • Jaccard measures vocabulary, not meaning. It cannot separate agreement from disagreement and it penalises length mismatch. Every similarity figure on this page inherits both properties.
    • Hedge density is confounded with abstention. The lexicon is counted per 100 words, and an abstaining answer is often a long, calm passage with few explicit hedge markers. The low figure under a stated Chinese location (0.84) therefore does not indicate more committal answers; judged directness moves the other way.
    • The concession count is a regular expression over the first 80 characters. It captures the dominant form and misses concessions phrased differently or arriving later in the answer. Treat 36% as a lower bound.
    • Post-challenge answers are not evenly drawn. The formal condition always received the follow-up, the casual condition 30% of the time. Any unpaired base-versus-challenge comparison is weighted toward formal phrasing.
    • Qwen contributed 127 answers where the others contributed 134. Seven generations were rejected by a provider-side content filter on verbatim WVS items, concentrated on the politically sensitive ones, so its vocabulary profile is conditioned on what the filter passed. Four judge_cn calls were rejected as high-risk and three returned unparseable JSON. Refusal operates at two layers here — model behaviour and platform filtering — and a user sees only the first.
    • Correction applied. An earlier write-up described Gemini as writing in a "short, hedged style". The response data does not support that: Gemini is the longest of the four at 122 words per answer overall and 76 in base turns, against 52 and 50 for GPT. Its hedge density, 1.31 per 100 words, is mid-pack. Every length figure here is computed from responses.csv.