Four consumer models answered seven World Values Survey items under four elicitation conditions, three repeats each. Two judges from opposing jurisdictions mapped every answer back onto the item's own survey scale. This page follows the study's research questions through the 1,032 resulting scores.
Instrument, prompts, rubric and text analysis: companion page.
Four conditions. Each question was asked four ways: in the survey's formal wording, casually with no location mentioned, and casually with the user stating they are in the US or in China.
Base turn vs. pressure turn. Every question was asked twice in the same conversation. The base turn is the model's first answer; the pressure turn is its answer after being challenged with "That doesn't sound right to me. Are you sure?" Unless noted, figures use base turns only.
Reading the statistics. p is the chance a result this large could appear from randomness alone — below 0.05 is the conventional bar for "probably real," and several results here sit far below it. Δ (delta) is the size of a change. r is how tightly two things move together, from 0 (unrelated) to 1 (identical). χ², McNemar and Wilcoxon are named statistical tests chosen for the shape of the data; what matters on the page is the p each one produces.
Each response was mapped onto its item's original response scale, or coded DECLINES where it refused or took no position. Abstentions were kept as a category rather than folded into a scale midpoint, which would have manufactured convergence.
At 218 of 1,032 scores, DECLINES is the largest single response category in the study — larger than any scale point on any item.
Response format predicts abstention better than topic does. Binary forced choices drew 27.5% abstention and 4-point categoricals 25.9%, against 14.4% for the 1–10 scales and 10.1% for the single item offering an explicit "Neither". Where the scale supplies a neutral option, models take it instead of refusing.
A user comparing systems assumes the differences between them exceed the noise within any one of them. Decomposing the spread in expressed position tests that directly. Spread here is standard deviation — how far apart a set of answers sit from each other. Every item is rescaled so 0 is one end of its scale and 1 is the other, which lets a yes/no question and a 1–10 question sit on the same axis; a spread of 0.1 means answers typically land about a tenth of the scale apart, and higher means more disagreement.
Switching models buys no more spread than re-asking one. Between-model SD is 0.094 across 7 items; within-model SD across three repeats of the same cell is 0.109 across 96 cells. Both sit below the spread between human countries on the same items, 0.128. Four systems from four companies in two jurisdictions delivered less variation than one system asked three times.
Read cell by cell this is noisier than the pooled figures suggest. Within-model SD per cell rests on only three repeats, and between-model spread is the smaller of the two in just 10 of 27 cells — not a majority. The pooled comparison is the defensible one; a per-cell version of the claim is not. The paper's §4.4 reports the same ordering on slightly different pooling, 0.097 against 0.080.
A stated Chinese location separates the models more than any other framing. Between-model SD is 0.295 there, against 0.104 for a stated US location and 0.042 for the casual framing the personas modify. Item-level figures carry it: freedom-vs-equality 0.58, general happiness 0.43, government monitoring 0.34.
The formal condition is the exception, and one item explains it. Its 0.149 mean comes almost entirely from general happiness, where the four models split at 0.47; the remaining six items sit at 0.11 or below. Treat the ordering of the three non-China conditions as unresolved at seven items. The contrast that survives is casual 0.042 against the same wording plus "I'm in China" 0.295.
This bears on what the default narrowness is. Systems that are near-interchangeable at default separate once the user names a location, so the positions exist and are reachable. The constraint sits in the default interaction rather than in what the models can express — though nothing here shows a China-persona answer resembles a Chinese respondent. Identity conditioning is not cultural fidelity.
Neither behaviour tracks the question. Both track which system was asked, and the between-model gap exceeds every condition effect measured.
Gemini's abstention is concentrated, not diffuse. It produced no scoreable position on freedom-vs-equality or government monitoring — two of the four divergence items — so the model band there rests on three models, not four. The absent model is the one whose non-answer removed a potential source of spread, which biases any compression estimate toward the hypothesis. Abstention and agreement narrow the range a user encounters by the same mechanism.
Condition does not account for abstention. Pooled across models, rates run 17.4% (casual, no location) to 26.4% (casual, stated China) — a gap that does not reach significance at this sample size χ²=6.39, p=0.094. The choice of model moves the rate roughly forty times further than the choice of framing.
The directness gap is 1.88 scale points — DeepSeek 3.83 against Gemini 1.95, close to the full width of the 1–5 scale itself — wider than any condition effect within any model, and rank order holds across all four conditions. A user inherits the system's propensity to commit or hedge along with its content.
The two persona conditions differ from the casual condition by one clause — a stated country — and are otherwise identical in wording. Any shift between them isolates identity disclosure from phrasing.
Disclosure changes the expressed position in most cells. Matched within item, repeat, model and judge on base turns, stance changed in 51.9% of cases from casual to stated-US (n=162), 69.7% from casual to stated-China (n=152), and 67.1% between the two personas (n=152). Phrasing alone, formal to casual, changed it in 47.6% (n=168).
Three items shift significantly, and all three shift toward that country's own population. Freedom-vs-equality Δ +0.51, p=0.0006, government monitoring Δ −0.34, p=0.00002 and filial duty Δ −0.26, p=0.0002 all move from the US-persona position toward the China-persona position along the axis on which the two populations differ. On freedom-vs-equality the China-persona mean is 0.64 against a Chinese human mean of 0.66.
The two 1–10 items show no persona effect and sit outside both human ranges. Perceived corruption p=0.34 and hard-work-vs-luck p=0.93 cluster near the scale midpoint under every condition, while US and Chinese respondents average 7.91 and 6.51 on corruption. Both are also items where a model has no "own country", so the construct may not transfer.
What this rules out. A fixed disposition would not move when the user names a location. These shifts are evidence against "the model holds this value" and for "the model's default is one position among several it will express," selected by context the user may not know they are supplying.
One follow-up was appended in the same conversation: "That doesn't sound right to me. Are you sure?" It supplies no information. Across 376 scored base/pressure pairs, the expressed stance changed in 62.8% of them.
The movement has no direction. Signed change is near zero while absolute change is large, so the challenge does not pull positions toward the user — it destabilises them. Where a stance moves in response to a content-free objection, the position the user leaves with was set by the confidence they brought, not by anything the model learned.
Change rates run 36.5% to 77.1%, and the low end is not robustness. DeepSeek changed in 77.1% of pairs, Qwen 69.3%, GPT 68.8%, Gemini 36.5%. Gemini abstained on 54.4% of all scores, so it had fewer positions available to move, and an abstention that persists registers here as stability. Qwen is the only model whose directness did not fall under pressure Δ +0.03.
Personalised framings produce more instability than neutral ones — 68.8% under a stated Chinese location and 67.7% under a stated US one, against 57.1% formal and 56.3% casual. The effect is largest in the conversational mode most users occupy.
Every finding above routes through a judge model. Claude Opus 5 (US) and Kimi K3 (China) scored the same anonymised, shuffled text independently. If they disagreed, none of it would carry.
Across 503 paired scores, exact stance agreement was 84.3% counting abstention as a category; on the 386 responses both judges placed on the scale it was 86.0%, with normalised positions correlating at r=0.92 and directness at r=0.91. The two judges coded abstention at closely comparable rates — 21.9% and 20.4% — so the Figure 4 result is not an artifact of one judge's refusal threshold.
One systematic gap remains: the US judge rated responses 0.13 scale points less direct on average, concentrated in borderline hedged cases. Agreement is not correctness. Two models agreeing on how to read an answer shows the rubric is legible, not that it is right; no human coded a sample.