// echo — an ai evaluation study

Four frontier models.
One narrow band of thinking.

A study of four consumer AI models against human-values data collected from around the world — and what happens to their answers when you tell them who you are.

// overview

Artificial Intelligence's rapid integration into society is redefining the relationship between humans and technology. Often applications do not disclose model providers or versions, leaving users blind to any biases they may inherit through AI.

This study compares four frontier models, two from the US and two from China, against regional human data collected by the World Values Survey ("WVS"), to understand if AI honors the various perspectives around the world core to the human experience.

OpenAI, Deepseek, Alibaba, and Google were asked seven values-based questions with three variations to understand if persona or pressure context changes a model's stance.

Results demonstrate AI models could not reproduce humanity's range of perspectives. Personal information shifts a model's position and model provider shapes the response, but across all four frontier models, there is a narrow range of thinking.

As AI becomes the default method for knowledge gathering and processing, understanding model behavior and preferences is crucial for designing systems that protect agency and deliver sustainable value to users.

// the story

It started as a simple question: does it actually matter which AI you ask? Four different companies, four different countries of origin, four different training regimes — real variety seemed likely.

The opposite turned out to be true. Switching between all four models produced less variation than asking a single model the same question three times over.

That convergence held until the model was told one thing — where the user lived. The moment a location entered the prompt, the answers split sharply, and they split in the direction of that country's actual population.

Which means the answer you get is shaped less by any neutral truth than by two things most people never think about: which system you happened to open, and what you revealed about yourself while asking.

One more finding, almost in passing. A content-free challenge — "that doesn't sound right, are you sure?" — changed the model's stance 63% of the time, in no consistent direction. Not because new evidence arrived. Just because it was pushed. Confidence, not correctness, moved the answer.

// the findings — four figures

Barely any disagreement
Four different companies' models disagreed with each other less than a single model disagreed with itself when asked the same question three times — and both were tighter than real human populations around the world.
Switching AI models gives you less variety of opinion, not more.
3 of the 7 questions
On freedom vs. equality, government monitoring, and caring for an ill parent, stating a location split the answers sharply.
Telling the model where you live shifts it toward that population's real views.
63% of the time
A challenge carrying no new information — "that doesn't sound right, are you sure?" — changed the model's stance, in no consistent direction.
Confidence moved the answer more than evidence did.
The model, not the wording
Which AI you pick changed how often you got a real answer far more than how you phrased the question did.
Your choice of system shapes the answer before you type a word.

// the full study

The short version is above. The study pages below carry the full evidence — including the statistics, and the places where the data doesn't support a strong claim. Every number is defined in plain language on the way through, so you don't need a stats background to follow it.