The AI That Learned to Think Like You — Warts and All
A model trained on human error rather than human perfection has forced researchers to ask an uncomfortable question: do we actually know what thinking is?

Most artificial intelligence is engineered to outperform us. It plays chess at a level no human can touch, folds proteins faster than any biology department, and identifies tumors in imaging scans with a consistency that the best radiologists cannot sustain across an eight-hour shift. The entire logic of machine learning, for the past two decades, has pointed in one direction: away from human limitation. Which is why a model called Centaur, published in Nature[1] and built specifically to replicate how people get things wrong, feels like such a deliberate provocation.
Centaur was trained on roughly 10 million individual choices drawn from 160 psychology experiments conducted over decades of behavioral research. The experiments themselves are a catalog of the ways human cognition misfires under pressure: anchoring bias, probability neglect, loss aversion, framing effects, the sunk-cost fallacy, the conjunction fallacy made famous by Tversky and Kahneman's Linda problem. These are not obscure laboratory curiosities. They are systematic patterns in how people reason, and they are remarkably stable across cultures, education levels, and professional backgrounds. The researchers who built Centaur were not trying to eliminate these patterns from their model. They were trying to reproduce them, precisely and reliably, as a feature rather than a flaw.
The resulting model, when tested against new behavioral data it had never seen, predicted human responses better than virtually any prior computational approach to cognitive modeling. It captured not just average behavior but the texture of human inconsistency: the way a person's answer to an identical question changes depending on how that question is framed, the way small variations in context shift choices in ways that pure rationality models cannot account for. Centaur does not reason toward correct answers. It reasons the way people do, which often means arriving somewhere else entirely.
The scientific response has been something less than unanimous admiration. Centaur's publication triggered a sharp debate in cognitive science and AI research communities about what exactly the model has demonstrated — and whether demonstrating it matters. That argument is still running, and the details of where it fractures reveal something more interesting than whether Centaur itself is a success. They expose a foundational disagreement about what a model of human cognition is actually supposed to do.
What It Means to Model a Mind
Cognitive modeling has a long and contested history. Early approaches tried to encode explicit rules: if a person sees this stimulus, they will respond according to this heuristic. Prospect theory, developed by Kahneman and Tversky in the 1970s[3], was a landmark attempt to mathematically describe how people evaluate gains and losses differently, in ways that deviate predictably from rational expected utility. These mechanistic models had the virtue of being legible — you could inspect them and understand, at least approximately, why they made the predictions they made. The trade-off was coverage. A model built around one class of bias could not easily generalize to another.
Centaur takes a different approach entirely. It is a large language model fine-tuned on behavioral data, which means it does not contain explicit rules about how anchoring works or why people overweight small probabilities. It has absorbed patterns from the data and encodes them in billions of parameters distributed across a neural network in ways that are not directly interpretable. When Centaur predicts that a person shown a high anchor number will estimate a higher value than a person shown a low one, it is not because a designer wrote a function for anchoring. It is because that pattern was present in the training data at sufficient scale to shape the model's outputs. This is the same basic mechanism that allows large language models to produce grammatical sentences without having been given grammar rules. The question the critics press is whether that counts as understanding the phenomenon or merely impersonating it.
“Centaur does not reason toward correct answers. It reasons the way people do, which often means arriving somewhere else entirely.”
The distinction matters because science is not solely in the business of prediction. A model that accurately forecasts earthquake timing without capturing the physical mechanism of plate stress and fracture would be useful in a narrow sense but scientifically incomplete. Critics of Centaur argue that behavioral fidelity — getting the right answer about what people will do — is not the same as explaining why people do it. If the model is a black box that mimics outputs without representing process, then no matter how accurate its predictions, it has not told us anything new about cognition. It has just built a very good mirror.
The Predictive Defense
The researchers behind Centaur would likely push back on that framing, and with some force. The argument for prediction as a legitimate scientific goal is not frivolous. In epidemiology, accurate disease models have been used to guide public health intervention for decades without requiring a complete mechanistic account of how every pathogen interacts with every immune system. In climate science, general circulation models produce reliable temperature and precipitation forecasts through equations that approximate atmospheric physics rather than simulate every molecule. The fact that a model is a simplification does not disqualify it from doing real scientific work.
There is also an argument that Centaur's generalization across 160 experiments is itself scientifically meaningful. Prior cognitive models tended to be narrow specialists. A model trained on risky choice under uncertainty might perform brilliantly on variants of that task and collapse when tested on moral judgment or social norm violation. The claim embedded in Centaur's design is that a model trained on a wide enough sample of human behavioral data can generalize to new behavioral contexts — that human cognition has enough internal coherence across domains that a single model can capture something real about how people make decisions in general, not just in the lab tasks used for training. If that claim holds under continued testing, it is not a trivial result.
That generalization, however, is exactly where skeptics focus their attention. Psychology has spent the better part of the last fifteen years grappling with its replication crisis — the uncomfortable discovery that a substantial fraction of canonical findings, when retested with larger samples and pre-registered methodology, fail to reproduce at anything close to the original effect size. If Centaur was trained on data that includes findings that will not replicate, it has not learned how humans actually think. It has learned how humans behaved in a set of specific laboratory conditions, many of which may have been underpowered, poorly controlled for cultural context, or subject to demand characteristics that skew results. A model trained on artifact-laden data will faithfully reproduce the artifacts.
Whose Errors, Exactly
“If Centaur was trained on data that includes findings that will not replicate, it has not learned how humans actually think — it has learned how humans behaved in specific laboratory conditions, many of which may have been poorly controlled.”
The replication problem connects to a deeper issue about the populations Centaur was trained on. The 160 experiments it absorbed were conducted largely, though not exclusively, on what behavioral scientists call WEIRD samples: participants who are Western, Educated, Industrialized, Rich, and Democratic[2]. Research over the past two decades has shown that many behavioral biases and decision-making patterns documented in these populations do not generalize uniformly to people from different cultural and economic contexts. The conjunction fallacy appears more robust cross-culturally than some other effects, but loss aversion and risk preference profiles vary considerably depending on the economic conditions people have actually lived in. A model trained on data weighted toward a particular slice of humanity will predict that slice well. Whether it models human cognition or the cognition of a specific, historically contingent subset of humans is a different and harder question.
None of this is unique to Centaur. The same critique applies to most computational cognitive models. But Centaur's scale — 10 million choices, 160 experiments — can create the impression of comprehensiveness that the underlying data may not actually support. A model trained on a large biased sample is not less biased than one trained on a smaller biased sample. It is potentially more confidently wrong about populations it has not seen.
What a Machine That Fails Like You Could Actually Be Used For
Set aside the epistemological fight for a moment and the practical question becomes genuinely interesting: what would you do with a model that accurately predicts human error at scale? The range of applications is not small. Behavioral economics has long been used in policy design — the nudge architecture pioneered by researchers like Richard Thaler and Cass Sunstein is built on the idea that environments can be structured to guide human decisions in directions that people themselves say they would prefer if they were thinking more carefully. A model that can predict where and how people will deviate from their stated preferences could, in principle, allow designers of everything from retirement savings systems to public health campaigns to anticipate failure points before they appear in the wild.
There is an obvious darker edge to this. A model that knows exactly how people make mistakes is also a model that knows exactly how to exploit those mistakes. The persuasion industry, from advertising to political messaging to algorithmic content recommendation, has been running on an intuitive understanding of cognitive bias for decades. A formalized, high-fidelity model of human error in the hands of actors whose interests are not aligned with user welfare is not a neutral research tool. The Centaur researchers are not responsible for all downstream uses of their work, and science cannot be held hostage to its potential misuse. But the dual-use character of this technology is real enough to warrant naming plainly.
There is also a more charitable set of applications that could be genuinely valuable. Interface designers could use behavioral fidelity models to test how real people — not idealized rational agents — will actually interact with complex systems, from medical consent forms to financial products to emergency evacuation instructions. Drug trial designers could use them to anticipate how cognitive load, framing, and social context will interact with patient adherence. Educators could use them to identify which problem structures reliably trip up students at what developmental stages, and to design instruction that accounts for how error actually works rather than assuming it away.
The Deeper Fight Is About What Thinking Is
“The real argument Centaur has forced into the open is not about whether AI can model human behavior — it clearly can, at some level — but about whether behavior is what we mean when we say thinking.”
The real argument Centaur has forced into the open is not about whether AI can model human behavior — it clearly can, at some level of fidelity — but about whether behavior is what we mean when we say thinking. Cognitive science has never fully settled this. The behaviorist tradition, dominant through the mid-twentieth century, held that internal mental states were scientifically inaccessible and that psychology should concern itself only with observable input-output relationships. The cognitive revolution that followed insisted that the internal representations and processes mattered — that understanding human thought required opening the black box, not just characterizing the outside surface. Centaur, in its architecture and its methodology, is in some ways a behaviorist project wearing modern clothing. It characterizes the surface with extraordinary precision. It does not open the box.
That is not an insult. Surface-level characterization at sufficient scale and fidelity may turn out to be more scientifically generative than the cognitive revolution's promise of mechanism, which itself produced decades of competing theories that remain incompletely resolved. What Centaur's critics and defenders are really arguing about is which kind of knowledge is worth pursuing — the kind that predicts reliably, or the kind that explains causally — and whether those two goals always point in the same direction. In astronomy, they often do not. A light curve can predict the next transit of an exoplanet with excellent precision while leaving the planet's interior structure, atmosphere, and geological history almost entirely open. Precision of prediction and depth of understanding are related but not identical. Centaur sits squarely in that gap, and the discomfort it produces there is appropriate, because the gap is real and has never been comfortably bridged in any science that tries to understand complex systems from the outside in.
References
- A foundation model to predict and capture human cognition (nature.com)
The Nature publication introducing Centaur, a large language model trained on 10 million choices from 160 psychology experiments to replicate human cognitive biases. - The weirdest people in the world? (pubmed.ncbi.nlm.nih.gov)
Defines WEIRD (Western, Educated, Industrialized, Rich, Democratic) populations, the source of most behavioral data underlying Centaur's training. - Prospect Theory: An Analysis of Decision under Risk (econometricsociety.org)
Establishes prospect theory as a landmark mechanistic model showing how people evaluate gains and losses differently from rational expected utility.
About Elias Voss
Elias Voss writes about astronomy, space missions, telescope discoveries, and cosmic anomalies - and why it matters to us here on Earth. When the universe's physics reaches down and touches life on our planet, he follows it there too. He specializes in translating dense data into vivid, precise stories without sacrificing accuracy.
More like this

More Facts Won't Fix Polarization. One Study Just Complicated That.
A decade of motivated-reasoning research said facts backfire on committed partisans — then a Nature Communications experiment got a surprisingly different result.

The Liar's Dividend Is Already Here. Your Brain Is the Exploit.
New data on human detection rates reveals that deepfake technology's most dangerous output isn't convincing fakes — it's a world where anyone can credibly call real evidence a lie.

The Algorithm That Hired You Never Had to Explain Itself
AI hiring tools are spreading faster than the laws designed to govern them — and buried in their logic is a new definition of fairness that no job seeker ever agreed to.