A preprint posted to arXiv on September 17 (2609.20543, sole author Tengfei Shao) pours cold water on the increasingly popular practice of simulating human group deliberation with groups of LLM agents. The researcher replayed 100 held-out human Wason selection-task groups with matched groups of LLM agents, ran them through the same protocol, and scored agents and people with identical code. The conclusion: agent groups systematically overstate consensus — and their "agreement" does not track accuracy.

Methodologic

[1]

ally, the paper is careful. Each participant's pre-discussion answer was seeded into a belief-anchored agent; discussion protocol and scoring code were identical for humans and agents; the study ran under a preregistration (37 pages, 4 figures). Human full-consensus rates were themselves measurement-dependent — ranging from 24.0% to 57.0% across scoring definitions — and about a fifth of human participants never posted, while agents almost always did.

The core result comes from two complementary post-unblinding sensitivity analyses. In the submit-based comparison (n=98), consensus-rate gaps between agent and human groups were 34.0 percentage points in chat mode and 43.9 in reasoning mode; in the participation-matched comparison (n=45), 34.1 and 44.4 points. The two routes remove different measurement asymmetries yet converge within 0.5 percentage points — which the author reads as evidence that the overstatement is behavioral, not an artifact of scoring.

Two further controls are more worrying. The gap persists without early stopping, and under a reparameterization that removes the memorizable answer, reasoning-mode agent groups approach unanimous agreement — mostly on incorrect answers. Simulated consensus does not track collective accuracy, and belief-anchored agent groups are biased estimators of the human group-outcome distribution. A group of LLMs agreeing is neither "correct" nor "human-like."

The boundaries deserve stating: single author, single task (Wason), single round of deliberation; whether the belief-anchoring scheme fairly represents human cognition is itself contestable; agents' near-always-posting behavior differs from human silence patterns. The paper does not show that AI group simulation is useless. It shows that if researchers treat consensus rates as a health indicator of collective cognition, they will be systematically misled — agent agreement mostly reflects architectural convergence, not persuasive reasoning.

For the fast-growing fields of AI social simulation and prediction markets, this is a methodological warning: when simulated groups reach consensus more easily than human groups, consensus itself is worth less as a signal. The author leaves the point in the conclusion: to evaluate simulated-group estimates of human deliberative outcomes, you need scoring-explicit baselines — prove they behave like people before you let them decide for people.

A word on why this matters beyond the single task. Wason selection tasks are the standard minimal setting for studying human reasoning; if agent groups cannot reproduce human disagreement patterns in the simplest deliberative setup, extrapolating to richer settings — forecasting markets, citizen assemblies, political deliberation — inherits the same bias, amplified. The paper does not claim agents cannot simulate humans at all; it claims the default assumption that they do, without scoring-explicit checks, is unjustified. Its contribution is procedural: a template for measuring, rather than assuming, the gap between simulated and real deliberation. That is a duller claim than "AI groups reach consensus," and a more useful one.

[1]
深夜研讨室,研究者背影站在白板前,白板左侧画着散开的人类观点点阵、右侧是一束几乎重叠的智能体观点曲线,研究者手持黑色马克笔正要在两者之间画一条连线
深夜研讨室白板前两组观点曲线对照的编辑级插画, AI 生成插画,非新闻照片