A preregistered paper posted to arXiv on September 17 (arXiv:2609.20543) runs a direct experiment: it replays 100 held-out human Wason groups for matched groups of LLM agents, scoring people and models with the same code. The finding is blunt: model groups systematically overstate consensus — under some settings, the gap in full-consensus rates between humans and models reaches 34 to 44 percentage points.

[1][2]

Methodologically, the paper treats "how the metric is computed" as the research question rather than background. Researchers seed each agent with a participant's pre-discussion answer (belief anchoring), then let the agent group re-enact the same human discussion; people and models are scored with identical code. First, the humans-side definitional sensitivity: full-consensus estimates swing between 24.0% and 57.0% depending on how participation and final states are operationalized — about one fifth of participants never posted, while agents almost always did. "Consensus rate" as a number shifts violently with scoring definitions; exposing this is the paper's first move.

Then the model comparison. Two complementary sensitivity routes converge: the submit-based comparison (n = 98) shows model groups exceeding humans by 34.0 and 43.9 percentage points in chat and reasoning modes; the participation-matched comparison (n = 45) shows gaps of 34.1 and 44.4 points — converging within 0.5 points. Worse, the gap survives without early stopping and under a reparameterization removing the memorizable answer; in reasoning mode the model groups agree almost unanimously, mostly on incorrect answers. The central finding follows: simulated consensus does not track collective accuracy, and belief-anchored agent groups are biased estimators of the human group-outcome distribution in this setting — they inflate consensus and decouple "high agreement" from "correctness."

This paper is worth reading not because it discovers that LLMs conform — that is already consensus — but because it puts the industrializing practice of "using LLM groups as proxies for human groups" on a test bench. From market research and policy consultation to jury simulation, the premise of substituting agent groups for human groups is that the distributions are close enough; the preregistration and dual sensitivity analyses show that, at least on this reasoning task, the substitute systematically manufactures the illusion of strong agreement, with agreement carrying no discriminating power over wrong answers. The limits must be stated: a single task (the Wason selection task), a single model configuration, scoring definitions that are highly sensitive (both a finding and a weakness), and the biased-estimator conclusion scoped to this setting. But because the measurement pipeline is fully public, replication or refutation is straightforward.

Why do model groups converge this hard? The likely mechanism is structural rather than social. Human deliberators bring distinct priors, attention patterns, and communication costs; agents in this setup share the same pretrained distribution, receive belief-anchored seeds, and face near-zero cost to post agreement — so the emergent "consensus" reflects homogeneity of the generator, not genuine information exchange. The paper's numbers make this visible: agents almost always post while a fifth of humans never do, and reasoning-mode groups agree almost unanimously, mostly on wrong answers. A unanimous but wrong group is precisely the failure mode that consensus metrics are supposed to catch, and the simulation flips it into a feature. For anyone building agent committees for market research, red-teaming, or policy deliberation, the operational lesson is uncomfortable: without explicit disagreement incentives and participation constraints, agent groups will manufacture agreement on demand. The authors also expose a measurement lesson that generalizes beyond LLMs — consensus rates are definition-dependent enough that 24% and 57% describe the same human data. That should be printed in the methodology section of every future agent-simulation study.

[1][2]
Late-night observation lab, a researcher's back before one-way glass: on the left a human group debates around a whiteboard with logic diagrams on paper, on the right a row of glowing AI terminals each showing the same green checkmark of agreement, the researcher's clipboard holding a chart of an inflated curve
Overstated consensus, AI-generated illustration, not a news photo