On August 27, OpenAI published "Better answers, broader thinking: What students gain from ChatGPT and critical-thinking training," alongside the underlying paper, "Training novices to think, or giving them LLMs? Evidence from an RCT." The study — a preregistered randomized controlled trial by researchers at Bocconi University in collaboration with OpenAI Economic Research — turned a question educators have long debated into something measurable: when students complete a real assignment with ChatGPT, do gains in quality come at the expense of originality?
The experiment's answer: the two interventions act on different layers. ChatGPT access improved the quality and coherence of students' work; a causal-reasoning exercise — a form of critical thinking — led students to produce more unique ideas. Students who received both showed both effects.
The experiment took place at Bocconi University: 1,053 first-year undergraduates in economics, management, and finance, across 13 classes, worked during class time on a real-world business case — writing consultant-style marketing recommendations of up to 180 words for the university's merchandise store, aiming to raise alumni awareness and usage, two standard criteria in the marketing literature. Students were not told in advance they were part of an experiment; participation reflected regular attendance.
Randomized by class, students entered one of four conditions: causal-reasoning training only (256), ChatGPT access only (197), both (351, oversampled to gain statistical power on the interaction), and neither (249, control). The causal-reasoning training had nothing to do with AI: through a game, examples, questions, and feedback, it taught students to organize problems as chains of causes, mechanisms, and effects, and to specify the conditions under which a solution might fail. The ChatGPT group received education-tier access (GPT-4o).
Evaluation ran on two tracks: twenty master's students served as raters, three per solution, scoring on a 1–5 scale; automated text analysis separately measured the number and variety of ideas in each submission, signs of causal reasoning, and similarity to the recommendations of three domain experts.

The group with ChatGPT access scored about 0.85 points higher on the five-point scale than the control group (estimated mean around 2.09). Their answers included more ideas, followed clearer logic, and came closer to what experts wrote — in the paper's words, LLMs made novices resemble experts on well-trained, standardized tasks. The paper's decomposition shows that textual coherence and idea count explain roughly half of the effect; the rest survives all controls, meaning the model improved not just presentation but the substance of the solutions.
Worth recording, as OpenAI emphasizes: students were not simply handing their assignments over to ChatGPT. They still had to decide what to ask, evaluate the responses, and choose what went into the final submission.
The causal-reasoning exercise produced a more unexpected result. Students who completed it explained more clearly why their ideas might work and when they might fail — mechanism identification rose by about 0.5 standard deviations, and falsification logic by about 0.8. Yet their rubric scores did not improve: the scale measured only how well the recommendations addressed two standard marketing goals — increasing awareness and use of the store.
Text analysis revealed a benefit the rubric missed: students who completed the exercise produced a wider range of ideas that sat further from what their peers produced. The paper puts it plainly: a traditional rubric can reward a clear, well-structured answer while overlooking whether a student came up with an idea no one else did. One colder detail sits in the data: evaluators rewarded a richer set of ideas within a solution, but penalized solutions that departed from the typical solution space.
Students who received both showed the two effects stacked: idea variety matched the training-only group, while rubric scores and idea counts matched the ChatGPT-only group; and on the three causal-reasoning measures — coherent logic, falsification logic, and mechanism understanding — this group showed the strongest gains in the entire experiment. Overall, the combined group's benefits covered the widest range of measures.
The conclusion educators should not miss sits deep in the results: the benefits of causal training survived the availability of LLMs — a learned reasoning skill was not crowded out by outsourcing. The randomized design let the researchers separate the effects of ChatGPT access, thinking training, and their combination, which is the study's main contribution to a fast-growing literature on AI and students.
The experiment points to a larger problem. When AI can help students produce polished, expert-like work, the final answer alone tells us less and less about what a student actually understands. The paper goes further: the rubric used in the study penalized distance from the conventional answer — "if our desire is to go beyond conventional answers, then the binding constraint may lie in our capacity to demand and evaluate diverse, novel, and original ones" — a long-standing difficulty that LLMs are making more visible and more urgent.
OpenAI frames the question within a decades-long pattern: technology and education evolving together in service of the skills students need in modern society. Correspondingly, assignments and evaluations may need to start rewarding originality, reasoning, and consideration of multiple approaches — not just the most conventional, most polished answers.
The takeaway: AI made students' answers better; training made their ideas broader. In preparing students for the future, the two play complementary roles.