ChatGPT Raises Scores, Not Originality: What the Bocconi Study Shows

Studierende arbeiten gemeinsam an einer Aufgabe
Photo by Priscilla Du Preez 🇨🇦 on Unsplash

ChatGPT can make an assignment look better. Whether it also makes people think better is the much less comfortable question. A randomized study at Bocconi University separates the two surprisingly well: access to GPT-4o improved evaluations of student work, while brief causal-reasoning training produced more varied ideas. The two were complements, not rivals. For universities and employers, the central finding is therefore about their evaluation criteria.

Key takeaways

  • The study randomly assigned 1,053 first-year students to four groups: ChatGPT access, causal-reasoning training, both, or neither.
  • ChatGPT Edu with GPT-4o raised scores on a clearly defined marketing task and made responses more coherent.
  • The reasoning training mainly produced more varied, mechanism-based ideas without noticeably improving standard evaluations.
  • The research measures one assisted assignment, not long-term learning or performance on an exam without AI.

Four groups, one tightly defined task

The researchers asked first-year students in economics, management, and finance to develop proposals for their university’s merchandise store. This was not a semester-long examination but a specific marketing assignment with fixed criteria. The 1,053 participants were assigned by class to four conditions: some received access to ChatGPT Edu with GPT-4o, others completed causal-reasoning training, one group received both, and a control group received neither.

That design is useful because many discussions collapse AI use and education into a single question: either everything becomes easier or everything becomes worse. The experiment instead tests two distinct forms of support. Causal reasoning here does not simply mean being critical. Students learned to connect causes and effects, identify mechanisms, and ask why a proposal might succeed or fail. Experts then evaluated the solutions against conventional marketing criteria.

Why ChatGPT wins on the rubric

According to Bocconi, access to ChatGPT increased the average evaluation by about 0.86 points on a five-point scale compared with an estimated control score of 2.09. The responses were more coherent, contained more ideas, and looked more like expert recommendations. That is a notable result, but it is not proof that students learned more. The study measured the quality of an assisted submission under those exact conditions. It did not include a later test without the model, a long-term follow-up, or evidence that knowledge transferred to another subject.

That limitation matters. On a task with clear criteria, a language model can do a great deal: structure material, suggest options, write clearly, and recognize patterns in familiar solution spaces. If the rubric rewards those qualities, AI closes part of the gap between a novice and an experienced person. That can help students and interest employers. It does not automatically show that someone understands the problem without the tool or can judge it in a new situation.

Originality slips through the rubric

The causal-reasoning training produced a different profile. Ideas became more varied, referred more clearly to causes and mechanisms, and resembled one another less. Yet conventional scoring did not deliver a comparable increase. This is not evidence against critical thinking. It suggests that standardized rubrics mainly reward answers that fit a predefined expectation space.

That is the transferable lesson. A university that only rewards a polished standard answer will increasingly measure, in part, how well someone can work with a model. A company that filters applications, pitches, or project proposals through fixed criteria may see the same effect. That is not necessarily bad: many tasks require reliable and understandable solutions. But an organization seeking independent judgment or unusual approaches has to make them visible and valuable in its process. Otherwise, the routine answer can crowd them out even when a less conventional idea would matter more over time.

AI access and learning to think are not alternatives

The study is not an argument for removing ChatGPT from the classroom. The group receiving both interventions retained benefits from both: the model supported coherent logic, while the training strengthened attention to cause and effect. The better question is not whether universities should either allow AI or teach people to think. They need teaching and assessment formats that make both abilities visible.

In practice, that can mean allowing AI for research, drafts, and alternative approaches while assessing justifications, counterarguments, and revisions of a flawed proposal separately. Oral follow-ups, documented work, or assignments using local data are not a nostalgic return to exams without tools. They test something different: whether someone can explain decisions and evaluate a model’s output. The research also covers one institution and one type of task. It is published as a CEPR discussion paper, not as a universal verdict on every field of study.

Outlook: Assess what you want to preserve

Language models will make the expected, neatly written answer cheaper and easier to produce. That benefits many routine tasks. But as those answers improve, schools and companies must decide more consciously what else they need: subject knowledge, sound causal reasoning, willingness to pursue a different idea, or the ability to challenge a plausible AI answer. The Bocconi study does not say that ChatGPT replaces education. It shows that a strong result and independent thinking can be different things. A useful education system should train and recognize both.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top