Contacts

Better evaluations with ChatGPT. Critical thinking broadens ideas

, by Ezio Renda
An experiment involving 1,053 students shows that AI can bring novices closer to experts on well-defined tasks. Causal reasoning, however, generates less conventional solutions. And it raises a question for universities and companies: are we sure we are assessing the right skills?

ChatGPT can help students improve the quality and coherence of their work, helping them to sound more expert. It also reinforces the logic provided by a training on causal thinking, but for producing more diverse ideas causal thinking remains the primary driver with ChaptGPT not penalizing the diversity of ideas, but not enhancing it either. That is the far-from-trivial distinction emerging from an experiment involving more than a thousand students: those who use AI receive higher evaluations, while those trained in causal reasoning produce ideas that differ more from those of their peers.

The findings come from Training novices to think, or giving them LLMs? Evidence from an RCT, a study by Hemanth Asirvatham, Chiara Betti, Rachel Brown, Arnaldo Camuffo, Aaron Chatterji, Chiara Fumagalli, Alfonso Gambardella, Myriam Mariani, Abhinev Pandey, Steve Ramos, Giovanni Salvucci and Vladimir Simic. The study grew out of a collaboration between Bocconi and OpenAI, with researchers also affiliated with Duke, Berkeley, SDA Bocconi and the ION Management Science Lab.

One thousand students, four groups

“We conducted a randomized controlled trial with 1,053 first-year students in economics, finance and management,” explains Alfonso Gambardella, professor at Bocconi and co-director of the ION Management Science Lab.

The students were divided into four groups: some received training in causal reasoning, some were allowed to use ChatGPT Edu with GPT-4o, some had access to both, while the remaining students had neither the training nor AI. All were asked to propose ways to increase awareness and use of Bocconi merchandise among alumni: a standard assessment in the marketing of new products and therefore a task particularly well suited to an LLM.

The answers were evaluated on a scale from 1 to 5, with higher marks associate to a greater level of awareness and usage according to this standard metrics. And this is where ChatGPT made a clear difference: access to the model increased scores by around 0.86 points compared with an estimated score of 2.09 for the control group, while also making students’ proposals more similar to those produced by experts.

This is not simply a matter of presentation. More coherent writing and a greater number of ideas explain part of the advantage, but not all of it: a substantial share of the effect remains even after controlling for these characteristics. In other words, AI also improves the substance of the solutions.

Originality does not necessarily earn a better evaluation

The picture changes when the focus shifts to the diversity of ideas. Training in causal reasoning encourages students to spell out causes, effects and underlying mechanisms more clearly and, above all, to produce solutions that are more different from those of other participants. ChatGPT, by contrast, mainly increases the number of ideas, not their originality.

The paradox is that these more original answers are not evaluated more highly. When predefined metrics are used, evaluations tend to reward answers that fit those metrics, but this also means penalizing solutions that move beyond the conventional solution space. The problem, then, is not that causal reasoning fails: it is that the evaluation system is calibrated around predefined answers and does not reward different ones.

This may be the paper’s most provocative implication. If an LLM becomes very good at producing the expected answer, continuing to assess students or employees primarily on their ability to provide that answer risks discouraging solutions that, in some cases, could be innovative. The solution is to do both, or to put it simply: we need even more teaching of basic principles in a world of LLMs because they complement each other in valuable ways.

An issue for managers and policymakers too

The implications directly concern companies. If recruitment, promotions, pitches and project approvals are based on codified criteria, AI can rapidly narrow the gap between junior employees and experts. But if an organization is looking for independent thinking and unconventional solutions, it needs assessment systems that genuinely reward them.

The same applies to universities and policymakers. The study does not suggest choosing between teaching people how to think and allowing them to use ChatGPT. Quite the opposite: students trained in causal reasoning continue to apply it even when they have access to AI, while combining that training with GPT strengthens some aspects of causal logic.

“This study offers an encouraging glimpse of what’s possible when AI and critical-thinking training work together. The opportunity ahead is to pair new technology with thoughtful teaching to help students think more deeply, explore more possibilities, and better prepare for the future. I look forward to seeing more research into how educators can put that promise into practice and thank Bocconi University for their partnership on this important work,” said Leah Belsky, Vice President of Education at OpenAI.

 “Understanding the impact of AI requires systematic data and experiments,” Gambardella stresses. “Our study is just one piece of a mosaic that needs to be completed with many more studies like ours.”

A mosaic in which the question is no longer simply what AI can do in our place, but which human capabilities we want to keep developing — and, above all, whether we are truly able to recognize them when we see them.

ALFONSO GAMBARDELLA

Bocconi University
Department of Management and Technology

MYRIAM MARIANI

Bocconi University
Department of Management and Technology

ARNALDO CAMUFFO

Bocconi University
Department of Management and Technology

CHIARA FUMAGALLI

Bocconi University
Department of Economics