OpenAI experiments have shown that when students use AI, they produce logically clear answers that are close to those of experts.



The debate over whether students should be allowed to use generative AI or whether it should be uniformly prohibited is ongoing worldwide. A new experiment conducted by OpenAI in collaboration with universities has revealed that there are certain benefits to using AI.

Training novices to think, or giving them LLMs? Evidence from an RCT
(PDF file)

https://cdn.openai.com/pdf/novices-and-llm-august-2026.pdf

Better answers, broader thinking: What students gain from ChatGPT and critical-thinking training | OpenAI
https://openai.com/index/what-students-gain-from-chatgpt-critical-thinking-training/



Researchers at Bocconi University in Italy assigned more than 1,000 first-year students the task of proposing marketing strategies for a store selling university merchandise. The students were randomly divided into four groups by class and given one of the following options: 'access to ChatGPT (GPT-4o),' 'causal inference training,' 'both of the above,' or 'none of the above.'

'Causal inference training' is training to connect causes and effects and explain why a particular solution works, or why it might not work. AI was not used in this training; the concept of causal inference was taught through games and feedback.



Student submissions were evaluated by humans on a 5-point scale. Separately, researchers performed automated text analysis to measure the number and diversity of ideas, the effectiveness of causal inference, and the similarity to model answers from three experts.

As a result, the group that used ChatGPT received an evaluation score that was almost one level higher than the other groups. This group's submissions contained more ideas, had clearer logic, and were more similar to the model answer. It should be noted that the students did not simply hand over the assignment to ChatGPT; they themselves decided what questions to ask, evaluated the answers, and ultimately decided what to include in their submissions.

On the other hand, the group that had received training in causal inference was able to clearly explain why their ideas would work and when they might fail. However, this sophistication of their answers did not match the evaluation criteria for this assignment, so they did not score as highly on the 5-point scale as the ChatGPT group.

The results of the automated text analysis also showed that the group trained in causal inference performed better in terms of originality of ideas.

OpenAI stated, 'The traditional evaluation criteria used in this experiment value clear and logical answers, but overlook whether the ideas were original and not something anyone else had thought of. In discussions surrounding education and AI, it's easy to frame it as a binary choice: should students learn to think for themselves, or should they learn how to use AI? However, this experiment demonstrates that there are different benefits to using AI versus not using it.'



The group that received both ChatGPT and causal inference training showed overall better trends than the two groups mentioned above. This group performed almost as well as the ChatGPT group in terms of the 5-point rating scale and the number of ideas, and was on par with the group that received only causal inference training in terms of originality of ideas.

OpenAI pointed out that 'if AI alone can produce the same results as experts, then traditional assessment methods make it difficult to understand what students actually understand. These results demonstrate the importance of evaluating originality, reasoning, and different approaches, rather than just formulaic or logical answers.'

in AI, Posted by log1p_kr