A randomized experiment at Bocconi University showed that students using GPT-4o earned significantly better grades on a business assignment compared to those without access. The study involved 13 sections of an introductory management course split into four groups: control, causal reasoning lesson, GPT-4o access, or both. Students were tasked with writing marketing recommendations for the university's merchandise shop within 180 words. Those using GPT-4o scored nearly a full point higher on a 1-to-5 scale, with more coherent arguments and closer alignment to expert recommendations. The authors attribute this to higher content quality rather than increased student knowledge.

The causal reasoning lesson did not improve traditional scores but encouraged more diverse and unusual solutions. Students who received the lesson scored slightly worse on average but provided more detailed explanations of how their proposals would work and under what conditions they might fail. They also generated more varied ideas that diverged from their peers. Combining the lesson with GPT-4o did not further boost traditional scores, but the two approaches complemented each other in enhancing causal reasoning and idea diversity.

The grading rubric favored conventional answers, with more ideas, coherent arguments, and idea diversity correlating with higher scores. However, stronger falsifiability and divergence from other students' ideas were associated with lower scores. The authors suggest that grading systems need to explicitly reward originality and reasoning if these aspects are to be valued. They also note that the study does not show whether AI use improved actual learning, as there was no follow-up test without AI assistance.

Source: thedecoder