Home
Scholarly Works
Assessing the Quality and Diversity of Student-...
Conference

Assessing the Quality and Diversity of Student- and AI-Generated Programming Exercises from Limited Examples: Empirical Insights from a Large-Scale Blind Evaluation

Abstract

As online learning platforms increasingly deliver personalized learning experiences, the demand for high-quality educational content grows. Creating such content presents a growing burden for instructors. Leveraging students for content co-creation (learnersourcing) and the use of large language models (LLMs) are two scalable and promising approaches for addressing this content creation challenge. We compare these approaches in a large introductory programming course (N=950), where students and an LLM independently created ‘code examples’ from the same limited set of exemplars. In a blind evaluation, students rated the correctness and helpfulness of both sets of resources. AI-generated examples were rated comparable in quality to student-generated examples, but more closely resembled the provided exemplars. Student-generated examples showed greater variation in length and syntax. This pattern is consistent with concerns about the homogenizing effects of LLMs. Overall, both learnersourcing and AI-based generation appear viable for producing supplementary programming materials, with indications that LLMs offer more consistency while students contribute greater diversity.

Authors

Denny P; Leinonen J; Hellas A; Sarsa S; Liut M; Khosravi H

Pagination

pp. 1-7

Publisher

Association for Computing Machinery (ACM)

Publication Date

April 30, 2026

DOI

10.1145/3828792.3828801

Name of conference

Proceedings of the 28th Western Canadian Conference on Computing Education

Labels

Sustainable Development Goals (SDG)

View published work (Non-McMaster Users)