Open-ended questions are effective for assessing students’ conceptual understanding and higher-order reasoning, but in large-scale educational contexts, such as university courses, their grading is time-consuming, which limits their practical use. Although recent work shows that large language models (LLMs) can automate grading under predefined or explicit criteria, failing to account for instructors’ implicit and individual grading standards risks overlooking a core component of authentic evaluation. This work addresses this gap by examining whether LLMs can infer and align with such latent criteria from a small number of graded examples. A few-shot grading setting is considered in which a teacher provides a set of scored student responses, and the model assigns scores consistent with the teacher’s evaluation scale. Several GPT-based models (GPT-4o, GPT-4.1, GPT-5, and their mini variants) are evaluated under few-shot conditions (k ∈ {1, 3, 5, 7} ) using student answers graded by a human instructor. Alignment with teacher grading is assessed using strict accuracy, relaxed accuracy with tolerance, mean absolute error, and quadratic weighted kappa, complemented by grading error analysis. Results show that all models benefit from few-shot examples, with gains emerging at low k and diminishing returns beyond approximately five examples. Higher-capacity models achieve stronger alignment, while smaller variants consistently exhibit lower performance. Remaining discrepancies are predominantly local, reflecting confusion between adjacent score levels rather than severe misgradings. These findings suggest that few-shot LLM-based grading can approximate a teacher’s evaluation criteria with limited instructor effort, providing scalable decision support for open-ended question evaluation.

Can Large Language Models Learn to Grade Like Teachers? A Few-Shot Study on Open-Ended Assessment

V. Scorza;G. Cassano;N. Di Blas
2026-01-01

Abstract

Open-ended questions are effective for assessing students’ conceptual understanding and higher-order reasoning, but in large-scale educational contexts, such as university courses, their grading is time-consuming, which limits their practical use. Although recent work shows that large language models (LLMs) can automate grading under predefined or explicit criteria, failing to account for instructors’ implicit and individual grading standards risks overlooking a core component of authentic evaluation. This work addresses this gap by examining whether LLMs can infer and align with such latent criteria from a small number of graded examples. A few-shot grading setting is considered in which a teacher provides a set of scored student responses, and the model assigns scores consistent with the teacher’s evaluation scale. Several GPT-based models (GPT-4o, GPT-4.1, GPT-5, and their mini variants) are evaluated under few-shot conditions (k ∈ {1, 3, 5, 7} ) using student answers graded by a human instructor. Alignment with teacher grading is assessed using strict accuracy, relaxed accuracy with tolerance, mean absolute error, and quadratic weighted kappa, complemented by grading error analysis. Results show that all models benefit from few-shot examples, with gains emerging at low k and diminishing returns beyond approximately five examples. Higher-capacity models achieve stronger alignment, while smaller variants consistently exhibit lower performance. Remaining discrepancies are predominantly local, reflecting confusion between adjacent score levels rather than severe misgradings. These findings suggest that few-shot LLM-based grading can approximate a teacher’s evaluation criteria with limited instructor effort, providing scalable decision support for open-ended question evaluation.
2026
Artificial Intelligence in Education. AIED 2026
978-3-032-29760-0
File in questo prodotto:
File Dimensione Formato  
AIED_short_paper_conference_2026 (1).pdf

embargo fino al 28/06/2027

: Post-Print (DRAFT o Author’s Accepted Manuscript-AAM)
Dimensione 310.77 kB
Formato Adobe PDF
310.77 kB Adobe PDF   Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11311/1324146
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact