Indexed metadata

AI-supported formative assessment of conceptual understanding in mathematics: LLM-based pre-coding of open student responses

Corinna Hankeln

Source record

Source: Crossref

Published: Sep 30, 2026

DOI: 10.21203/rs.3.rs-10200491/v1

Open original source ↗

Source abstract

Abstract Formative assessment of students' conceptual understanding in mathematics benefits from open-ended tasks that reveal deep insights into student thinking, yet the practical burden of analyzing such responses means that their diagnostic potential frequently goes unrealized. This paper investigates to what extent prompt-based large language models (LLMs) can reliably identify theory-derived categories in authentic student responses to open mathematical items without task-specific fine-tuning, and how performance varies across models and coding formats. Drawing on two datasets (13,082 codings from 1,763 students across 21 items (natural numbers and fractions), and 1,186 responses to a problem posing item coded with holistic and analytic rubric formats across four open-source models), LLM performances were evaluated against one expert and two non-expert human codings. Across contents, content-related categories achieve a mean weighted F1-score of 0.818. For the in-depth analysis of a problem posing item, subcategory analysis reveals three distinct difficulty patterns: subcategories requiring the recognition of something positively present are identified reliably both by humans and LLMs; those requiring arithmetic verification, recognition of an absent feature, or resolution of ambiguous linguistic signals prove hard for models and a less experienced rater, but are navigated reliably by a rater with deeper subject-matter expertise. Crucially, the subcategories operationalizing deep conceptual understanding of multiplication are among the most reliably identified features. We derive implications for prompt architecture, task design, and human-in-the-loop assessment, including the role of linguistically explicit task design and subject-matter-grounded prompt development, arguing for LLMs as catalysts within the formative assessment cycle rather than autonomous assessment agents.

Evidence graph

No public relationships recorded yet.

Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.

AI-supported formative assessment of conceptual understanding in mathematics: LLM-based pre-coding of open student responses — Mathematical Frontier Network