Do Cognitive Classifications Predict Empirical Item Behaviour? Aligning Bloom’s Taxonomy with the Psychometric Properties of National Centralized Mathematics Items across Grades 4–12, with an AI-Assisted Classification Benchmark
Rafat Y M Altaweel, Moustafa Kamal Moussa
Source abstract
Centralized testing systems routinely build test blueprints on the assumption that an item’s targeted cognitive level, coded with Bloom’s revised taxonomy, is monotonically related to its empirical difficulty and discrimination: lower-order items should be easier, and higher-order items harder and more discriminating. This assumption is rarely tested empirically in K–12 national assessment, and almost never at scale in mathematics. Using item-analysis records for 265 four-option multiple-choice items drawn from centralized mathematics examinations (End-of-Term 1, 2025 2026) across Grades 4–12 and both the General and Advanced tracks, this study examined (a) the alignment between the expert human cognitive classification and the empirical psychometric behaviour of items, and (b) the agreement between the expert classification and an independent large-language-model (LLM) classification produced by Claude (Anthropic). Difficulty (p-value) differed significantly across cognitive bands (Kruskal–Wallis H = 11.26, p = .004), with higher-order items being the hardest; the effect, however, was weak (ε² = .043; Spearman ρ = −.16), and a Jonckheere Terpstra test confirmed a significant decreasing trend across the ordered bands (J = 7943, z = −2.56, p = .006). Discrimination did not differ across cognitive bands (H = 0.54, p = .77), and difficulty and discrimination were essentially uncorrelated (ρ = −.05). Human–AI agreement was only fair (Cohen’s κ = .22 at the six-level resolution and κ = .17 at the three-band level; raw agreement was 51.7% and 58.1%, respectively), with the model systematically over-assigning the “apply/analyse” band. No significant item-level difference emerged between the General and Advanced tracks. The findings offer partial, weak empirical support for the cognitive-difficulty assumption, no support for a cognitive–discrimination link, and a clear caution against replacing expert human coders with current LLMs in high-stakes item classification. Implications for blueprint validation, item-writer training, and AI-assisted assessment workflows are discussed.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.