Bloom’s Taxonomy plays a central role in assessment design by helping instructors align evaluation tasks with learning objectives. However, applying Bloom’s framework in practice, especially in programming education, requires substantial effort and often leads to divergent interpretations among educators. This study explores the extent to which Large Language Models (LLMs) can support the automated classification of programming assessment items across Bloom’s cognitive levels. We evaluate Bloom-based classification on a dataset of items from an introductory undergraduate Computer Science course, covering four cognitive levels (Remember, Understand, Apply, and Analyze), comparing proprietary LLMs with open-source alternatives. Our methodology considers two strategies: zero-shot prompting and a council-based ensemble approach. Results show that proprietary models achieve accuracies of up to 78%, while locally deployable models reach 59–73%, within the range of human inter-annotator variability. By lowering the technical and economic barriers to adoption, our approach can assist educators in designing balanced assessments. Dataset and framework are available at: https://doi.org/10.5281/zenodo.19332485.

Ferrato, A., Limongelli, C., Schicchi, D., Taibi, D. (2026). Large Language Models for Automated Bloom’s Taxonomy Classification in Computer Science Assessment. In Lecture Notes in Computer Science (pp.262-270). GEWERBESTRASSE 11, CHAM, CH-6330, SWITZERLAND : Springer Science and Business Media Deutschland GmbH [10.1007/978-3-032-29760-0_29].

Large Language Models for Automated Bloom’s Taxonomy Classification in Computer Science Assessment

Ferrato A.
;
Limongelli C.;Schicchi D.;Taibi D.
2026-01-01

Abstract

Bloom’s Taxonomy plays a central role in assessment design by helping instructors align evaluation tasks with learning objectives. However, applying Bloom’s framework in practice, especially in programming education, requires substantial effort and often leads to divergent interpretations among educators. This study explores the extent to which Large Language Models (LLMs) can support the automated classification of programming assessment items across Bloom’s cognitive levels. We evaluate Bloom-based classification on a dataset of items from an introductory undergraduate Computer Science course, covering four cognitive levels (Remember, Understand, Apply, and Analyze), comparing proprietary LLMs with open-source alternatives. Our methodology considers two strategies: zero-shot prompting and a council-based ensemble approach. Results show that proprietary models achieve accuracies of up to 78%, while locally deployable models reach 59–73%, within the range of human inter-annotator variability. By lowering the technical and economic barriers to adoption, our approach can assist educators in designing balanced assessments. Dataset and framework are available at: https://doi.org/10.5281/zenodo.19332485.
2026
9783032297594
Ferrato, A., Limongelli, C., Schicchi, D., Taibi, D. (2026). Large Language Models for Automated Bloom’s Taxonomy Classification in Computer Science Assessment. In Lecture Notes in Computer Science (pp.262-270). GEWERBESTRASSE 11, CHAM, CH-6330, SWITZERLAND : Springer Science and Business Media Deutschland GmbH [10.1007/978-3-032-29760-0_29].
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11590/559483
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? 0
social impact