Large Language Models in Mathematics Education: Practitioners as Pedagogical Prompt Engineers

Authors

  • Sima Caspari-Sadeghi Østfold University of Applied Sciences

DOI:

https://doi.org/10.33423/7csntx42

Keywords:

higher education, critical AI Literacy, large language models, mathematics education, pedagogical prompt engineering

Abstract

Effective classroom integration of large language models (LLMs) requires domain-specific Prompt Engineering (PE) competence among teachers and students. Existing PE literature remains largely technical or cross-disciplinary, offering limited guidance for mathematics education, where LLMs continue to struggle with mathematical reasoning. This paper proposes a pedagogically grounded framework that positions educators as pedagogical prompt engineers who scaffold students’ interactions with LLMs through discipline-specific prompting practices. The framework provides structured classroom exemplars to enhance mathematical reasoning, conceptual understanding, and critical evaluation of AI-generated outputs. By fostering critical AI literacy and human-centered inquiry, it supports deeper, reflective, and ethical engagement with intelligent technologies.

References

Ahn, J., Verma, R., Lou, R., Liu, D., Zhang, R., & Yin, W. (2024). Large language models for mathematical reasoning: Progresses and challenges. arXiv preprint arXiv:2402.00157. DOI: https://doi.org/10.18653/v1/2024.eacl-srw.17

AI for Education. (2024). Prompt framework for educators: The Five "S" model. Retrieved from https://www.aiforeducation.io/ai-resources/the-five-s-model

Bandura, A., & Jeffrey, R.W. (1973). Role of symbolic coding and rehearsal processes in observational learning. Journal of Personality and Social Psychology, 26(1), 122–130. DOI: https://doi.org/10.1037/h0034205

Boonstra, L. (2025). Prompt engineering [White paper]. Innopreneur. Retrieved from https://www.innopreneur.io/wp-content/uploads/2025/04/22365_3_Prompt-Engineering_v7-1.pdf

Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.

Caspari-Sadeghi, S. (2026a). Adaptive learning systems: Leveraging generative artificial intelligence for optimal adaptivity. In B. Wasson & C.E. Tømte (Eds.), Handbook of learning analytics and adaptive learning in compulsory education (pp. 21–34). Edward Elgar Publishing. https://doi.org/10.4337/9781035330676.00009 DOI: https://doi.org/10.4337/9781035330676.00009

Caspari-Sadeghi, S. (2026b). AI literacy for teacher educators: A holistic curriculum for capacity-building in higher education. Frontiers in Education. DOI: https://doi.org/10.3389/feduc.2026.1745768

Caspari-Sadeghi, S. (2025). Technology-enhanced and formative assessment (Unpublished habilitation thesis). University of Passau.

Cain, W. (2024). Prompting change: Exploring prompt engineering in large language model AI and its potential to transform education. TechTrends, 68, 47–57. DOI: https://doi.org/10.1007/s11528-023-00896-0

Correia, A.-P., Hickey, S., & Xu, F. (2025). Realizing the possibilities of large language models: Strategies for prompt engineering in educational inquiries. Theory Into Practice, 64(4), 434–447. DOI: https://doi.org/10.1080/00405841.2025.2528545

Dang, H., Mecke, L., Lehmann, F., Goller, S., & Buschek, D. (2022). How to prompt? Opportunities and challenges of zero- and few-shot learning for human-AI interaction in creative applications of generative models. arXiv preprint arXiv:2209.01390.

Eager, B., & Brunton, R. (2023). Prompting higher education towards AI-augmented teaching and learning practice. Journal of University Teaching and Learning Practice, 20(5). DOI: https://doi.org/10.53761/1.20.5.02

Federiakin, D., Molerov, D., Zlatkin-Troitschanskaia, O., & Maur, A. (2024). Prompt engineering as a new 21st century skill. Frontiers in Education, 9, 1366434. DOI: https://doi.org/10.3389/feduc.2024.1366434

von Garrel, J., & Mayer, J. (2023). Artificial intelligence in studies—Use of ChatGPT and AI-based tools among students in Germany. Humanities and Social Sciences Communications, 10, 799. DOI: https://doi.org/10.1057/s41599-023-02304-7

Kim, J., Lee, H., & Cho, Y.H. (2022). Learning design to support student-AI collaboration: Perspectives of leading teachers for AI in education. Education and Information Technologies, 27(5), 6069–6104. DOI: https://doi.org/10.1007/s10639-021-10831-6

Knoth, N., Tolzin, A., Janson, A., & Leimeister, J.M. (2024). AI literacy and its implications for prompt engineering strategies. Computers and Education: Artificial Intelligence, 6, 100225. DOI: https://doi.org/10.1016/j.caeai.2024.100225

Lo, L.S. (2023). The CLEAR path: A framework for enhancing information literacy through prompt engineering. The Journal of Academic Librarianship, 49(4), 102720. DOI: https://doi.org/10.1016/j.acalib.2023.102720

Miao, F., & Cukurova, M. (2024). AI competency framework for teachers. UNESCO.

Ng, D.T.K., Su, J., Leung, J.K.L., & Chu, S.K.W. (2023). Artificial intelligence (AI) literacy education in secondary schools: A review. Interactive Learning Environments, 1–21.

Park, J., & Choo, S. (2024). Generative AI prompt engineering for educators: Practical strategies. Journal of Special Education Technology, 40(3), 411–417. DOI: https://doi.org/10.1177/01626434241298954

Pepin, B., Buchholtz, N., & Salinas-Hernández, U. (2025). A scoping survey of ChatGPT in mathematics education. Digital Experiences in Mathematics Education, 11(1), 9–41. DOI: https://doi.org/10.1007/s40751-025-00172-1

Sahoo, P., Singh, A.K., Saha, S., Jain, V., Mondal, S.S., & Chadha, A. (2024). A systematic survey of prompt engineering in large language models: Techniques and applications. arXiv preprint arXiv:2402.07927.

Schoenfeld, A.H. (1992). Learning to think mathematically: Problem solving, metacognition, and sense-making in mathematics. In D.A. Grouws (Ed.), Handbook for research on mathematics teaching and learning (pp. 334–370). Macmillan. DOI: https://doi.org/10.1108/978-1-60752-874-620251019

Schorcht, S., Buchholtz, N., & Baumanns, L. (2024). Prompt the problem: Investigating the mathematics educational quality of AI-supported problem solving by comparing prompt techniques. Frontiers in Education, 9, 1386075. DOI: https://doi.org/10.3389/feduc.2024.1386075

Surowiecki, J. (2004). The wisdom of crowds: Why the many are smarter than the few and how collective wisdom shapes business, economies, societies, and nations. Doubleday.

Tassoti, S. (2025). Assessment of students' use of generative artificial intelligence: Prompting strategies and prompt engineering in chemistry education. Journal of Chemical Education, 101(6), 2475–2482. DOI: https://doi.org/10.1021/acs.jchemed.4c00212

Velásquez-Henao, J.D., Franco-Cardona, C.J., & Cadavid-Higuita, L. (2023). Prompt engineering: A methodology for optimizing interactions with AI-language models in the field of engineering. Dyna, 90(230), 9–17. DOI: https://doi.org/10.15446/dyna.v90n230.111700

Walter, Y. (2024). Embracing the future of artificial intelligence in the classroom: The relevance of AI literacy, prompt engineering, and critical thinking in modern education. International Journal of Educational Technology in Higher Education, 21, 1–29. DOI: https://doi.org/10.1186/s41239-024-00448-3

Wang, B., Rau, P.L.P., & Yuan, T. (2023a). Measuring user competence in using artificial intelligence: Validity and reliability of artificial intelligence literacy scale. Behaviour & Information Technology, 42(11), 1324–1337. DOI: https://doi.org/10.1080/0144929X.2022.2072768

Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., & Zhou, D. (2022). Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.

Wardat, Y., Tashtoush, M., Alali, R., & Jarrah, A. (2023). ChatGPT: A revolutionary tool for teaching and learning mathematics. Eurasia Journal of Mathematics, Science and Technology Education, 19(7), 1–18. DOI: https://doi.org/10.29333/ejmste/13272

Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35, 24824–24837. DOI: https://doi.org/10.52202/068431-1800

Xiao, Y., Hou, X., Ye, R., Kazemitabar, M., Diana, N., … Stamper, J. (2025). Improving student–AI interaction through pedagogical prompting: An example in computer science education. arXiv preprint arXiv:2506.19107.

Yen, A.-Z., & Hsu, W.L. (2023). Three questions concerning the use of large language models to facilitate mathematics learning. arXiv preprint arXiv:2310.13615. DOI: https://doi.org/10.18653/v1/2023.findings-emnlp.201

Zamfirescu-Pereira, J.D., Wong, R.Y., Hartmann, B., & Yang, Q. (2023). Why Johnny can't prompt: How non-AI experts try (and fail) to design LLM prompts. Proceedings of the CHI Conference on Human Factors in Computing Systems, 1–21. DOI: https://doi.org/10.1145/3544548.3581388

Zheng, H.S., Mishra, S., Chen, X., Cheng, H., Chi, E.H., Le, Q.V., & Zhou, D. (2023). Take a step back: Evoking reasoning via abstraction in large language models. arXiv preprint arXiv:2310.06117.

Downloads

Published

2026-09-12

Issue

Section

Articles

How to Cite

Caspari-Sadeghi, S. (2026). Large Language Models in Mathematics Education: Practitioners as Pedagogical Prompt Engineers. Journal of Higher Education Theory and Practice, 26(4). https://doi.org/10.33423/7csntx42