Niels Doorn
(Co-)promotors: prof.dr. T. Vos (OU), prof.dr E. Barendsen (OU/RU), dr. B.M. Marín (Universitat Politècnica de València) and dr. M.R. van Diggelen (NHL Stenden Hogeschool)
Open Universiteit
Date: 13 November 2026
Thesis: PDF will follow
Summary
Software testing is one of the most effective techniques to ensure software quality, yet in computer science education it remains under-represented and difficult to teach. Students often graduate without the ability to design and evaluate tests effectively, a problem rooted in the cognitive complexity of testing and the lack of established didactic approaches. This gap between education and industry is widely recognised and calls for research into how learners and experts make sense of testing activities. This skill gap creates a disconnect between academia and industry: graduates struggle with test design, exploratory practices, and critical evaluation of software, all of which are highly valued in practice. Understanding and improving how students and experts use sensemaking during tests is the key to bridge this gap. The ultimate solution would be an educational ecosystem where software testing is no longer a separate, neglected topic but a natural and integrated part of computer science learning from the very beginning. In this vision, students do not see testing as extra work or an afterthought, but as a core activity of programming and software engineering. Instead of approaching testing with a purely rationalist, code-driven mindset, they learn to balance rationalism with empiricism: treating tests as small scientific experiments in which hypotheses are formulated, explored, and evaluated against evidence. Such a solution would combine several complementary elements. Programming education would routinely use approaches like Test Informed Learning with Examples, where testing is embedded into every exercise and example. Students would participate in puzzle-based learning activities that build their skills in questioning, modelling, experimenting, and communicating about tests. Serious games would provide immersive, hands-on opportunities to practice exploratory testing strategies, scaffolded by testing tours and Socratic questioning, helping them experience the reflective and investigative side of the discipline. At the advanced level, students would be exposed to expert reasoning patterns through simulations and role-play, learning to align their approaches with the practices of professionals in industry. Ultimately, this integrated framework would form “AI-resilient” testing professionals: graduates who can critically evaluate, guide, and improve software quality even in a future where machine code is increasingly generated by AI. They would not only be technically skilled, but also capable of making sense of complex systems, adapting to uncertainty, and applying critical thinking to software behaviour. In doing so, the persistent gap between academia and industry would be closed, and testing would assume its rightful place at the heart of computer science education. This thesis addresses the problem by investigating sensemaking in software testing and exploring how empiricism, exploration, and reflection-in-action can be better supported in education. Through a series of empirical studies and design-based research interventions, the work develops new ways of embedding testing into current curricula. Test Informed Learning with Examples shows how testing can be introduced seamlessly and subtly in programming courses. Puzzle-based learning is used to engage students in questioning assumptions, modelling, experimenting, and communicating about tests. A serious game design demonstrates how exploratory testing can be scaffolded with testing tours and Socratic questioning to foster reflective practice. This thesis integrates a series of empirical investigations, design-based research interventions, and conceptual frameworks. Each contribution is anchored in a published or submitted study, collectively articulating a coherent narrative on how to improve testing education through sensemaking and empiricism. Collectively, these studies contribute to an expanding corpus of knowledge that transitions the pedagogy of testing from a rationalist, code-centric paradigm toward an empiricist approach, founded on experimentation and evidence. The thesis concludes by outlining future work on expanding serious games and puzzle collections, further refining pedagogical interventions, and continuing to bridge the persistent gap between academic teaching and the demands of professional software testing practice.