An Empirical Investigation Into Human Judgments of Security Threats

Winnie Bahati Mbaka

(Co-)promotors: prof.dr.ir. Fabio Massacci (VU, UniTrento), dr. Katja Tuma, (TU/e)
Vrije Universiteit Amsterdam
Date: 7 September 2026

Summary

In an increasingly interconnected digital landscape, security risk management is critical in ensuring that the confidentiality, integrity, and availability of critical assets such as data are maintained. A key component of this process is Threat Analysis and Risk Assessment (TARA), systematic methods for assessing security risks. To assist in these efforts of building secure systems, a myriad of threat analysis methodologies already exist, including STRIDE, privacy threat modeling with LIND(D)UN, PASTA, attack trees, approaches based on misuse cases from requirements engineering, and CORAS.

Past empirical research has focused its efforts on measuring the performance of threat modeling techniques using use cases evaluated by researchers or involving students and practitioners. However, these studies often evaluate the analyst’s ability to correctly identify threats using measures of success such as true positives or false positives. Yet, in some cases, substantial variations in the identified threats, given the same use case, have been reported. Such instances highlight the existence of underlying influences, such as human factors or contextual conditions, on the outcomes of threat analysis.

Although there has been widespread application of threat analysis and risk assessment methodologies, several persistent challenges still exist. First, TARA techniques are based on expert judgment, which is prone to cognitive, optimism, or confirmation biases, raising concerns about the reproducibility of the findings and the subjectivity of their quality. Second, the issue of threat explosion has been observed, especially during the application of TARA to complex systems models, compounded by the limited availability and varying quality of analysis materials used for threat validation. Lastly, there is a lack of a clearly defined “definition of done” during threat analysis activities, a gap that new technologies such as Large Language Models (LLMs) are increasingly being adopted to address, in an effort to expedite a rather time-consuming process.

This thesis investigates the STRIDE methodology as an example to explore these three challenges in controlled experimental settings. First, we investigate the reliability of human assessors in reviewing security outcomes in the presence of a defined assessment criterion by closely replicating the performance and precision of STRIDE variants for threat elicitation. Second, we designed a randomised factorial survey to investigate bias in evaluating a security case study with ethical implications, considering human factors such as gender, seniority, and level of education. Lastly, we investigated the effect of analysis materials, specifically data flow diagrams and sequence diagrams, alongside Large Language Models (as AI assistants), on threat validation.

In summary, our findings point to several key insights. The choice between STRIDE variants has no practical impact on performance, suggesting that organizational security needs should drive the selection of threat analysis techniques. Regarding human factors, we found no evidence that gender, seniority, or education level influenced the evaluation of security mitigations, indicating that analysts tend to assess proposed solutions on their merits. Finally, regarding analysis materials, we found that less is more. Textual materials, such as threat descriptions, were perceived as most useful, while additional graphical or AI-generated information did not improve the actual performance of threat validation.

Scroll to Top