Threat Modeling Insider – September 2026

Threat Modeling Insider Newsletter

55th Edition – September 2026

Welcome!

Welcome to this month’s edition of Threat Modeling Insider!

In this issue, Marco Morana from Threat Modeling Academy, explores the question of how we can objectively measure the quality of AI generated threat models.

Next, on the Toreon Blog, Asma Oualmakran continues on her previous blog, showcasing how threat modeling fits within the CRA.

There’s plenty of other actionable insight ahead, so settle in and let’s get started!

Threat Modeling Insider edition

Welcome!

Welcome to this month’s edition of Threat Modeling Insider!

In this issue, Marco Morana from Threat Modeling Academy, explores the question of how we can objectively measure the quality of AI generated threat models.

Next, on the Toreon Blog, Asma Oualmakran continues on her previous blog, showcasing how threat modeling fits within the CRA.

There’s plenty of other actionable insight ahead, so settle in and let’s get started!

In this edition

Guest Article
From AI-Generated to Defensible: Measuring the Quality of AI-Powered Threat Models

Toreon Blog
Threat Modeling for CRA Compliance: The minimal Viable Model

Curated content
Models are getting better at Threat Modeling…

Tips & tricks
Treat threat modeling as a journey of understanding, not a one-time snapshot

Training update
An update on our training sessions.

Guest Article

From AI-Generated to Defensible:

Measuring the Quality of AI-Powered Threat Models

Founder, Threat Modeling Academy | Field CISO | Author & Instructor

Abstract

LLM-based and increasingly agentic threat-modeling tools can now augment significant portions of threat modeling from analyzing architectures and identifying assets to generating Data Flow Diagrams (DFDs), threats, attack scenarios, attack trees, adversarial tests, and mitigations. As these capabilities become more sophisticated and autonomous, an equally important challenge emerges: how do we objectively measure the quality of the threat models they produce?

The number of threats generated is not, by itself, an indication of quality. Nor is the percentage that results in remediation, because threat validity, risk acceptance, and remediation are different decisions. More importantly, these measures tell us little about relevant threats the AI failed to identify.  I approach this challenge through scope and input validation, expert-established ground truth, True Positive/False Positive/False Negative classification, precision, recall, F1, and Human-in-the-Loop validation. The objective is not simply faster threat modeling, but continuously measuring and improving its completeness, correctness, and practical usefulness.

The Evaluation Challenge

Consider an AI-powered threat-modeling process that generates 100 threats, but only five are ultimately prioritized for additional risk treatment, such as implementing new security controls or making architectural design changes. Looking only at this outcome, it may be tempting to conclude that the remaining 95 threats represent noise because their associated risks were not considered significant enough, based on likelihood and impact, to require further action.

But this conclusion confuses threat-model quality with risk-treatment decisions. The fact that a threat does not require additional remediation does not mean that the threat was incorrectly identified, irrelevant, or a False Positive. A generated threat may be valid and applicable to the architecture, while existing controls already reduce its associated risk to an acceptable level.

Before risk-treatment decisions can tell us anything useful, the AI-generated threats must first be independently validated for accuracy, relevance, and coverage. Only then can we distinguish between a valid threat that does not require further treatment and an invalid or irrelevant threat that represents actual noise

Measuring Threat Model Quality

Answering two fundamental questions (1) how many AI-generated threats are valid and relevant, and (2) how many relevant threats the AI failed to identify or incorrectly generated because of hallucinated, or out-of-context assumptions requires a trusted reference against which the AI-generated results can be evaluated.

In my training [1], I refer to this as establishing first and foremost a threat model “ground truth”, with an expert kept Human-in-the-Loop (HIL) that is to create a baseline to validate the AI-generated threat model against the actual architecture, defined scope, and expected threat landscape.

As illustrated in Figure 1, AI powered threat modeling can follow a workflow of evaluation that combines quantitative measurement with expert validation.

Screenshot 2026 09 30 at 14.24.53
Figure 1 – AI-Powered Threat Model Evaluation Workflow

Validate the Context Before Measuring the Threats

After the LLM generates the first threat model, each generated threat is compared individually against the expert-established ground truth to determine whether it is relevant, technically plausible, and applicable to the architecture and scope being modeled.

Before comparing the LLM-generated threat model against the established ground truth, however, we must first validate the inputs provided to the AI threat-modeling tool. Inaccurate, incomplete, or out-of-context inputs can produce equally inaccurate or incomplete outputs, a classic example of the Garbage-In, Garbage-Out (GIGO) principle.

Missing architectural details such as APIs, authorization flows, data stores, RAG pipelines, AI agents, MCP servers, or tool integrations can directly lead to missed threats. If an architectural element is missing from the LLM’s context, the threats associated with it may be missing from the threat model as well [2]. These omissions frequently lead to False Negatives, where legitimate threats remain unidentified simply because the corresponding architectural elements were never presented to the model.

Measuring Threat Model Accuracy and Coverage

Once the threat-model ground truth has been established and the initial AI-powered threat-model context and baseline have been validated, the AI-generated threats can be evaluated quantitatively. One practical approach is to classify each threat through Human-in-the-Loop review as a True Positive (TP) when it is correctly identified and applicable to the architecture and scope, or as a False Positive (FP) when it is irrelevant, out of scope, or based on incorrect or hallucinated context.

Equally important are False Negatives (FN): relevant threats that the AI failed to identify or adequately consider. These missed threats can be identified by comparing the AI-generated threat model against the established ground truth, including applicable threat-actor tactics, techniques, and procedures (TTPs), knowledge bases such as MITRE ATT&CK and ATLAS, threat intelligence, known incidents, and lessons learned relevant to the scope of the threat model.

Together, TP, FP, and FN provide a quantitative foundation for evaluating both the accuracy and coverage of AI-generated threats. These classifications can then be used to calculate three key measures:

  • Precision = TP / (TP + FP) measures the proportion of AI-generated threats that were validated as relevant and applicable.
  • Recall = TP / (TP + FN) measures how effectively the AI identified the relevant threats represented in the established ground truth.
  • F1 Score = 2 × (Precision × Recall) / (Precision + Recall) provides a balanced measure of overall performance by considering both accuracy and coverage.

These metrics provide a stronger basis for measuring threat-model quality than simply counting findings or remediation actions.

Figure 2 illustrates how the measurements of precision, recall, and F1 can expose these differences. An LLM may identify one threat category with high accuracy while providing weaker coverage for another. High precision with lower recall can indicate relevant results but missed threats; higher recall with lower precision can indicate broader coverage accompanied by more false positives. A higher F1 Score provides an overall measure of threat-model performance, indicating that the AI is achieving a stronger balance between accurately identifying relevant threats and providing sufficient coverage of the established threat landscape.

Screenshot 2026 09 30 at 14.26.56
Figure 2 – Measuring LLM Threat Modeling Performance on OWASP LLM T10

Consider the following evaluation scenario:

  • Ground Truth: 28 threats
  • LLM Generated: 30 threats
  • Correctly Identified: 26
  • Incorrectly Identified: 4
  • Missed Threats: 2

From these values:

  • Recall = 26 / (26 + 2) = 92.9%
  • Precision = 26 / (26 + 4) = 86.7%
  • F1 Score = 89.7%

These results indicate that the LLM provides excellent threat coverage while generating a manageable number of false positives. The evaluation also highlights opportunities for improving prompt engineering or architectural context to further increase precision.

Continuous Evaluation Across the LLM Threat Modeling Workflow

Evaluation cannot wait until the threat model is complete. With LLM-augmented threat modeling, validation can occur after each significant prompt interaction before its output becomes context for the next session. When the LLM extracts assets, the security architect validates their accuracy, completeness, and scope. When it generates or consumes a DFD, components, actors, flows, and trust boundaries are validated. Generated threats are subsequently validated for architectural relevance and accuracy, providing opportunities to identify TP, FP, and FN as the threat model develops.

The process follows a continuous sequence in which each prompt produces an LLM output that is reviewed and validated by the security architect. Based on that validation, the practitioner can refine the inputs, add missing architectural or security context, improve the DFD, and correct inaccurate assumptions. Once validated, the results are saved within the session and become trusted context for the next cascading prompt. In this way, each subsequent interaction builds upon progressively validated and enriched information, improving the quality of the threat model as it evolves.

This corresponds to the “Improve Prompt / Context / DFD” feedback loop in Figure 1. Each session preserves validated knowledge from previous interactions, allowing subsequent prompts to build on progressively improved context rather than starting from scratch. It also prevents errors from cascading. A hallucinated asset, missing API, incorrect trust boundary, or misunderstood authorization flow introduced early can otherwise propagate into irrelevant threats, incorrect attack paths, inappropriate controls, or missed risks.

This philosophy is why at the Threat Modeling Academy (TMA) we developed the TMA AI-Powered Threat Modeling Playground [3] as a training assistance environment. Practitioners progressively build threat models through session-based, cascading prompts while reviewing, challenging, correcting, and validating AI-generated outputs.

The objective is not simply to become better at prompting an LLM. It is to use AI augmentation to become better at threat modeling and produce a defensible threat model whose assumptions, architecture, threats, attack scenarios, control gaps, and risk conclusions can withstand expert peer review and security leadership scrutiny.

Making AI Powered Generated Threat Models Defensible

The challenge TODAY is no longer whether LLM-augmented or agentic based threat modeling tools can generate threat models quickly and automatically. They can. This challenge becomes even more important as systems continuously evolve. Architectures change, new integrations are introduced, trust boundaries shift, new attack techniques emerge, and security controls are modified. As these changes occur, the threat model and the assumptions on which it was built must be revisited. AI-augmented threat modeling can help practitioners keep pace with this rate of change by accelerating the reassessment and evolution of the threat model.

But greater automation does not reduce the need for human validation. As LLM and agentic threat-modeling tools become more capable and autonomous, Human-in-the-Loop evaluation becomes more important, not less. Continuous validation provides the evidence needed to ensure that an evolving AI-augmented threat model remains accurate, relevant, and defensible. Accepting an AI-generated threat model out of the box, without expert validation of its architecture, assumptions, threats, and coverage, risks creating false confidence. If inaccuracies, hallucinations, or missed threats are later exposed during peer review, security architecture review, or leadership scrutiny, confidence in both the threat model and the AI-assisted process can quickly erode.

A defensible threat model therefore requires evidence of what the AI identified correctly and, critically, what it missed. Human-in-the-Loop evaluation, supported by ground truth and measurable indicators such as precision and recall, provides a way to establish that evidence rather than simply trusting the generated output.

References

  1. Threat Modeling Academy https://threatmodeling.academy/ “AI-Powered Threat Modeling: Modernizing Security Analysis with LLM Augmentation”
  2. Avocado Systems https://www.avocadosys.com/gigochallenge/ “The GIGO Challenge in AI-Based Threat Modeling—and How to Solve It.”
  3. Threat Modeling Playground. https://demo.esadecimale.it/  “An iterative, session-based approach in which security practitioners progressively build, validate, and refine AI-assisted threat models, including Human-in-the-Loop validation and identification of false positives and false negatives.”

Learn to integrate AI into your threat modeling process.

Handpicked for you

Threat Modeling for CRA Compliance: The minimal Viable Model

In Threat modeling as a strategic path to CRA compliance we explained why threat modeling fits the Cyber Resilience Act so well. It covers the risk assessment that Article 13(2) asks for, it moves security into design instead of bolting it on before release, and it produces a good part of the technical documentation you need anyway.

This post is about the how. What does a threat model need to contain to stand up as CRA evidence? How do you map the countermeasures you pick to the essential requirements in Annex I? And what do you do with the gaps and residual risk that remain? We walk through a worked example at the end so you can reuse the approach on your own product.


Curated Content

Models are getting better at Threat Modeling…

Qwen 3.8 27B achieved a record 84.6 score on TM-Bench, making it the strongest open-weight model tested for STRIDE threat modeling on RTX 4090-class hardware. The benchmark evaluates threat coverage, completeness, technical validity, and JSON structure across 30 application designs, with Claude Sonnet 4.6 providing the grading. Despite keeping the same 27B parameter count and model family as Qwen 3.6, the new version improved by 18 points in roughly three months. Completeness remains the key differentiator, with Qwen 3.8 identifying 87% of ground-truth threats.

Threat modeling is not just for enterprise

Threat modeling is not just an enterprise compliance activity; it is a structured way to ask what you are building, what can go wrong, what to do about it, and whether you did enough. The article explains that methods like STRIDE, PASTA, LINDDUN, and MITRE ATT&CK are different layers or lenses for applying that thinking, with STRIDE often being the most practical starting point for smaller systems. Using a public GitHub/Terraform/Cloudflare setup as an example, it shows how threat modeling can lead to concrete improvements like pinning GitHub Actions, tightening token permissions, enabling secret scanning, and documenting why existing controls are sufficient.

Handpicked for you

Threat Modeling for CRA Compliance: The minimal Viable Model

In Threat modeling as a strategic path to CRA compliance we explained why threat modeling fits the Cyber Resilience Act so well. It covers the risk assessment that Article 13(2) asks for, it moves security into design instead of bolting it on before release, and it produces a good part of the technical documentation you need anyway.

This post is about the how. What does a threat model need to contain to stand up as CRA evidence? How do you map the countermeasures you pick to the essential requirements in Annex I? And what do you do with the gaps and residual risk that remain? We walk through a worked example at the end so you can reuse the approach on your own product.


Curated Content

Models are getting better at Threat Modeling…

Qwen 3.8 27B achieved a record 84.6 score on TM-Bench, making it the strongest open-weight model tested for STRIDE threat modeling on RTX 4090-class hardware. The benchmark evaluates threat coverage, completeness, technical validity, and JSON structure across 30 application designs, with Claude Sonnet 4.6 providing the grading. Despite keeping the same 27B parameter count and model family as Qwen 3.6, the new version improved by 18 points in roughly three months. Completeness remains the key differentiator, with Qwen 3.8 identifying 87% of ground-truth threats.

Threat modeling is not just for enterprise

Threat modeling is not just an enterprise compliance activity; it is a structured way to ask what you are building, what can go wrong, what to do about it, and whether you did enough. The article explains that methods like STRIDE, PASTA, LINDDUN, and MITRE ATT&CK are different layers or lenses for applying that thinking, with STRIDE often being the most practical starting point for smaller systems. Using a public GitHub/Terraform/Cloudflare setup as an example, it shows how threat modeling can lead to concrete improvements like pinning GitHub Actions, tightening token permissions, enabling secret scanning, and documenting why existing controls are sufficient.

TIPS & TRICKS

Treat threat modeling as a journey of understanding, not a one-time snapshot

The Manifesto explicitly values this over treating it as a point-in-time compliance artifact.

Book a seat in our upcoming trainings & events

Our trainings & events for 2026

3-Day Training: AI Whiteboard Hacking aka Hands-on Threat Modeling Training, in-person, OWASP Global AppSec EU, Vienna Austria

22-24 June 2026

Threat Modeling Practitioner training, hybrid online, hosted by DPI, US Cohort

June 2026

AI Whiteboard Hacking aka Hands-on Threat Modeling Training, TROOPERS, Heidelberg

22-23 June 2026

2 Day Training: Beyond Whiteboard Hacking: Embracing AI-Assisted Threat Modeling, in-person, OWASP Global AppSec USA, San Francisco

3-4 November 2026

1-day Workshop “Old School Threat Modeling”, hosted by OWASP BeNeLux, Utrecht NL

27 November 2026

Book a seat in our upcoming trainings & events

Our trainings & events for 2026

3-Day Training: AI Whiteboard Hacking aka Hands-on Threat Modeling Training, in-person, OWASP Global AppSec EU, Vienna Austria

22-24 June 2026

Threat Modeling Practitioner training, hybrid online, hosted by DPI, US Cohort

June 2026

AI Whiteboard Hacking aka Hands-on Threat Modeling Training, TROOPERS, Heidelberg

22-23 June 2026

2 Day Training: Beyond Whiteboard Hacking: Embracing AI-Assisted Threat Modeling, in-person, OWASP Global AppSec USA, San Francisco

3-4 November 2026

1-day Workshop “Old School Threat Modeling”, hosted by OWASP BeNeLux, Utrecht NL

27 November 2026

Threat Modeling Practitioner training, hybrid online, hosted by DPI, Europe Cohort

September 2026

Upcoming Events/Webinars

Webinar – Threat Modeling in the Age of Mythos:
What Threat Modelers Need to Do Differently

June 30 | 8 AM – 8:45 AM (EDT), Webinar together with QA

Join Toreon and our partner QA for a timely discussion on how AI-enabled security capabilities, including Anthropic’s Mythos, are changing the way organizations approach application security, product security, and secure-by-design practices.

QA Webinar Thumbnail 1

Conference – ThreatModCon Vienna
26 – 27 June | Meliá Vienna

Gather with the world’s leading security architects and engineers in the heart of Vienna. Join us for a full day of uncompromising, deeply technical focus.

Start typing and press Enter to search

Shopping Cart