Evaluation cannot wait until the threat model is complete. With LLM-augmented threat modeling, validation can occur after each significant prompt interaction before its output becomes context for the next session. When the LLM extracts assets, the security architect validates their accuracy, completeness, and scope. When it generates or consumes a DFD, components, actors, flows, and trust boundaries are validated. Generated threats are subsequently validated for architectural relevance and accuracy, providing opportunities to identify TP, FP, and FN as the threat model develops.
The process follows a continuous sequence in which each prompt produces an LLM output that is reviewed and validated by the security architect. Based on that validation, the practitioner can refine the inputs, add missing architectural or security context, improve the DFD, and correct inaccurate assumptions. Once validated, the results are saved within the session and become trusted context for the next cascading prompt. In this way, each subsequent interaction builds upon progressively validated and enriched information, improving the quality of the threat model as it evolves.
This corresponds to the “Improve Prompt / Context / DFD” feedback loop in Figure 1. Each session preserves validated knowledge from previous interactions, allowing subsequent prompts to build on progressively improved context rather than starting from scratch. It also prevents errors from cascading. A hallucinated asset, missing API, incorrect trust boundary, or misunderstood authorization flow introduced early can otherwise propagate into irrelevant threats, incorrect attack paths, inappropriate controls, or missed risks.
This philosophy is why at the Threat Modeling Academy (TMA) we developed the TMA AI-Powered Threat Modeling Playground [3] as a training assistance environment. Practitioners progressively build threat models through session-based, cascading prompts while reviewing, challenging, correcting, and validating AI-generated outputs.
The objective is not simply to become better at prompting an LLM. It is to use AI augmentation to become better at threat modeling and produce a defensible threat model whose assumptions, architecture, threats, attack scenarios, control gaps, and risk conclusions can withstand expert peer review and security leadership scrutiny.