Toreon Office | Grotehondstraat 44 1/1 - 2018 Antwerpen | +32 3 369 33 96
Written by Jordan Hardy
By Asma Oualmakran, Sebastien Deleersnyder, Maxim Baele
In Threat modeling as a strategic path to CRA compliance we explained why threat modeling fits the Cyber Resilience Act so well. It covers the risk assessment that Article 13(2) asks for, it moves security into design instead of bolting it on before release, and it produces a good part of the technical documentation you need anyway.
This post is about the how. What does a threat model need to contain to stand up as CRA evidence? How do you map the countermeasures you pick to the essential requirements in Annex I? And what do you do with the gaps and residual risk that remain? We walk through a worked example at the end so you can reuse the approach on your own product.
While almost everyone agrees that threat modeling is useful, we do sometimes experience initial apprehension when teams hear it is coming their way. That hesitation usually comes from what people have lived through: three-day workshops, a tool licence, an outside facilitator and a forty-page report that nobody ever fully read through. If that is your picture of threat modeling, doing it for every product and every release sounds unrealistic.
We see two patterns that don’t work. The first is the compliance model. Someone writes it because a checklist asks for one, it lands in a folder, and the engineers never look at it again. The second is the exhaustive model. Every component is in it, every threat is listed, every score is debated in a meeting. It is out of date by the time it’s finished, and nobody volunteers to do it again.
Both sit outside the development process, so neither keeps pace with the product, and keeping pace is exactly what the CRA expects.
Four points in the regulation indicate how you should run the practice.
It has to be continuous. Article 13(2) says the outcome of the risk assessment must be taken into account during “the planning, design, development, production, delivery and maintenance phases” of the product. A single assessment before launch does not cover that, and the support period you commit to is normally at least five years (Article 13(8)).
It has to be written down. Article 13(3) and 13(4) require the assessment to be documented, updated as appropriate during the support period, and included in the technical documentation under Article 31 and Annex VII. What was said in a workshop does not count unless it ends up on paper with a date on it.
It has to take a position on every requirement. This is the part teams underestimate most. Article 13(2) requires the risk assessment to state whether, and if so how, each security requirement in Annex I, Part I, point 2 applies to your product, and how you implemented it. Where a requirement does not apply, Article 13(4) asks for a clear justification. So your threat model has to end in a stated position per requirement, not in a list of findings.
The supply chain runs alongside it. Article 13(5) requires due diligence on components you integrate from third parties. That produces different artefacts (SBOMs, vulnerability triage, supplier records) and is best handled as a parallel track that feeds the threat model, not as part of it.
The reporting obligations in Article 14 have applied since 11 September 2026 and the rest of the regulation applies from 11 December 2027, so the practice you set up now is the one you will be running once the full regulation is in force.
We are used to threat modeling from an information security point of view, where impact is rated against the company and its risk appetite: downtime, data exposure, reputation, cost. The CRA requires you to take the perspective of your product’s user instead. That way of thinking comes from product safety, the same tradition CE marking is built on, and anyone who has worked with the Machinery Directive or its successor, the Machinery Regulation, will recognise it. You look at the product’s intended purpose, its reasonably foreseeable use and the people who depend on it, not at what an incident costs you.
The threats you find barely change, but who gets hurt does: a denial of service on a telemetry endpoint can be a minor nuisance for you and a real problem for the person whose device stops working. So when you write your doomsday scenarios for a CRA threat model, leave your own business losses out and write them from the user’s side. Three to five scenarios is usually enough, and everything you rate later should point back to one of them.
We use the four-question framework, which we teach as DICE. The questions stay the same. What you do with each answer shifts a little.
What are we working on? This is the context of your threat model, and under the CRA it maps directly to things you have to describe anyway. Start from the product’s core functionality, which is also what determines whether your product falls into one of the important or critical categories. Add its intended purpose and reasonably foreseeable use, and to a lesser extent the foreseeable misuse, which feeds straight into the doomsday scenarios you write at this stage. All of it belongs in the general description of the product that Annex VII asks for in your technical documentation, so write it once and reference it from both places. Then draw the product as the CRA defines it: the components you control, their data flows and the trust boundaries between them. A backend that you design and develop yourself, or have developed under your responsibility, and that the product needs to perform one of its functions counts as remote data processing and is part of the product, so it belongs in scope. Third-party services and libraries appear as external entities only.
What can go wrong? Use whatever method your engineers already know. STRIDE and Attack Trees work fine, and the CRA itself does not prescribe a method, but STRIDE is the one that keeps coming back in CRA-related guidance, so if your team has no strong preference it is a safe place to start. Whatever you pick, every threat should lead to at least one user-harm scenario, and if it doesn’t, park it.
What are we going to do about it? Turn each countermeasure into a backlog item with an ID, an owner and acceptance criteria you can test. Tag it with the Annex I requirement it supports at the moment you create it. Doing it at that point is far cheaper than reconstructing the mapping at the end. A mitigation that only exists in the threat model document has not really been decided.
Did we do a good job? Within a single session, check that every threat has a decision (mitigate, accept, avoid or transfer) and that every mitigation has proof that it works, such as a test, a code review or a closed pentest finding. Then roll everything up per Annex I requirement, so that for each one you can see which countermeasures cover it, what evidence shows it, its status and what is left open.
The CRA also pushes this question up a level, to the process around the model. Threat modeling a product once is fairly easy. Doing it continuously and improving it as you go takes a lot more, and the CRA expects the model to stay current for as long as you support the product. At a minimum, every new feature should come with a quick check on whether the threat model needs an update. If it does, you go back to design, update the model, select your requirements and produce the evidence again. The same question applies to the activities that consume the model’s output. In OWASP SAMM terms, the architecture assessment you do during verification, the security requirements you selected and your vulnerability triage all generate evidence, and each of them should tell you whether the threat model gives them the input they need. When it doesn’t, the model has to grow. Expect it to take a few years before most organisations can say with confidence that their threat modeling process runs smoothly.
Once impact is about the user, the scoring has to follow. A 4×4 matrix gives enough resolution to separate risks that deserve different treatment, and it is simple enough for a developer to use in a design review without a risk analyst in the room.
Likelihood works the way you are used to: how exposed the entry point is, what preconditions an attacker needs, what skill and access it takes, and whether public exploits or tooling exist. Impact is rated against your user-harm scenarios.
| Likelihood / Impact | 1 Negligible | 2 Contributory | 3 Significant | 4 Critical |
|---|---|---|---|---|
| 4 Near-certain | Low | Medium | High | Critical |
| 3 Likely | Low | Medium | High | Critical |
| 2 Unlikely | Low | Low | Medium | High |
| 1 Remote | Low | Low | Medium | Medium |
Because every impact score names the scenario it comes from, the rating doubles as the written justification Article 13(3) expects.
To make this concrete, here is a compact example of a wall-mounted EV charger for private homes, with firmware on the charger, a mobile app and a cloud backend built and operated by the vendor. The app uses Bluetooth for setup and the cloud for everything after that. The charger also talks to an energy provider’s API to charge when electricity is cheap.
Scope. Charger firmware, mobile app, cloud API and the firmware update service are in scope. The cloud backend is in scope because the vendor builds and operates it and the charger cannot be controlled remotely without it. The energy provider’s API, the cloud hosting platform and the open-source libraries are external entities, handled in the supply chain track.
User-harm scenarios.
Threats, ratings and countermeasures. A first session on the baseline produced about twenty threats. Seven of them show the pattern well.
| ID | Threat | Scenario | L x I | Risk | Countermeasure | Annex I, Part I(2) |
|---|---|---|---|---|---|---|
| T1 | Bluetooth setup uses a static PIN, so anyone in range can pair during installation | D1 | 3 x 3 | High | M1: unique pairing secret per device (QR on label), pairing only after a physical button press, pairing locked afterwards | (b), (d), (j) |
| T2 | Cloud API does not check that the charger belongs to the logged-in account (broken object-level authorisation) | D1, D4 | 3 x 4 | Critical | M2: ownership check on every charger endpoint, authorisation tests in CI, non-guessable charger IDs, alerts on bulk commands | (d), (i), (l) |
| T3 | Firmware updates are not signature-checked, so a malicious image can be installed | D1, D3, D4 | 2 x 4 | High | M3: signed images, secure boot, A/B partitions with automatic rollback, anti-rollback counter | (c), (f), (h), (k) |
| T4 | Per-minute charging data is kept in the cloud indefinitely and exposed in a breach | D2 | 2 x 3 | Medium | M4: store hourly aggregates only, 13-month retention, encryption at rest, deletion from the app | (e), (g), (m) |
| T5 | Denial of service on the cloud API (or an outage) stops the charger from charging | D3 | 3 x 3 | High | M5: local fallback mode that keeps charging with the last known schedule, rate limiting at the API edge | (h) |
| T6 | A compromised fleet-admin account sends a mass start or stop command | D4 | 2 x 4 | High | M6: MFA and two-person approval for fleet commands, fleet-wide rate limit, randomised start delay and maximum current enforced in firmware | (d), (i) |
| T7 | Debug port left open on production boards, and one key shared across the fleet can be extracted | D1, D4 | 2 x 4 | High | M7: debug interfaces disabled in production, per-device keys (drops impact to contributory: one unit, not the fleet) | (e), (j), (k) |
Two things stand out. T2 is the classic one: a single missing check that turns one user’s bug into a fleet-wide problem, which is why it scores critical. And look at where the T6 mitigations live. The limits that protect the grid sit in the firmware, so they still hold if the cloud is compromised. Asking “what if the component I trust most is the one that fails?” is often where the best countermeasures come from.
Now roll it up. Annex I, Part I, point 2 lists thirteen requirements, (a) to (m). For each one, you record what covers it, the evidence, the status and what is left. This table is the part of the technical documentation that answers Article 13(2) most directly.
| Requirement | Covered by | Evidence | Status | Gap or residual risk |
|---|---|---|---|---|
| (a) No known exploitable vulnerabilities when placed on the market | Supply chain track: SCA in CI, SBOM triage; pentest of firmware and cloud API | Dated scan reports, pentest report | Partially met | No release gate that blocks on open exploitable findings. Add before first release. |
| (b) Secure by default configuration, reset to original state | M1, factory reset | Default configuration baseline, test cases | Partially met | Factory reset removes pairings but keeps Wi-Fi credentials. Fix planned in firmware 2.4. |
| (c) Security updates, automatic by default with opt-out, notification, option to postpone | M3 | Update design, signing procedure, rollback tests | Partially met | Automatic updates are on, but the app cannot notify users or let them postpone. Backlog item raised. |
| (d) Protection from unauthorised access, reporting possible unauthorised access | M1, M2, M6 | Authorisation tests in CI, IAM review | Partially met | Access control is in place, but owners are not told about new pairings or logins. Add notifications. |
| (e) Confidentiality of stored, transmitted and processed data | TLS to cloud, encrypted Bluetooth after pairing, M4, M7 | Configuration export, pentest | Met | Residual risk R2 (physical extraction from one device). |
| (f) Integrity of data, commands, programs and configuration, reporting corruption | M3, signed cloud commands | Secure boot tests, code review | Met | Integrity failures shown as an error state in the app. |
| (g) Data minimisation | M4 | Data inventory, retention job configuration | Partially met | Retention job not yet running in production. |
| (h) Availability of essential functions, also after an incident; DoS resilience | M3 rollback, M5 | Failover and load tests | Met | Residual risk R3 (hosting provider dependency). |
| (i) Limit negative impact on other devices and networks | M2, M6 | Firmware tests, design record | Met | Residual risk R1 (signing key compromise). |
| (j) Limit attack surfaces, including external interfaces | M1, M7, local setup service shut down after pairing, unused endpoints removed | Production checklist, port scan | Met | None open. |
| (k) Reduce incident impact with exploitation mitigation mechanisms | M3, M7, compiler hardening, memory protection unit, hardened containers | Build flags, image scans | Met, with justification | The microcontroller platform has no address space randomisation. Justified in the risk assessment. |
| (l) Record and monitor relevant internal activity, with user opt-out | Cloud audit log for commands and configuration changes | Logging configuration | Not met | No on-device log and no opt-out for users. Needs design work before release. |
| (m) Secure permanent removal of data and settings; secure transfer | M4, factory reset, account deletion in the app | Test cases | Partially met; transfer not applicable | Cloud backups keep data for 35 days after deletion. The product offers no data transfer to other products, so that part is not applicable (justification recorded). |
Three lessons from this table.
First, the threat model covers (b) to (k) well, because those requirements describe attack paths and their defences. Requirements (a), (l) and (m), and the user-facing parts of (c) and (d), rarely come out of a threat modeling session on their own. They describe lifecycle activities and product features, so you have to walk through the full list on purpose.
Second, “partially met” and “not met” are fine during development, as long as each has a backlog item and a target release. They are not fine when the product is placed on the market. By then every row needs to read “met” or “not applicable” with a justification. The gap column is your pre-release backlog.
Third, “met” does not mean “zero risk”. Annex I, Part I, point 1 asks for an appropriate level of cybersecurity based on the risks. A requirement can be met while some residual risk remains, as long as you have written down what it is and why it is acceptable.
The CRA gives you no numeric threshold for acceptable risk. You have to argue that the level you reached is appropriate for the risk to the user and in line with the state of the art. In practice that means four things per accepted risk: what it is, why you accept it, what reduces it, and who reviews it when.
| ID | Residual risk | Why it is accepted | What reduces it | Owner and review trigger |
|---|---|---|---|---|
| R1 | Compromise of the firmware signing key | Cannot be removed completely; likelihood remote with current controls | Key held in an HSM, two-person signing, rotation and revocation plan | Head of engineering; yearly and on any change to the signing pipeline |
| R2 | Secrets extracted from one charger by someone with physical access | Intended use is a private home; per-device keys limit it to one unit | Debug interfaces disabled, tamper-evident enclosure | Product owner; re-assess if the charger is sold for public or shared car parks |
| R3 | Hosting provider outage | Provider SLA, and charging continues locally | M5, status shown in the app | Operations lead; on any provider change |
| R4 | Owners who switch off automatic updates stay exposed to fixed vulnerabilities | The CRA requires that users can opt out | Reminders in the app, security advisories, clear user information | Product owner; every release |
Note the review trigger on R2. Selling the same charger for shared car parks changes the reasonably foreseeable use, which changes the risk. Tying residual risks to the assumptions behind them tells you when to reopen the model.
Some residual risks also belong in front of the user. Annex II requires you to tell users about known or foreseeable circumstances that may lead to significant cybersecurity risks. R4 is a good example: if users can switch off updates, the user information should explain what that means for them.
A good threat model covers a large part of what the CRA asks for, but not everything, and it helps to be clear about the rest so nobody assumes it is handled:
The threat model feeds some of these. It helps your triage decide whether a reported vulnerability is actually exploitable in your product, and your update pipeline is worth modeling as a system in its own right. Still, these are processes rather than design decisions, and each deserves its own treatment. We are working on a CRA playbook that covers them in depth. Subscribe to the Threat Modeling Insider to get it as soon as it is out.
Most of the code in a modern product was written by someone else, and it is routinely left out of threat models. How to handle it is the subject of Expanding scope to the software supply chain. Article 13(5) makes ignoring it a compliance problem as well as a security one.
Once teams notice the gap, the reflex is to pull every library into the model. We advise against it, because you cannot mitigate a flaw in code you don’t control, and a model that tries grows until nobody maintains it. Run the supply chain next to the threat model instead.
Start with an SBOM generated in the build, not assembled by hand. Read it for risks that never show up as CVEs: abandoned projects, single-maintainer dependencies, licences that don’t belong in a commercial product. They matter for your CRA position, because Annex I, Part II commits you to security updates, and an unmaintained upstream makes that hard to keep.
Include the people, too. The Shai-Hulud npm worm in 2025 spread through compromised maintainer accounts and stolen publishing tokens, not through a bug in anyone’s code. When you assess a vulnerability, look at whether it is reachable in your product, not just at its CVSS score, and rate it on the same user-harm scale as everything else. Record what you checked, when and against which source.
Then hand the leftovers back to the threat model as explicit assumptions. In the charger example, one assumption reads: “The energy provider’s tariff API authenticates our cloud and cannot send commands to chargers.” If that assumption breaks, T6-style threats come back and the position on (d) and (i) needs another look. Writing it down is what makes that link visible.
A practice that only a security specialist can run won’t scale to a regulation that touches every release. Five things make the difference.
Hook into moments you already have. Design reviews, architecture decisions, new integrations and major dependency upgrades all have a security angle. Adding the four questions there costs far less than scheduling separate sessions.
Model the change, not the whole system. Build the baseline once, then work in deltas. For a moderate change against an existing model, an hour is usually enough. Bring the developer who owns the change, a tester, and someone who knows the product’s users.
Keep the artefacts next to the code. Diagrams, threats and the Annex I mapping belong in the repository and are versioned with the code. A layout like this works well:
threat-model/
scope.md # product definition, intended use, foreseeable misuse
dfd.drawio # data flow diagram with trust boundaries
scenarios.md # user-harm scenarios D1..Dn
threats.yaml # threats, ratings, decisions, Annex I tags
annex-i-mapping.md # position per Annex I, Part I(2) requirement
residual-risk.md # accepted risks, owners, review triggers
assumptions.md # supply chain and environment assumptions
A single threat entry can then carry everything the mapping needs:
- id: T2
title: Cloud API does not check charger ownership
scenarios: [D1, D4]
likelihood: 3 # public API, IDs were guessable
impact: 4 # remote control of any charger
risk: Critical
decision: mitigate
countermeasures:
- id: M2.1
ticket: CHG-412
text: Ownership check on every charger endpoint
verified_by: authorisation tests in CI
- id: M2.2
ticket: CHG-413
text: Rate limit and alert on bulk commands
annex_i: ["I.2(d)", "I.2(i)", "I.2(l)"]
last_reviewed: 2026-09-18
With the Annex I tags in the threat file, the mapping table becomes a report you generate rather than a document you maintain by hand. The git history gives you a dated, attributable record of how the assessment evolved with the product, which is strong evidence you never had to produce separately.
Let the team own the model, and have security review it. Early team-made models will miss things, and that is a fair price: a model that is 80% as good and updated with every release beats a perfect one produced once, and the quality gap closes as teams practise.
Grow it, don’t launch it. Pick one product, build a lean baseline, run it through two or three real changes, fix what didn’t work, then move on to the next team. Our Threat Modeling Playbook describes how to build that capability step by step.
After a few months of working this way, most of the risk assessment part of your technical documentation already exists:
Because it all lives in the repository, you also get the change history, and that history is your proof that you assess cybersecurity risk throughout the product’s life.
The case for threat modeling under the CRA is strong, but it only pays off if the practice survives its first quarter. Teams that threat model only to satisfy the regulation tend to produce the minimum that passes review, while teams that do it because it improves their design decisions end up with a safer product and with something the regulation accepts anyway.
So scope the model to what you control. Handle what you integrate on a separate track. Rate impact from the user’s side. Tag every countermeasure with its Annex I requirement, and write down the gaps and the residual risk you accept. Then attach the work to decisions your team is already making, and keep the first round small enough that the second one actually happens.
Book a discovery call for Toreon’s threat modeling training and turn the ENISA playbook into a working product-security practice.
Start with Threat modeling as a strategic path to CRA compliance for the business case.
Want your teams to run this themselves? Toreon’s threat modeling training puts the practice in their hands, so CRA evidence becomes a by-product of how you already design.
Asma is a principal cybersecurity consultant passionate about securing systems and enhancing development practices. With expertise in code analysis and scanning technologies, she specializes in identifying vulnerabilities throughout the software development lifecycle. Asma has conducted research into leveraging generative AI for security improvements, exploring how artificial intelligence can enhance threat detection and automate vulnerability assessment. As a trusted advisor to development teams, she combines technical depth with practical strategies to help organizations build robust security into their development processes.

