Dit artikel is momenteel in het Engels beschikbaar. Als je het liever samen in het Nederlands bespreekt, nemen we dat graag mee in een gesprek.

NIS2 has pushed many organisations into the same uncomfortable conversation: the security programme may exist, but how much of it has actually been tested against the way incidents unfold? Policies, risk registers, supplier questionnaires, and tooling inventories all have their place. They do not prove that a real attacker path would be noticed, contained, or prevented from reaching a critical service.

This article is not legal advice. National implementation, sector guidance, and your own risk context matter. But from an offensive security point of view, the practical question is straightforward: what evidence would convince you that the controls protecting your important services work under pressure?

Start with critical services, not with test types

"Critical services" is a useful starting phrase, not the whole NIS2 story. The testing plan should also consider the systems that support those services: identity, remote access, cloud administration, suppliers, backup administration, logging, and the operational processes people rely on during an incident. Sometimes the most important test target is not the service itself, but the path that could quietly control or disrupt it.

A weak testing plan starts with a menu. Web pentest. Infrastructure pentest. Phishing. Red team. Cloud review. The labels are familiar, so they feel like progress. The stronger plan starts somewhere less tidy: which services would hurt the organisation if they were disrupted, manipulated, exposed, or used as a stepping stone?

For one organisation, that may be an external customer portal tied to operational data. For another, it may be identity infrastructure, remote access, a managed service provider connection, a logistics platform, a payment workflow, or a cloud environment that quietly holds the keys to several internal systems. Once those services are named, offensive testing becomes easier to aim.

Translate risk questions into tests

The point of testing is not to collect activities. It is to answer questions. Can a vulnerability in an exposed system become internal access? Can a compromised user reach sensitive data? Can identity misconfiguration turn a small foothold into broad privilege? Would defenders see the steps that matter? Can the incident process make decisions quickly enough when the evidence is incomplete?

Different tests answer different parts of that picture:

  • A penetration test gives depth on a defined asset, application, network segment, or cloud environment.
  • An attack-surface review helps when nobody is fully sure what is reachable from the outside.
  • An internal identity assessment shows how far an attacker could move after an initial foothold.
  • A phishing or initial-access validation tests whether a realistic entry path still leads somewhere meaningful.
  • A red-team or adversary simulation tests a broader path, including detection, escalation, and decision-making.
NIS2 pressure Useful offensive-security evidence
"Do we know what is exposed?" Attack-surface review, external validation, and ownership cleanup for unknown or stale assets.
"Can one exposed issue become business impact?" A scoped attack-path review that tests whether external exposure can reach identity, data, or operational systems.
"Would we see the important steps?" Detection validation around privilege escalation, lateral movement, data discovery, and persistence.
"Are suppliers part of the path?" Testing or procedural validation of remote support, federation, escalation contacts, and evidence access.
"Did remediation change anything?" Retest notes, closed attack paths, updated detections, and a record of what remains accepted or deferred.

None of these is automatically better than the others. The right choice depends on the evidence gap. Buying a red team when nobody has recently looked at exposed assets or identity hygiene can be an expensive way to rediscover basics. Buying only narrow pentests when the real concern is detection and response can leave leadership with false comfort.

Treat "effectiveness" as a practical word

Security controls can fail in quiet ways. A control may be deployed but scoped too narrowly. Alerts may exist but go to the wrong queue. Logging may be available but retained for too little time. An EDR rule may fire but not trigger containment. A privileged account may be monitored but still usable from the wrong workstation. These are effectiveness problems, not simply tooling problems.

Offensive testing is useful because it connects those details into a path. It can show whether controls work together or merely exist next to each other. That is the difference between saying "we have monitoring" and showing what the monitoring did when the attacker moved from initial access to privilege escalation, lateral movement, data discovery, or persistence.

Supplier and third-party edges deserve attention

Many organisations have better visibility over their own core systems than over the edges around them. Managed services, outsourced platforms, remote support paths, SaaS integrations, shared identity providers, and inherited trust relationships can become part of the attack path. If testing never touches those edges, the programme may prove the safest part of the environment while leaving the most convenient route untested.

This does not mean every supplier needs to be attacked. It does mean scope should reflect reality. Sometimes the right test is contractual and procedural: who can make changes, who receives alerts, who owns escalation, and what evidence can be collected during an incident? Sometimes the right test is technical: remote access, identity federation, exposed admin surfaces, or cloud trust boundaries.

Useful evidence has a remediation trail

The output of offensive testing should be readable months later by someone who was not in the room. It should explain what was tested, why it mattered, what the tester attempted, what succeeded, what failed, what was out of scope, and what changed afterwards. A report without remediation tracking is weak evidence. It proves activity, not improvement.

The remediation trail should separate urgent attack paths from background hygiene. Both matter, but they do not deserve the same sequence. If a finding gives a realistic path into a critical service, it should not sit beside low-risk clean-up work as if all rows in a spreadsheet are equal.

Build a cadence, not a single event

NIS2 pressure can tempt teams into a one-off assessment. That may help in the short term, but it will not stay convincing for long. Environments change. Suppliers change. Identity changes. Cloud permissions drift. New services appear. Old exceptions remain. Testing needs a rhythm that follows that reality.

A practical cadence might combine continuous attack-surface awareness, periodic testing of critical applications, identity-path reviews after major changes, targeted phishing or initial-access validation, and occasional broader simulations when the basics are mature enough. The exact mix will differ by organisation. The principle is the same: evidence should follow risk, not calendar theatre.

The useful starting point

If you are under NIS2 pressure, resist the urge to buy the most impressive-sounding test first. Start by naming the services that matter, the assumptions you currently make about them, and the evidence you would need to trust those assumptions. Offensive testing becomes much more valuable once it is tied to those questions.

That is also a better conversation with leadership. Instead of "we ran a pentest because NIS2", you can say: "we tested the paths most likely to affect these services, found these gaps, fixed these first, and scheduled follow-up where evidence was still weak." That is a stronger story because it is not just a story.


If you are trying to turn NIS2 pressure into a testing plan that is actually defensible, we can help scope it. The useful starting point is usually your critical services, not a generic package list.