An alert can identify an attack correctly and still arrive too late to limit the damage. That is the most useful starting point for a security buyer reading Anthropic’s September threat report. AI lets smaller crews do more work in less time. Several also ran their operations on stolen AI accounts, leaving other businesses to handle the abuse.

This matters particularly if your organisation runs production AI integrations, relies on externally maintained applications, or outsources detection and response. Simply using an AI chatbot does not create all of these exposures. The relevant questions concern the credentials, permissions and services connected to it.

Before changing your testing programme, ask what your current scope would establish. Would it show whether you can contain detectable activity in time? Would anyone notice stolen AI access? And would the findings give your team enough evidence to fix either problem?

Anthropic’s September 2026 report covers operations disrupted between December 2025 and August 2026. The cyber cases provide reasons to revisit those questions.

Containment before impact

The report describes state actors, criminal crews and a lone hacktivist delegating reconnaissance, software development and data processing to agents. Their capabilities are not identical, but smaller teams can cover more ground. A response process built around slower, sequential activity may leave too much time between the first useful signal and containment.

Familiar tools appear throughout the cases. Their names alone tell you little about what a victim could detect: some run on attacker infrastructure, while others have legitimate uses. Your monitoring needs to collect the relevant evidence and put it in front of someone who can act.

The Russian espionage case adds another concern. Anthropic describes a workflow that rebuilt malware after security products detected it. Catching one sample gives limited assurance about the wider activity. Behavioural coverage and the response to an alert still need examination.

For a client, we would want an exercise timeline showing when activity became observable, when an alert reached the responsible team, and when containment became possible. Delays need an explanation. Was evidence missing, did triage stall, or did the responder lack authority to isolate the affected system?

That evidence makes the next decision clearer. A collection gap needs engineering work. An escalation delay needs a process or ownership change. Neither is adequately described by a finding that says “improve monitoring”.

Stolen AI access

This is the report’s most direct lesson for businesses adopting AI. ShinyHunters affiliates switched workloads onto victims’ keys. The European hacktivist used keys obtained from exposed containers. Another actor compromised an AI vendor’s evaluation sandbox and obtained production API keys.

The provider initially sees requests under the victim’s account. The account owner then has to separate legitimate usage from abuse and revoke access. A business does not need to be an espionage target for its stolen credentials to support someone else’s campaign.

AI keys should already be covered by credential management. These cases explain why checking only for unexpected spending is insufficient. Misuse can happen without a large bill, and a spending alert is of little use if nobody knows which application owns the key or what revoking it would interrupt.

A useful assessment would document where production AI credentials are held, who owns them and how access can be withdrawn. It should also establish whether available usage records are sufficient to investigate unexpected activity. Where the provider does not expose enough information, that limitation belongs in the findings.

For integrations, review what the connected system can access and which credentials it holds. The sandbox case is a reason to question whether an evaluation environment needs production access at all. Any resulting finding should identify the permission or isolation issue, its owner and the evidence needed to confirm a fix.

What the assessment should produce

In our November article on Cobalt Strike and MCPs, we argued that a control interface does not supply operational judgment. The newer cases show more work being automated. Human decisions still matter, but team size is becoming a less reliable guide to how much an attacker can attempt.

That does not mean every organisation needs a new engagement labelled “AI red teaming”. Start with the unanswered question. An application or cloud assessment can examine credential exposure and integration permissions. An exercise involving your defenders can establish what they see and how quickly they respond.

When reviewing a proposal, look for outputs you can use:

  • An observed timeline from the first available evidence to escalation and containment, with delays explained.
  • Documented gaps in AI credential ownership, usage visibility or revocation, tied to the affected service.
  • Findings assigned to a responsible team, with a clear way to verify that the agreed change worked.

A report that only says a compromise was possible leaves those decisions unresolved. Agree on the evidence you need before the exercise starts.

Supplier access and recovery

The European case makes this concrete. Anthropic reports that one French-speaking operator gained internal access to at least 14 of 42 tracked targets, including political organisations, media and their SaaS providers. It also describes poisoned backups. Restoring those backups could return the compromise to the recovered environment.

For businesses using similar applications or suppliers, the practical questions are familiar but specific: what access does the provider hold, who can revoke it during an incident, and how will restored systems be checked before they are trusted? Include those questions where they apply to the systems being assessed.

For organisations in scope of NIS2, this work can also inform risk assessments and evidence of control effectiveness under Article 21. The report helps inform the scope; it does not turn these priorities into a prescribed compliance checklist.

About the evidence

This is Anthropic’s investigation of misuse involving its services, informed by platform visibility and other intelligence. Attribution and confidence in individual figures remain subject to those limits. Unverified actor claims should not be treated as confirmed impact, and published indicators do not establish when victims detected an incident.

If you have a pentest or detection-and-response exercise planned, talk to us about reviewing its scope. We can work through which of these exposures apply to your environment and what evidence the engagement should produce.