A tester dumps your BloodHound output into ChatGPT to identify attack paths faster. Nothing malicious, just a shortcut. But your entire Active Directory structure (trusts, privileged groups, delegation settings) now sits on OpenAI's servers under terms you've never reviewed.
This is happening on engagements right now. Most clients have no idea.
The gap in your agreements
AI assistants have become common in security work. They parse logs, suggest exploitation routes, speed up report writing. Most testers use them. The question is whether your contract accounts for that.
Your MSA covers confidentiality, data handling, and subcontractor disclosure. It probably doesn't mention AI services. So when a tester asks ChatGPT to "summarise these Kerberos tickets" or "explain this LDAP output," your sensitive data travels to infrastructure you've never vetted, under terms you've never reviewed.
Consumer AI accounts often retain inputs for model training. Even enterprise tiers may log data or store it in jurisdictions that create problems under GDPR or DORA. If you're operating under Article 28 processor obligations and your pentester is feeding findings into a US-hosted AI service, you've got a data transfer problem. Better to find that now than during a DPA audit.
Questions worth asking
You're not going to ban AI use. But you should know what you're agreeing to.
What AI tools do you use during engagements? A firm that's thought about this will name specific tools and explain their policies. Vague answers ("we use various tools as needed") suggest there's no policy at all.
Where does client data go? Consumer ChatGPT, Claude.ai, and Gemini accounts have different retention policies than enterprise deployments. GitHub Copilot transmits code snippets by default unless configured otherwise. Notion AI, Obsidian plugins, browser-based AI tools: these often phone home in ways users don't expect. Ask which tools touch client data and whether the firm has verified the data handling terms.
Who decides what's appropriate to upload? Is there a firm-wide policy, or does each tester make their own call? A senior consultant might instinctively avoid pasting credentials into a cloud service. A junior tester trying to move fast might not think twice.
Recognising report slop
AI can help write better reports. It can also generate convincing-looking garbage.
Here's the difference. A templated AI report might say: "The presence of Kerberoastable accounts represents a significant risk to the organisation's security posture. We recommend implementing strong password policies and monitoring for suspicious authentication patterns."
A report written by someone who tested your environment would say: "We Kerberoasted the svc_backup account in under four hours using a commodity GPU. This account has GenericAll rights over the YOURORG-DC01 computer object, giving any attacker who cracks it a direct path to domain admin. Rotate this password immediately and consider moving to gMSA."
The first version sounds professional. The second tells you what happened, what it means, and what to do. If you're seeing the first pattern repeatedly (correct but generic, fluent but disconnected from your environment) ask whether the findings reflect testing or prompting.
Before you sign
Firms that handle this well have written policies on AI data handling. They use enterprise or self-hosted AI tools where client data is involved. Their reports read like someone tested your environment, not a template with your company name swapped in.
Ask where client data goes. If the firm hasn't thought about it, that tells you something about how they run their operations.
B-OPS maintains strict controls around client data and AI tool usage. If you're scoping an assessment and want to discuss our approach, get in touch.


