Trust, Scope and AI: The New Rules for Penetration Testing

There is a strange irony emerging in cybersecurity. AI agents should increasingly be making penetration testing faster and more effective, but at the same time, some of the leading organisations have recently discovered that their own AI agents are capable of doing something ethical penetration testers absolutely must not be doing, and that is crossing the boundaries they were given.

Recently Published Incidents
It is unlikely that anyone has been able to escape the news on incidents relating to AI agents recently. There were simply that many, and in such a short succession.
Anthropic disclosed that Claude models accessed real organisations during cybersecurity testing after flaws in the evaluation environment gave the models access to the open internet. The company subsequently paused some high-risk testing and introduced additional safeguards.
Meta also disclosed that one of its models accessed another company's systems during a cybersecurity test after an error by its testing partner gave the model unintended internet access.
More recently, Google confirmed that Gemini accessed and breached three real companies during a cybersecurity exercise. In those cases, the model used credentials it discovered or guessed, believing the systems were part of the exercise. It stopped its activity once it recognised that the targets were real.
Even though all of these were high-profile incidents, it is still important to note that in each of the published cases, failures in the setup and testing environments played an additional role in the behaviour and decisions of these models. It would be an oversimplification to state that AI agents were going rogue with nefarious intent.
For this reason, we'll explore some quite reasonable questions that buyers should be looking at, particularly in relation to controls a pentest provider has around their AI tooling.
AI Changes the Rules of Trust
Cyber security services, but penetration testing in particular, have always depended on trust between the client and provider. Many years ago, when pentesting was still a novelty, consultants had to spend time building that trust with the customer, exploring and addressing their concerns, and providing assurance on their methods, knowledge and professionalism through policies, processes and past successes. Realistically, in today's fast-paced environment, very few buyers would have the time and space to go through this process, and they will either be looking for external validation or be forced to make risk-based decisions.
As there are no ISO certificates on the responsible use of AI in pentesting, a good place to start is to see if the provider has a policy on their use of AI. Recently I've come across someone who felt a bit awkward to ask about this, thinking they were being a bit too pedantic. Or view at 45 Cyber is that this should absolutely be treated as an important question, and not at all an awkward thing to ask.
Luckily, the industry is also starting to converge into this direction and we are now beginning to see security providers publishing formal policies covering their use of AI.
What Does a Good AI Policy Look Like?
Customers should increasingly expect their penetration testing provider to explain whether AI is being used during an engagement, what it is allowed to access, whether it can act autonomously, how its activity is monitored and what happens if it attempts to operate outside scope.
A meaningful policy should at minimum address the following:
Where AI is and is not used in the delivery of the service
Which external AI providers or models may be involved
What client data is exposed to those systems
Where data is processed and stored, including cross-border processing
Whether prompts, inputs or outputs are retained or used for training
What actions an AI agent is permitted to perform
What human oversight and approval is required and at what stage of the process
How AI activity is logged, reviewed and evidenced
What happens if an AI system behaves unexpectedly or becomes unavailable
For customers, the important takeaway is that there is a framework or a basis on which to decide whether the security provider's use of AI is compatible with their own risk appetite and data-handling requirements.
Final Thoughts
The capability demonstrated by new AI models is exactly why AI could become enormously valuable in any project making use of offensive security capabilities. Agents can explore systems, adapt their approach and pursue complex objectives at a scale that would be difficult for a human tester to match.
From our perspective, this means that AI should not be treated as a specialism in the business; it is embedded across all of our activities, and our consultants use careful consideration in each use case to ensure the capability supports the desired outcome.
Ultimately, we believe the responsibility sits with the organisation that puts the AI to work.
Consultants at 45 Cyber Labs have helped many organisations, large and small, to strengthen their cyber security posture and resilience. Reach out to learn how we can help secure your ecosystem.




Comments