Security audits in 2026: the AI agent helps, whether we like it or not
AI is changing IT security at a remarkable pace. In this article I look at how the everyday work of a security auditor differs from what it was a few years ago, and at the opportunities and problems that come with the change.

How a security audit used to work
A typical job meant going through the source code of a web application in detail and systematically tracking down as many of the known types of vulnerability as possible. An experienced auditor could get through roughly 10,000 lines of code a day, depending on the language and how complex the code was. Larger projects were audited by sampling, with tools searching for known patterns to cover the rest. Ideally the findings were then confirmed on a test system, or a penetration test was run alongside the audit. Depending on the setup, we also looked at the infrastructure, neighbouring systems and the configuration of the web server and operating system.
At the end, each vulnerability was rated by risk and written up in a final report, together with recommendations on how to fix it.
AI in 2022/2023: the first steps
ChatGPT or a similar assistant was always open in the background, helping to research current vulnerabilities, to use tools or to analyse small snippets of code. Whatever it came up with was checked for plausibility and completeness before it went anywhere near the actual work.
When it came to the report, the AI helped to polish the wording. The substance was still written by hand.
IDE integration in 2024/2025: AI in the code editor
With tools like Cursor, Windsurf, Antigravity or Copilot in VS Code, the AI chat moved into the IDE itself. A simple prompt such as "Find security issues" was enough for a first, rough pass over the code. It didn't catch everything, but it was a decent place to start. Writing reports became much easier thanks to autocomplete, and so did running tools: MCP (Model Context Protocol) provided the first interfaces to pentesting tools.
Fixing a vulnerability was often just a matter of a prompt like "Fix security issue XYZ" or "Add unit test for security issue". Later came plan mode, which lets you review the AI's proposed solution before anything is changed. That made a noticeable difference to code quality.
Agent-based workflows in 2026: autonomous audits
With Codex, Claude Code, Opencode or Mistral Vibe CLI, AI agents are taking on more and more work by themselves. To keep the workflow consistent and well defined, you give them "skills": ready-made sets of instructions that describe the task and the rules the AI has to follow. Claude Code, for example, already ships with a skill called /security-review.
If you want something tailored to your needs, you can start from an existing skill, such as the security-review skill from Sentry, and adapt it with a targeted prompt:
Research current web security standards and how to find vulnerabilities. Use owasp.org as a basis. Update this security-review skill. Focus specifically on applications written in PHP with framework XYZ.
The result is a skill tuned to the case at hand and a solid basis for a source code audit. That is only the first step, though. Then comes the iterating: improving the skill, refining it and running other models over the code (Codex after Claude, say) to get more precise results.
Sub-agents can also be used to split the work by type of vulnerability, for example one sub-agent for all injection issues.
There are several ways to validate the findings. The simple one is to ask for a proof of concept (PoC) that tries to exploit the vulnerability with tools that are already available, such as curl. The more advanced one is to connect pentesting tools like Burp Suite directly via MCP.
This way many vulnerabilities turn up already rated by risk. The auditor's job is now to judge how relevant each finding is, in particular whether it is a real problem or whether a business requirement, such as the application's business logic, justifies it.
In some cases the AI even opens a merge request with the proposed fixes, so the development team gets a ready-made starting point and can merge the changes after a short review. The line between security audit and development is getting blurry.
Vulnerabilities AI finds more efficiently
AI is particularly good at spotting:
- Insecure Direct Object References (IDOR) and Broken Object Level Authorization (BOLA): manipulated IDs that give access to data you shouldn't be able to see. Checking every relevant place by hand, in the code or during a pentest, takes a long time. The AI finds more of the problem cases, and finds them faster.
- Information disclosure and leftovers in the Git history: sensitive data such as passwords or tokens in the Git history or in frontend code is picked up automatically. Doing that thoroughly by hand in a reasonable amount of time is next to impossible.
Conclusion: a big step forward, but no replacement
Classic vulnerabilities like XSS or SQL injection are found almost without exception, and more complex ones such as IDOR or BOLA are found reliably too.
The AI still has its limits, though. It doesn't see the big picture: context, and how things fit together, are still for people to work out. And rating and prioritising risks takes expertise and experience.
So the role of the security auditor is shifting. Less time goes into tracking down and documenting vulnerabilities by hand, and more is left for:
- talking to stakeholders such as development teams and management
- security awareness training
- fitting the findings into the client's existing business processes
AI doesn't replace the auditor. It's a powerful tool that lets them concentrate on what matters: strategy, context and people.
A German version of this article was also published on LinkedIn on 7 August 2026.