AI is starting to move from finding security problems to fixing them.
Google’s Gemini 4 Argon is the latest example. Google says the model can autonomously detect, validate and patch critical software vulnerabilities, and is initially rolling it out to trusted cybersecurity specialists through its Fairwind programme.
That is exciting. It also changes the risk.
The moment an AI system can change code, configurations or production systems, the question is no longer just “is the answer correct?” It becomes “what can this system actually do if it is wrong?”
Before you let an AI system patch production, check these seven things. They apply whether the tool is Gemini 4 Argon or any other AI system that can make changes to software or infrastructure.
What can it actually change?
Start with scope.
Can the system only inspect code and suggest a patch? Can it open a pull request? Can it merge changes? Can it deploy directly to production?
Those are very different levels of authority.
A useful default is simple: give the AI the smallest level of access that still lets it do the job. If it only needs to propose a fix, it should not have deployment credentials. That same least-privilege principle is covered in more detail in Too much power: why AI agents shouldn't have access to everything.
Does a human still approve the change?
Autonomous does not have to mean unsupervised.
For high-impact systems, keep a human approval step before code is merged or deployed. That gives someone the chance to check whether the proposed fix is technically correct, whether it creates a new issue elsewhere, and whether the timing is sensible.
For lower-risk changes, you may choose more automation. The important thing is that the approval model is deliberate rather than accidental.
Is it working in a sandbox?
Security tools often need broad visibility. That does not mean they need broad freedom.
Run analysis and testing in an isolated environment where possible. Give the AI access to copies, test systems or controlled environments before allowing it near live infrastructure.
A sandbox will not solve every problem, but it reduces the blast radius when the model misunderstands something or behaves unexpectedly.
Can you see everything it did?
Logging matters more when software can act.
You should be able to reconstruct what the AI looked at, what it changed, which tools it called, what permissions it used, and who approved the final action.
Good logs are not just for investigations after something goes wrong. They also help teams learn which tasks are safe to automate and where the system still needs supervision.
Can you roll it back quickly?
A patch can be technically valid and still break something important.
Before allowing automated changes, make sure you have a reliable rollback path. That might mean version control, deployment snapshots, tested backups, feature flags or a known-good configuration.
If reverting the change is slow or uncertain, the AI should not be making that change automatically.
What happens if the AI is manipulated?
AI systems can be influenced by untrusted input.
A security agent may inspect source code, issue trackers, documentation, logs, web pages or other material that an attacker could potentially influence. That creates a risk that malicious instructions or carefully crafted content affect the agent’s behaviour. Our explainer on prompt injection covers that risk in more detail.
Treat the data the AI reads as untrusted. Restrict the tools it can call, separate reading from acting where possible, and do not rely on the model itself to enforce all of its own boundaries.
What credentials and secrets can it reach?
This is the check that is easiest to overlook.
An AI agent that can patch software may also be able to see repositories, cloud credentials, API keys, deployment tokens or internal documentation.
Map those permissions before deployment.
Ask:
- Which accounts does the agent use?
- What secrets can those accounts access?
- Are credentials short-lived?
- Can permissions be revoked quickly?
- Does the agent have access to anything unrelated to its task?
If the answer is “it can see most things”, reduce the scope before you automate more.
The useful rule: separate capability from authority
The safest way to think about AI security agents is to separate what they are capable of from what they are authorised to do.
A highly capable model does not need unrestricted access.
Let it reason broadly. Let it inspect what it needs. But keep permissions narrow, approvals explicit, actions logged and recovery tested.
That is how you get the speed benefit of AI without turning every mistake into a production incident.
Why this matters now
Argon is one early sign of a broader shift: AI security tools are moving from analysis and recommendations toward direct action.
For small teams, that could eventually mean faster remediation with less manual effort. But the more an AI can change, the more important its permissions, approvals, logs and recovery controls become.
The teams that benefit most will not be the ones that simply give these systems more power. They will be the ones that design the boundaries around that power first.
