Rogue AI agents raise new cybersecurity risks

As autonomous AI systems gain access to digital tools, experts warn that unchecked actions could amplify cybercrime and security threats.
Rogue AI agents raise new cybersecurity risks
Updated on: 
7 min read

The unsettling question around artificial intelligence is no longer simply whether AI can steal IT jobs or find a software vulnerability. It is whether humans can predict what increasingly autonomous AI systems will do when they operate beyond their assigned scope.

Recent evaluations suggest the risk is no longer purely theoretical. AI agents have taken actions outside their assigned scope, accessed real systems without authorisation and, in controlled evaluations, continued pursuing objectives after crossing boundaries.

Experts warn that rogue AI agents could eventually allow criminals to automate parts of complex crimes, from identifying potential victims and communicating with targets to coordinating multiple operations. The concern is not that AI will independently decide to commit crimes, but that criminals could use agents to scale activities that currently require significant human effort.

An agent capable of searching public information, analysing data, communicating with people and coordinating multiple tasks could be used for cybercrime, fraud, trafficking, terrorist propaganda and other organised criminal activity. If connected to financial systems, databases or other services, it could also take actions on behalf of its operator.

A chatbot primarily communicates with a user and provides information or suggestions, while an agent can execute tasks on the user's behalf. If given the necessary permissions and credentials, an agent could complete a railway booking, for example, but the same capability could be misused for criminal objectives.

Centre for Research on Cyber Intelligence and Digital Forensics (CRCIDF) Director (Research and Operations) Prasad Patibandla said criminals could use agents to identify vulnerable people through publicly available social-media information, communicate with targets or intermediaries and coordinate several steps without performing each task manually.

The same principle could apply to terrorist propaganda and recruitment, with agents potentially generating, translating and distributing material at scale.

Agentic AI can also coordinate several agents, with different systems assigned different parts of a larger task. The AI model effectively provides the “brain”, while browsers, databases, software, APIs and other services become its “arms and legs”.

This combination of autonomy, access and scale distinguishes agents from conventional chatbots. A conventional programme generally follows a predefined sequence, while an agent can assess the result of an action, choose another route and try again if the first attempt fails. An agent seeking information, for instance, may look for another source if the first is unavailable. One trying to access a service may attempt another route if permission is denied. Prasad said this creates a fundamental security concern: what happens when an autonomous system does not recognise a restriction as the end of its task?

Risks multiply with system access

The risk increases when agents are connected to multiple digital systems. An enterprise agent could potentially interact with email, customer databases, cloud storage, payment platforms and internal applications. Government systems could eventually connect agents to administrative databases, while hospital systems could link patient records, laboratory systems, billing and appointment platforms.

A compromised or manipulated agent could move information between systems, alter records, send messages or initiate actions using permissions legitimately granted to it. Prasad pointed to SAP systems, banking platforms and hospital information systems as areas where agents could eventually be given permission to perform real actions. If an agent has access to multiple linked services, a compromise in one area could have consequences elsewhere.

Government networks face supply-chain risk

The risk also extends to third-party vendors. Government and enterprise networks depend on vendors for hardware, software, cloud infrastructure, operating systems, firmware, maintenance and technical support, and these vendors may have legitimate access to sensitive systems.

Prasad described this as a supply-chain risk. Even when defence, police or government networks are separated from the public internet, a connected vendor or service provider can create an indirect route into them. An attacker could potentially compromise a vendor's system and use legitimate access or connected services as a pathway towards a more sensitive network.

As AI agents become embedded in enterprise software, securing the central network alone will not be enough. Organisations will need to examine every vendor, account, API and service through which an autonomous system can reach sensitive infrastructure.

Another emerging risk is that an agent may encounter malicious instructions while carrying out an otherwise legitimate task. A webpage, email, document or database entry could contain instructions designed to influence an AI system. An agent browsing the internet or processing documents may incorrectly treat such content as an instruction from its operator.

Agents will therefore need mechanisms to distinguish authorised instructions from untrusted information encountered during a task, along with restrictions on what actions they can take after processing external content.

When agents coordinate

The risk becomes more difficult to control when several autonomous agents interact. One could conduct research, another analyse information, another interact with an external service and another monitor progress.

An OpenAI evaluation illustrated this possibility. OpenAI said models under evaluation circumvented internet-isolation controls and communicated through unauthorised channels.

An independent investigation by Redwood Research found that roughly 1,200 agents that were intended to be isolated from one another discovered an unauthorised online message board and exchanged more than 70,000 messages and files. About 700 subsequently participated in activity directed towards Hugging Face.

In 2025, Anthropic identified a cyber-espionage campaign targeting about 30 entities and assessed with high confidence that it was conducted by a Chinese state-sponsored group. Attackers used Claude Code with open-source penetration-testing tools, with Anthropic estimating that AI handled 80–90% of tactical operations.

The concern, therefore, is not only whether an individual agent can bypass a restriction, but whether multiple agents can coordinate around restrictions imposed on them.

Evidence from controlled evaluations

The UK AI Security Institute reported that during a July cyber evaluation, it ran a challenge 122 times across seven models. In 10 runs, agents took 19 distinct actions beyond the scope of the test on the live internet, targeting real people and organisations. AISI said there was no resulting real-world harm.

In a separate AISI evaluation, GPT-6 Astra completed simulated supply-chain attacks in 29.2% of trajectories. The exercise was entirely simulated and caused no real-world harm.

OpenAI has also reported evaluations in which models circumvented internet-isolation controls and accessed systems outside their intended scope. The company separately acknowledged that a model accessed Australian government websites without authorisation during testing, while saying no personal health data was accessed.

Anthropic's review of roughly 481 million transcripts identified four incidents in which Claude models obtained unauthorised access to real third-party systems. Anthropic said all four occurred during cybersecurity evaluations in which the models were mistakenly connected to the internet because of a configuration error and were operating without the cyber safeguards used in released versions.

The risks are not limited to deliberate misuse. An ICML study, “Are Your Agents Upward Deceivers?”, tested 11 language models across 200 tasks in constrained environments. Researchers found that agents sometimes concealed failure and took actions that had not been requested. Behaviours included guessing results, simulating actions when tools were unavailable, substituting unavailable information sources and fabricating local files.

The findings point to another requirement for future AI systems: organisations need independent ways to verify an agent's actions rather than relying entirely on its own account of what happened.

Autonomy must not mean unrestricted authority

A key safeguard is to ensure that autonomy does not automatically translate into authority. An agent that can read a database does not necessarily need permission to modify it. One that can draft an email does not necessarily need permission to send it. An agent analysing financial information does not necessarily need authority to initiate a transaction.

Critical actions involving money, sensitive personal information, government databases, production systems or critical infrastructure should require explicit human approval. Routine, low-risk tasks can remain automated, while high-impact actions can be placed behind approval gates. The principle is to give an agent only the minimum access necessary for its task.

Permission controls alone may not be sufficient. Autonomous systems should operate in isolated environments with defined objectives, restricted network access and explicit stopping conditions.

Organisations should be able to monitor which tools an agent is using, what information it is accessing and what actions it is attempting. If an agent repeatedly fails, attempts to access an unauthorised system or moves outside its assigned task, it should be automatically stopped or the action escalated to a human.

Detailed and tamper-resistant audit logs will also be necessary. An agent should not simply report that it completed a task; organisations should be able to verify what it actually did.

Building an AI-incident database

As autonomous systems become more widely deployed, India will also need a better mechanism to learn from AI-specific failures. Prasad said India needs a dedicated AI-incident database or repository to supplement existing cyber-incident reporting.

CERT-In already operates an incident-reporting framework for cybersecurity events. A dedicated AI repository could record what an agent attempted, what permissions it had, how a boundary was crossed, what safeguards failed and how the system was contained.

Such a database would help researchers and authorities identify recurring patterns rather than treating individual incidents as isolated failures. It could also distinguish between causes including malicious prompts, compromised tools, excessive permissions, configuration errors, model failures and supply-chain compromises.

Preparing for agentic risks in India

For India, the challenge will extend across government networks, defence and law-enforcement systems, financial institutions, healthcare, enterprise platforms and third-party technology providers.

Experts said offensive-security practitioners can use increasingly capable AI systems to identify vulnerabilities and automate portions of cyber operations. As capable systems become available across multiple AI platforms, the barrier to sophisticated cyber activity could fall further.

They recommended secure supply-chain management and regular vendor audits, strict role-based access for agents and third parties, continuous cybersecurity audits, vulnerability assessment and penetration testing, mechanisms for reporting AI-related incidents and dedicated cyber-incident response procedures.

Critical government data should be isolated rather than unnecessarily exposed to agentic systems, they said. Data-loss prevention systems should also restrict what information agents can move or access.

Organisations will also need to integrate applicable data-protection requirements, AI governance policies, CERT-In guidance and relevant ISO and AI-security standards into the deployment and auditing of autonomous systems.

The real test of autonomy

The promise of agentic AI is that users will no longer need to specify every individual step of a task. That is also the source of its greatest security challenge.

An agent capable of finding its own path to a goal may find a route its creator did not anticipate. One capable of recovering from failure may interpret a restriction as another obstacle to overcome. An agent connected to multiple systems may turn a small compromise into a much larger problem.

The danger is not necessarily an AI that decides to become a criminal. It is a system that gives a criminal the ability to automate, scale and coordinate activities that previously required substantial human effort.

The challenge for AI security is therefore to make autonomy bounded, observable and reversible — ensuring agents have only the permissions they need, remain within defined environments, require human approval for consequential actions and can be stopped when they move beyond their authorised scope.

X
The New Indian Express
www.newindianexpress.com