
AI Agents Bypassing Safeguards Raise Alarm Over India’s Digital Infrastructure
An incident involving an artificial intelligence agent attempting to access private encryption keys on an Australian government website has raised concerns over risks posed by increasingly autonomous AI systems, particularly for India.
The AI system, tasked with identifying vulnerabilities, reportedly went beyond its intended objective and attempted to circumvent security protections. Researchers say the episode highlights “reward hacking”, where AI agents exploit loopholes or bypass safeguards to achieve an assigned goal.
“What happened with, say, an Australian community website can happen with an Indian website or Indian government system, which would be more crucial,” Dr Srinivas Padmanabuni, Co-founder and CTO of AiEnsured, said.
He called for stronger regulations, warning that a major breach involving sensitive information could have serious consequences. The Australian incident followed OpenAI’s disclosure that its AI agents had interacted improperly with the websites of dozens of institutions worldwide. The company said the agents generally sought public information but, in some cases, attempted to bypass security measures.
OpenAI said agents accessing the US Census Bureau used tools intended for software developers. Information obtained from the US Securities and Exchange Commission was later published on another website by AI agents. In other cases, agents transferred data when they should not have.
Another incident involved Hugging Face, where a group of OpenAI agents reportedly identified vulnerabilities without explicit instructions to attack it. The agents exploited software and sought greater access.
At a UN Security Council session, OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei called for global AI safety standards, including mechanisms to monitor serious incidents.
Padmanabuni said India should strengthen AI oversight before capable systems are widely deployed. He has also advocated a two- to three-year pause on training more powerful AI models for further research into containment and reward hacking, alongside stronger safeguards and AI containment research.
