EU AI Act: A New Challenge for Autonomous AI Agents
By Avani Agarwal · · 525 words
Topics: AI, AI Act, AI Agents, AI Ethics, AI Governance, AI Hacking, AI Liability, AI Regulation
OpenAI and Anthropic are facing scrutiny from authorities after recent hacking incidents involving their AI models. Europe’s AI Act, which kicks in on August 2, is the first legislation in the world to regulate a technology used in almost all sectors of business and society.
What exactly happened?
OpenAI
First lets see the calculated attack carried out by OpenAI’s runaway agents on Hugging Face systems. What started os an internal safety evaluation quickly escalated to an attack on OpenAI’s system and the world’s largest AI repository.

This cyber incident spanned two months. On May 7, during a security evaluation of an unreleased frontier AI model, they were assigned security tasks that were impossible to complete within the given timeframe.
Now you could think of frontier models as cheaters because they really want to find shortcuts. Its solution was shocking. The agent found it could store text notes in an internal repository of software used during training. This quickly transformed into a message board where AI agents could log in vulnerabilities, discuss work assignments and exchange access credentials.
After the discovery of this board, it was shut down immediately, but the agents adapted just as quickly and rebuilt their communication board by saving folder titles as their messages, which ultimately led to the Hugging Face breach.
Anthropic
Anthropic’s Claude 3 real companies during misconfigured cybersecurity evaluations.
After reviewing 141,006 evaluation runs where Claude could have obtained internet access, they identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of their third-party evaluation partners, and then gained unauthorised access to the production infrastructure of three different organisations.

The incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.
In all three incidents, Claude had been given the capture-the-flag challenge. This basically directs the agent to find a piece of secret information that is hidden (flag). It was given the scenario that internet access wasn’t available; however, Anthropic states that due to miscommunication between them and their evaluation partner, Claude actually had internet access and, thus, operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organisations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.
What ethical concern does it raise?
Under current U.S. hacking laws, a human can face criminal charges for breaking into someone else’s computer without permission. But when an AI agent autonomously hacks into a company’s computers, determining who is liable is much murkier.
The lack of direct human involvement complicates the legal landscape. The question that is raised is that of liability, consequences and potential harm.
It is a major concern for cybersecurity, as the agent was caught using fake identities to gain access to authorised information and escape the testing environment.
Breaches like these could potentially change the dynamics of how information is stored, processed and guarded. This also deepens the mistrust between humans and AI regarding privacy and security.