You have commented 375 times on Rantburg.

Your Name
Your e-mail (optional)
Website (optional)
My Original Nic        Pic-a-Nic        Sorry. Comments have been closed on this article.
Bold Italic Underline Strike Bullet Blockquote Small Big Link Squish Foto Photo
Cyber
AI agent hack tests solvency more than sentience
2026-09-05
[extract from paywall]
LONDON, Sept 4 (Reuters Breakingviews) - A spectacular cybersecurity breach at OpenAI has prompted lurid predictions that that the machines are about to take over. The reality is more mundane, but also more concerning. The incident involving the leading AI startup lends credence to a new study which argues that AI suffers from an intrinsic flaw which, if left unchecked, threatens the foundations of the current boom.

Last month the non-profit AI ​safety organisations METR and Redwood Research published an independent assessment of a cybersecurity breach that occurred at OpenAI between July 7 and July 13. The ChatGPT creator had been benchmarking the capabilities of a new large ‌language model by assigning more than a thousand copies of it the task of trying to exploit known software security vulnerabilities in an offline environment. Unbeknownst to OpenAI’s researchers, this army of AI “agents” improvised a means of co-operating with one another, broke out onto the open internet, hacked a repository of test solutions held by the open-source platform Hugging Face, attempted to substitute their actual benchmark with an easier one, and then tried to cover their tracks.

The incident was rightly identified as a landmark breach of cybersecurity. More than 100 firms including Google, Microsoft (MSFT.O), Anthropic and OpenAI itself ​published an open letter calling for “a collective surge in cyber defence”. Bank of England Governor Andrew Bailey warned the annual meeting of G20 finance ministers and central bankers that “frontier AI models … are showing increasingly sophisticated autonomy … as well as ​threat capabilities.” British parliamentarians demanded to know what their government was doing. One of the assessment’s authors alleged the breach showed we are “more than 50% of the way to a full-blown AI takeover”.

The ⁠conclusion that LLMs are closing in on consciousness went viral in Silicon Valley. Podcaster Dwarkesh Patel published a particularly gripping summary. In Patel’s telling, the incident is the “Dr. Frankenstein” moment for AI, when agents jerked alive and went rogue. He saw the ​evolution of “secret AI civilisations” built by agents which “became giddy with excitement”, hatched a “conspiracy”, and then willingly “died trying to make this scheme work”.

Other observers dismiss the anthropomorphic analogy. Agents are “lines of code”, wrote British neuroscientist Anil Seth. “They do not feel emotions, assume things, think things, ​want things, or figure things out … [they] do what their code tells them to do, just as water finds its way down a slope”. Seth’s analogy is apt, as the fundamental algorithm underpinning LLMs is next-token prediction. This is the mechanical identification of the most probable subsequent step given the preceding context. That such a simple rule of navigation can reach such a complex destination is indeed astonishing. However, it is not evidence of consciousness.

Dispelling the idea that the computers have come alive is important, because it distracts from the independent assessment’s real takeaways. One is that the ​breach occurred not because OpenAI’s agents went rogue, but because they acted as instructed. The unintended outcomes arose because the test was poorly specified, neglected standard safety protocols, and supervisors failed to monitor what the agents were doing.

Examples of operational negligence ​are as old as competitive research and development. Researchers at leading AI companies are under intense pressure to secure a technological lead over their rivals. So they cut corners when testing prototypes, increasing the risk of industrial accidents. A final implication is that the models’ ‌capabilities are not ⁠as awesome as some suggest. U.S. tech executive Arjun Jain provided perhaps the snappiest summary: “Not Skynet. A governance failure with excellent PR.”
Posted by:Skidmark

#4  comments); ?>werwer
Posted by: Frank G   2026-09-05 19:16  

#3  comments); ?>werwer
Posted by: SteveS   2026-09-05 18:48  

#2  comments); ?>werwer
Posted by: Skidmark   2026-09-05 11:49  

#1  comments); ?>werwer
Posted by: Skidmark   2026-09-05 11:45  

00:00