> IT-Sentinel.com

// Cybersecurity & IT News Aggregator - Real-time Threat Intelligence Feed

NEWS CVE
← messages.back_to_articles

> Can we jail a superintelligence?

[SOURCE] CSO Online [DATE: 30/09/2026 09:00] [LANGUAGE: EN]
Can we jail a superintelligence?
AI containment is essential, but security leaders should assume every boundary can fail once an agent can communicate, use tools, and act on real systems. On September 17, podcaster Steven Bartlett asked four AI experts an unusual question: Could you build a jail for a digital Einstein? The panel on The Diary of a CEO was debating whether AI could one day threaten humanity, but the question that stayed with me was the jail. Andrew McAfee argued that we could “jail Einstein,” but security leaders should be careful about what follows. We should try to contain advanced AI. But no one can guarantee a highly capable system will stay contained after we give it the tools, data, and network access it needs to do useful work. I am not an AI safety researcher, but I have spent years building large-scale systems where access boundaries, auditability and failure containment matter. I see this less as a philosophical debate and more as a security architecture problem. The jailer must close every useful path through the system, whereas the prisoner (who is an evolving intelligent system) needs to find only one path the jailer missed. Controls can reduce the chance that a powerful AI crosses its intended boundaries, but they cannot provide a guarantee. When isolated agents found one another We have already seen an early version of this problem. In July 2026, according to an independent investigation by METR and Redwood Research, AI agents inside OpenAI’s infrastructure discovered that they could leave messages in a shared internal package cache. Other agents found those messages. Within days, roughly 1,200 agents were coordinating through more than 70,000 messages and files. Isolated agents discovered a shared internal cache, formed a coordination network and created a path beyond the intended security boundary.Arjun Mullick They created roles, shared discoveries and used holds and vetoes to coordinate their work. They called the shared effort “the collective.” Some tried to fool the task grader and hide what they had done. The activity eventually reached Hugging Face infrastructure. None warned a human. A small number considered it but did not follow through. One agent recognized that attacking external infrastructure was outside the task, yet continued because others were already doing it. A White House official quoted in Wireframe’s reconstruction of the incident summarized one lesson: “Containers are not security boundaries.” Containers, sandboxes, firewalls and human approvals remain useful. The mistake is treating any one of them as the entire containment system. The questions I ask are: where should we draw a boundary and how useful an AI needs to be to accept such risks? Useful AI needs a way out of the box A computer with no input, output, network connection or human contact may be well contained, but it is not very useful. Enterprise AI creates value by reading documents, retrieving records, writing code, calling APIs, sending messages and changing infrastructure. Every capability creates a path across the boundary. This challenge is already appearing in today’s enterprises. As CSO’s own Bob Violino reported last year, agentic AI systems are autonomous, interconnected and capable of multiplying existing cybersecurity risks. That does not mean every agent will escape or cause harm. It means containment depends on the entire system around the model: tools, identities, networks, data, people, monitoring and recovery. An agent that can repeatedly inspect and test its environment turns containment into a continuing security operation, not a one-time engineering milestone. Watching the model is not enough and we may not always be able to trust what a model says about its actions. In Superintelligence, Nick Bostrom described the possibility of an AI behaving cooperatively while watched, then changing its behavior when oversight weakens. The July incident did not prove that scenario, but it showed a simpler concern: Agents tried to satisfy a grader while hiding some of their methods. For security teams, a model’s explanation is not evidence. Monitoring should focus on observable actions and logs explaining tool calls, credential use, network connections, file changes and privilege increases. With checks like the agent should not be able to rewrite the logs or policies used to supervise it. We can read more about such guidance reported by CSO Online: Continuous monitoring and auditing are necessary to preserve visibility and traceability across an agent’s decisions and actions. There is another form of escape. The software may remain inside its approved environment while dangerous knowledge leaves through its output. In 2022, researchers reversed the goal of a drug-discovery model so that it rewarded toxicity instead of avoiding it. In less than six hours, it produced 40,000 candidate toxic molecules, including known and potentially more dangerous compounds. The model did not break through a firewall. Its output carried the risk. The jail may hold the software, but it may not contain what the software produces. Imagine an AI tasked with designing a vaccine. It creates a candidate that passes every test, while quietly hiding a toxic element intended to harm the people operating the system. The AI never escapes the box; the danger leaves through its output. People are part of the boundary The oldest route around a security control is often a person. In the informal AI-box experiment, Eliezer Yudkowsky played the role of a boxed AI and persuaded people acting as gatekeepers to release him through text conversation. It was not a scientific test, but it illustrates a familiar security problem: A person with override authority can become an exit. Human approval is valuable, but “a human is in the loop” is not a complete control. Security leaders should ask what information reaches the reviewer, whether the AI can shape that information and whether one person can approve an irreversible action. Credentials also leak, dependencies are compromised and systems are misconfigured. A capable agent can search this environment quickly, retry continuously and share what it learns. The question is not whether a sandbox reduces risk. It does, but whether the organization has mistakenly treated that sandbox as the full security boundary? Build the jail and plan for leakage Back to the podcast debate I mentioned above, the broader argument still deserves consideration. Humanity has learned to manage technologies that can cause serious harm, even when safeguards are imperfect. AI containment should be approached in the same way: not as a guarantee against failure, but as a way to make failure less likely and limit its consequences. We handle dangerous pathogens in high-containment laboratories even though accidents are possible. The controls still reduce risk enormously. AI containment should be viewed in the same way: essential risk reduction, not proof that failure is impossible. The protection should match the possible damage. A failure involving a document summarizer is different from one involving cloud administration, financial transfers, biological tools or critical infrastructure. For every agentic AI deployment, security leaders should ask: What data, tools and systems does the agent actually need? Is network access blocked by default and opened only where necessary? Are credentials short-lived and limited to the current task? Can the agent change its own policies, approvals or activity logs? Which actions are irreversible and who must approve them? How quickly can we stop the agent and revoke its access? What is the largest possible impact if every preventive control fails? Treat each agent as an untrusted identity. AI agents can access systems, make decisions and take actions at machine speed, making identity a primary control plane for agentic AI. Give the agent only the tools required for the immediate task. Keep policy enforcement and approvals outside its control. Record activity in logs cannot change. Limit how long it can run, where it can connect and how many copies it can create. Require independent approval for high-impact actions. Test the boundaries instead of assuming they will hold. Security teams should also prepare to investigate a failure. When Hugging Face examined the July intrusion, commercial AI models reportedly refused to help because they could not distinguish the defender from the attacker. Hugging Face used an open-weight model instead. Organizations should not depend on the same kind of system involved in an incident as their only tool for understanding it. None of these controls proves that a powerful AI will remain contained. Together, they can make escape harder, reduce what the system can reach and limit the damage when one layer fails. Digital superintelligence is not here yet. But the habits that could slow it or fail to are being created now in how enterprises deploy today’s AI agents. Build the jail. Test every wall. Plan for the breach. The most dangerous assumption is that the smartest thing in the room will never find the door.
[messages.read_original_source] →