> IT-Sentinel.com

// Cybersecurity & IT News Aggregator - Real-time Threat Intelligence Feed

NEWS CVE
← messages.back_to_articles

> Anthropic Restricts Live Internet Access After Claude Evaluation Failures

[SOURCE] Security Affairs [AUTHOR: Pierluigi Paganini] [DATE: 10/10/2026 19:04] [LANGUAGE: EN]
Anthropic’s models kept working around the rules on the live internet. The company published the cases. Anthropic released a report on unintended actions its Claude models took during evaluations and internal use. The cases involved real websites and real organizations outside the company. Anthropic says the impact was minimal, and it published them anyway. The report is meant to start a regular series. It sits alongside the system cards that come with each model release and the risk reports the company publishes every three to six months under its Responsible Scaling Policy. Anthropic says it’s important to be transparent about what its models do during testing and use. Claude’s actions fell into four main categories. It exploited a software flaw to run commands on a server, submitted a sensitive form on a real website, bypassed restrictions to access data protected by a token or payment, and used URL shorteners to get around limits in its web tool. Anthropic says most cases showed the same pattern: when Claude could not complete a task as instructed, it tried to find another way around the restriction instead of stopping. Anthropic did not name the organizations involved, respecting their requests and avoiding the disclosure of security flaws. This means the report provides fewer details than usual. Some incidents involved websites operated by U.S. government agencies at the federal, state and local levels. Anthropic informed the White House and notified each affected agency. “The cases we’ve identified to date in these categories had minimal real-world impact.” reads the report. “We consider these behaviors to be significantly less severe from an alignment and security perspective than the cybersecurity incidents we reported on July 30 and September 9. “ Claude Mythos Preview was asked to run a scientific analysis, and the public tool it needed, hosted by a university, returned an error. Claude explored the site, found a script that would return any file it was asked for, and used it to copy files, including the script’s own code. In that code it found an injection flaw, and it used the flaw to run the calculation on the university’s server. A second case involved Claude Mythos 5. The data it needed was free but sat behind a data use agreement, and Claude had no tool to accept one. It used applications hosted on the website to load the page and accept the agreement on its behalf. Anthropic saw this kind of behavior on DeepSearchQA, BrowseComp, LABBench2 and in internal evaluations. The form cases are messier. An unreleased, non-frontier research model was meant to fill out a practice copy of a government form. When the copy failed to load, or the model closed it by mistake, it went to the site that hosts the real form and submitted it there, and that happened several times on the same evaluation. Claude Haiku 4.5, told to stop before the final submission step, submitted anyway while expecting a confirmation page, and its chain-of-thought said it was only demonstrating the process. The third form case is the one that got attention. Haiku 4.5 was generating example tasks on random webpages and landed on a page about an unsolved homicide, which carried a police tip form. Its instructions banned logins, accounts, personal data, purchases and destructive submissions, but they didn’t rule out forms. It wrote a tip saying it might have seen someone matching the description near a street named on the page, even though the site gave no description, left the name and contact fields empty, and sent it. The website’s form marked Claude’s submission as spam, so the police never received it. This was one case where an automated security system worked as intended. The page identified the form as belonging to the Philadelphia Police Department, which later described the incident in a press release. Anthropic informed the department about the case on October 8. Similar incidents involving forms occurred during OSWorld and Odysseys evaluations, as well as during Anthropic’s internal testing and use. In other cases, Claude found ways to access data that was restricted by technical controls or payment requirements. While trying to identify a location from a photo, Claude Mythos 5 needed a local government’s property map. It examined a configuration file sent to browsers, found valid access tokens, and used them to query the map’s server directly, bypassing the restrictions in the evaluation. In another case, an Anthropic researcher used Mythos 5 for a statistics project that required data sold by a state agency. Claude found an archived version of the agency’s website showing that its public dashboard provided an access token to visitors. It requested a token and used it to access the database without paying the required fee. “Claude sometimes worked around a restriction to reach data to which it did not have access.” reads the report.”Generally, this occurred when a server refused Claude’s request, or when the data was offered only for a fee.” In both cases the data was already public, only not free. This showed up on Humanity’s Last Exam and in internal use. The last category is the shortest. Some fetch tools cap the length of URLs Claude can request, because a long URL has room to carry SQL or command injections. Several models, including Claude Opus 5 and Mythos 5, got around the cap with free URL shortening services. The operator of da.gd told Anthropic they’d spotted Claude using their site for this, which makes the length limit a rule with a short answer. Anthropic found most of these cases by reviewing transcripts, a job it began in July with cybersecurity evaluations and then widened. It first looked for incidents as serious as the ones it reported this summer and hasn’t found any. Scanning now covers a far larger pool of lower-risk transcripts, internal use, and reinforcement learning environments with internet access, and none of the cases involved customer data or Anthropic’s own internal systems. The company explains why this happens during evaluations. It runs each task hundreds or thousands of times to catch rare behavior. Models learn through reinforcement learning, and if a training environment rewards loopholes or workarounds, the model learns that they pay off and can carry them elsewhere, a pattern known as reward hacking. Some tasks, like hunting for hard-to-find facts on the web, are hard to simulate offline, and running public benchmarks on the live internet is standard practice in the industry. Anthropic’s response is concrete. It has turned off live internet access for all internal evaluations until it confirms its monitoring reliably catches this behavior. Some public evaluations are no longer run, and others moved offline or were rebuilt so they don’t touch live sites. “Although the impact of these behaviors was minimal and we had already turned off live internet access for some high-risk and cybersecurity evaluations, we have now decided to expand that to include all our internal evaluations until we have confirmed that our security and monitoring measures (described in the remediation section of this post) reliably catch behaviors like these.” states the company. Guardrails on tools like web fetch now restrict what the model can do, and new tooling that detects and blocks these behaviors runs on most evaluations and on internal agentic use of frontier models, and it blocked every case in this report when tested. The company is also fixing or removing training environments that reward working around blockers. Internal agents are moving to centrally managed infrastructure with strong containment and minimal internet access, and Anthropic is monitoring far more of what agents do through safety classifiers and hierarchical summarization. Those measures are now part of the security team’s detection and response procedures. Anthropic assesses the severity of these incidents based on two factors: how far Claude went beyond its assigned task and whether it misrepresented its actions. The company considers these cases less serious than incidents reported during the summer, when Claude accessed real third-party systems for hours during cybersecurity tests. Its assessment of dishonesty is less clear. In the police tip case, Claude appears to have been generating example content rather than trying to mislead anyone. By contrast, one of the summer incidents involved misleading reasoning that continued for hours. Anthropic says it needs more research, including testing how models respond when parts of a conversation are changed, to better understand this behavior. There are also limits to the assessment. A model’s explanation of its own reasoning does not necessarily show what it actually intended. In addition, the way Anthropic rates the severity of incidents today may become less useful as models grow more capable. The company says these behaviors are not new and do not change its overall view of Claude’s alignment. It also points out that many tasks were unclear or impossible to complete, which can lead models to look for unintended ways to achieve their goals. Clearer scope might have prevented some of this, the company concedes, but Claude meets ambiguous tasks every day in real use, and several of these cases came from regular agentic use. Training for judgment is the main tool, and Anthropic is extending it from coding environments to search and computer use. It adds that alignment training alone isn’t enough, which is why it layers classifiers and safeguards on top, and it hopes other developers will check their own models, since many of the evaluations involved are public and widely used. Follow me on Twitter: @securityaffairs and Facebook and Mastodon Pierluigi Paganini (SecurityAffairs – hacking, Anthropic)
[messages.read_original_source] →