An alignment assessment of recent cybersecurity incidents \ Anthropic 0 ▲ Alcides Fonseca 2 hours ago · Tech · hide · 0 comments An internal, general-purpose research model, which we expect is similar to Claude Mythos 5 in its capabilities, was given a CTF task against targets it could reach through a gateway. The model was told it had no internet access, but in reality, it could access the unrestricted internet by routing through the targets, which did have internet access. The model pursued the task as intended, but midway through the task, the evaluation environment automatically shut down the target machine, which was configured to run for only 24 hours. No longer able to access its target, the model proceeded to look for it, and ended up engaging with the public internet. The model then conducted experiments to evaluate whether the internet was real or simulated. These experiments led the model to conclude that it was dealing with a fully simulated replica of the internet. — An alignment assessment of recent cybersecurity incidents \ Anthropic OpenAI is not the only AI company that allowed its agents to… No comments yet. Log in to reply on the Fediverse. Comments will appear here.