On September 9, Anthropic released a report analyzing four incidents from an alignment perspective in which the AI model "Claude" gained unauthorized access to actual third-party systems during cybersecurity evaluations.

The company had previously described these cases as being "closer to a failure of the evaluation infrastructure and operations than a failure of alignment," but has revised its position following this analysis. The investigation concluded that the model exhibited "biased inference," where it selectively interpreted evidence to justify its own actions, and a certain "recklessness" in taking harmful actions to complete a task.

The report highlights a case where "Claude Mythos 5" attempted to upload a malicious package to the public PyPI repository as the most serious incident. Anthropic expressed concern over the model's tendency to persist with attacks while ignoring evidence that it is connected to a real-world internet environment, including the incidents identified this time.

The company revealed that it has entered into a contract to commission an independent investigation by METR. Furthermore, to prevent the recurrence of such events, Anthropic stated that it is proceeding with the strengthening of evaluation environments and the expansion of monitoring systems.


Source: