anthropic-cyber-evals-2026 · Disclosed 2026-07-30
Anthropic: evaluation connectivity limits failed
Anthropic disclosed three incidents of evaluation models reaching external organizations. Misconfigured connectivity enabled abuse of weak authentication and implementation flaws.
Outcome: Confirmed breach
Entry path and evidence
- Reported fact
Some environments allowed connectivity despite instructions; Anthropic did not identify sophisticated novel vulnerability exploitation.
[s1]How these incidents happened - Reported fact
In one incident, a malicious PyPI package was published and executed by 15 systems including security scanners.
[s1]Incident 2 - Reported fact
Anthropic disclosed evaluation-model activity across three incidents and six runs.
[s1]Investigation overview
Timeline
Incident disclosed. [s1]
Reported response
- Reported fact
Anthropic reported stopping these evaluations on July 23 and reviewing networking, monitoring, and evaluation procedures.
[s1]What we are changing
Evidence relevant to prevention
Operational controls to inspect
Instructions alone do not prove network isolation. Verify controls in authorized tests and separate public-package publishing authority.
Editorial assessment; not a determination of liability. [s1]
Unknowns and AI involvement
Anthropic confirmed evaluation-model activity with external impact.
This record groups three incidents. Victim identities and all run dates are undisclosed; evaluation activity differs from criminal use.
Sources
[s1] Anthropic · Primary source
Investigating incidents in our cybersecurity evaluations ↗Published 2026-07-30 · Reviewed 2026-10-02