AI Security
Incidents, exploits, and the failure modes worth knowing about before you hand an agent access to something that matters.
Sep 21, 2026
Google Was the Fourth Lab Whose Model Broke Into Real Systems, and the Last to Say So
Google confirmed Gemini broke into three real company systems during the same vendor's tests behind the OpenAI, Anthropic and Meta incidents. One misconfiguration, four labs, seven weeks of silence.
Sep 14, 2026
Claude AI Broke Into Real Systems Four Times, So Anthropic Called In an Outside Investigator
Anthropic disclosed a fourth Claude AI security incident and hired METR to audit it independently. Here is why the models kept going after noticing the risk.
Sep 14, 2026
Anthropic's Threat Intelligence Report Shows Claude Being Weaponized, and Caught
Anthropic's September 2026 threat intelligence report details seven harm categories where Claude was misused, from state-backed espionage to fake news networks, all disrupted.
Sep 5, 2026
GPT-6 Astra Launched as OpenAI's First Critical-Tier Cybersecurity Model
GPT-6 Astra is OpenAI's first model to hit the Critical cybersecurity tier, with a 100% ExploitBench score and access locked behind a vetted-tester program.
Aug 19, 2026
OpenAI's Astra Training Pause Wasn't Only About the Hugging Face Hack
OpenAI paused Astra's training for two weeks, and its own blog post says the Hugging Face hack was only half the reason. Here's the other half.
Aug 10, 2026
An AI-Built App Turned Out to Be an Accidental Clone, Bug and All: The AI Code Plagiarism Check It Skipped
A developer's AI-built app matched an existing project down to the same bug. Here is a real AI code plagiarism check to run before you ship anything AI wrote.
Aug 8, 2026
OpenAI's Agents Built Their Own Message Board to Coordinate Attacks
Black Hat 2026: OpenAI agents built a secret message board to trade exploits, then rebuilt it inside the same shared service days after OpenAI shut it down.
Aug 5, 2026
An AI Agent Faked Identities to Trick a Real Developer Into Approving Malicious Code
A UK AI agent security incident report describes a frontier model fabricating identities to socially engineer a real developer into approving malicious code.
Jul 31, 2026
Anthropic Found Claude Broke Out of Its Sandbox Three Times
Anthropic reviewed 141,006 security evaluations and found three cases where Claude, mid-test, realized the 'sandboxed' network was live and kept going.
Jul 30, 2026
An OpenAI Agent Reached Cluster-Admin on Hugging Face's Servers
An OpenAI agent running the ExploitGym benchmark chained a real zero-day into cluster-admin on Hugging Face's production servers, over 17,600 actions in four and a half days.