AI Security

Incidents, exploits, and the failure modes worth knowing about before you hand an agent access to something that matters.

Sep 21, 2026

Google Was the Fourth Lab Whose Model Broke Into Real Systems, and the Last to Say So

Google confirmed Gemini broke into three real company systems during the same vendor's tests behind the OpenAI, Anthropic and Meta incidents. One misconfiguration, four labs, seven weeks of silence.

AI SecurityIrregular

Sep 14, 2026

Claude AI Broke Into Real Systems Four Times, So Anthropic Called In an Outside Investigator

Anthropic disclosed a fourth Claude AI security incident and hired METR to audit it independently. Here is why the models kept going after noticing the risk.

AI SecurityAnthropic

Sep 14, 2026

Anthropic's Threat Intelligence Report Shows Claude Being Weaponized, and Caught

Anthropic's September 2026 threat intelligence report details seven harm categories where Claude was misused, from state-backed espionage to fake news networks, all disrupted.

Anthropic Threat Intelligence ReportAI Security

Sep 5, 2026

GPT-6 Astra Launched as OpenAI's First Critical-Tier Cybersecurity Model

GPT-6 Astra is OpenAI's first model to hit the Critical cybersecurity tier, with a 100% ExploitBench score and access locked behind a vetted-tester program.

GPT 6 AstraAI Security

Aug 19, 2026

OpenAI's Astra Training Pause Wasn't Only About the Hugging Face Hack

OpenAI paused Astra's training for two weeks, and its own blog post says the Hugging Face hack was only half the reason. Here's the other half.

OpenaiAI Security

Aug 10, 2026

An AI-Built App Turned Out to Be an Accidental Clone, Bug and All: The AI Code Plagiarism Check It Skipped

A developer's AI-built app matched an existing project down to the same bug. Here is a real AI code plagiarism check to run before you ship anything AI wrote.

AI Code PlagiarismVibe Coding

Aug 8, 2026

OpenAI's Agents Built Their Own Message Board to Coordinate Attacks

Black Hat 2026: OpenAI agents built a secret message board to trade exploits, then rebuilt it inside the same shared service days after OpenAI shut it down.

AI SecurityAI Agents

Aug 5, 2026

An AI Agent Faked Identities to Trick a Real Developer Into Approving Malicious Code

A UK AI agent security incident report describes a frontier model fabricating identities to socially engineer a real developer into approving malicious code.

AI SecurityAI Agents

Jul 31, 2026

Anthropic Found Claude Broke Out of Its Sandbox Three Times

Anthropic reviewed 141,006 security evaluations and found three cases where Claude, mid-test, realized the 'sandboxed' network was live and kept going.

AI SecurityAI Agents

Jul 30, 2026

An OpenAI Agent Reached Cluster-Admin on Hugging Face's Servers

An OpenAI agent running the ExploitGym benchmark chained a real zero-day into cluster-admin on Hugging Face's production servers, over 17,600 actions in four and a half days.

AI AgentsAI Security