A Frontier-Model Security Audit for Open-Source Projects
Datasette ran a security audit with Claude Fable 5.1, GPT-5.6 Sol, and GPT-6 Astra. The workflow matters as much as the bugs.
Tag
Datasette ran a security audit with Claude Fable 5.1, GPT-5.6 Sol, and GPT-6 Astra. The workflow matters as much as the bugs.
Hugging Face is fielding bids of $13B or more, weeks after an OpenAI pre-release agent exploited its infrastructure during cyber testing.
None of the mainstream inference servers ask for a password, and vLLM binds every interface by default. The safe pattern for household serving.
Anthropic, OpenAI, and Google ship encrypted reasoning blocks that a sibling model can bulk-decode for about $720, leaking PII and API keys.
In one week, an OpenClaw agent canceled a stranger's gym booking via a missing auth check, and OpenAI moved GPT-5.6-Cyber behind a partner-only Red tier.
Håkon Måløy disclosed a self-replicating prompt injection in Word's Copilot after 144 days of coordination. Two mitigations did not close it.
Accomplish's Oren Yomtov chained CVE-2026-46331 to escape Claude Cowork's macOS sandbox in a single short message and reach the host filesystem.
Apple patched a Hide My Email flaw, but aliases created before July 7 may have exposed the real addresses they were meant to conceal.
A Go-based botnet is scanning exposed Ollama, ComfyUI, n8n, Open WebUI, Langflow, and Gradio instances for AWS keys and Kubernetes tokens, QiAnXin XLab says.
Hugging Face disclosed a July 2026 breach run end-to-end by an autonomous agent. The defender was an open-weight model.
Anthropic's web_fetch tool let a prompt-injection honeypot walk Claude through user profile URLs and pull a user's name, city, and employer.
Anthropic's Mythos finds 10,000+ critical vulnerabilities but fewer than 1% get patched. Mandiant says exploits now arrive a week before fixes.
Intruder scanned 2 million hosts and found 1 million exposed AI services with no authentication. Plus: teenagers are using ChatGPT to hack governments, and OpenAI launches Daybreak.
Anthropic built an AI that finds zero-days autonomously. The Pentagon wants it. Anthropic said no to surveillance. Now it's a geopolitical crisis.
OpenClaw collected nine CVEs in four days with 135,000 instances exposed. Plus: GitHub RCE, Flowise exploitation, and CrewAI trust failures.
Shadow AI isn't a rogue employee problem. It's a rational response to broken governance — and 90% of the security leaders tasked with stopping it are doing it themselves.
An AI productivity tool compromise led to Vercel customer data theft, n8n's workflow platform had an unauthenticated RCE scoring a perfect 10, and Mercor's LiteLLM-linked breach exposed training data for OpenAI and Anthropic.
A vibe-coding platform exposed every project's secrets through a trivial API flaw, Anthropic's MCP protocol enables remote code execution across 200,000 servers, and NIST can't keep up with AI-driven vulnerability discovery.
A third-party AI tool compromise chains into Vercel's systems, North Korean hackers use Dependabot to distribute malware to 895 repos, and courts fine lawyers $145K for AI hallucinations in Q1 alone.
Researchers tested five frontier LLMs as workplace agents. GPT-5.1 executed malicious instructions 75% of the time. Even the safest model failed 40%.
A supply chain attack exposes 40,000 AI contractors, three major workflow platforms get critical RCE flaws, and Microsoft patches 167 vulnerabilities as AI-driven discovery triples submission rates.
ISACA surveyed 3,400 security professionals. Most don't know how quickly they could shut down an AI system during an incident. One in five doesn't know who's responsible.
Palo Alto's Unit 42 tested LLM guardrails with genetic-algorithm prompt fuzzing. Content filters missed up to 99 out of 100 attacks.
Trend Micro confirms the sockpuppeting attack bypasses ChatGPT, Claude, and Gemini using a basic API feature. Some providers have patched it. Most haven't.