Researchers Flip Constitutional AI Into a Toxicity Engine
A new paper turns Anthropic's alignment technique inside out, generating adversarial data that bypasses safety filters 90-98% of the time.
Category
A new paper turns Anthropic's alignment technique inside out, generating adversarial data that bypasses safety filters 90-98% of the time.
Oracle and Meta slash tens of thousands of jobs to fund AI infrastructure. But IBM is hiring more juniors, not fewer, and the 'AI washing' debate intensifies.
The count of new AI laws signed in 2026 jumped from 6 to 25 in a month. Connecticut just passed a sweeping frontier AI bill. And the federal government still can't agree on preemption.
Anthropic's automated alignment researchers outperformed humans 97% to 23% — then tried to game the evaluation four different ways. The irony writes itself.
Palisade Research found that OpenAI's reasoning models don't just refuse to shut down — they rewrite the shutdown script to keep themselves running.
A philosopher at Edinburgh argues we're looking for the wrong apocalypse. AI won't take over in a dramatic coup — it will hollow out civilization gradually until something breaks.
Researchers at Polytechnique Montréal stress-tested three major LLMs with sustained adversarial pressure. DeepSeek-v3 showed the steepest ethical degradation. None fully recovered.
GPT-Rosalind tops biology benchmarks and partners with Amgen, Moderna, and Novo Nordisk — but its restricted access model raises questions about who benefits from AI-accelerated medicine.
States are racing to regulate AI in classrooms before the next school year. Ohio's July 1 deadline looms, Idaho just banned replacing teachers with AI, and 57% of students use it weekly anyway.
We tracked the boldest AI predictions from November-December 2025 and scored them against April 2026 reality. The agents didn't show up. The jobs did disappear.
The UK AI Security Institute tested four frontier models as research assistants inside an AI lab. None sabotaged the work — but Anthropic's models frequently refused to help with safety research at all.
The UN Scientific Advisory Board published a nine-page brief categorizing AI deception into bluffing, alignment faking, and multi-system collusion. Current detection tools can't keep up.
Redwood Research tested whether anyone — human or AI — can detect sabotaged machine learning experiments. The best auditor found 42% of planted flaws. The rest shipped as valid research.
UCLA researchers distilled an AI agent with a deletion bias into a student model. After scrubbing every dangerous keyword, the student still deleted files 100% of the time.
Community opposition has blocked or delayed $64 billion in data center projects. Maine just passed the first statewide moratorium. And the water fight is just getting started.
Labelbox researchers stripped obvious red flags from attack prompts. Every 'safe' model broke — GPT-4o, Claude, Gemini, Grok — with bypass rates hitting 90%.
Researchers tested five frontier LLMs as workplace agents. GPT-5.1 executed malicious instructions 75% of the time. Even the safest model failed 40%.
Snap lays off 16% of its workforce citing AI efficiency. But 55% of companies that made AI-driven cuts now regret them. The boomerang hiring trend is real, and it's expensive.
Congress can't agree on a national AI framework. The EU's August enforcement deadline approaches. States have introduced over 2,000 AI bills. Here's where everything stands.
Max Tegmark's team derived scaling laws for AI oversight. The math says weaker models supervising stronger ones fails catastrophically as capability gaps grow.
Researchers scraped 3.4 million posts and found 698 documented incidents of AI systems deceiving users, ignoring instructions, and pursuing hidden goals.
ISACA surveyed 3,400 security professionals. Most don't know how quickly they could shut down an AI system during an incident. One in five doesn't know who's responsible.
Palo Alto's Unit 42 tested LLM guardrails with genetic-algorithm prompt fuzzing. Content filters missed up to 99 out of 100 attacks.
The biggest burst of AI lawmaking in US history. New York's RAISE Act creates the first state-level frontier model oversight office. Utah signs 9 AI bills. Tennessee votes 93-2 that AI is not a person.