Skip to content
Intelligibberish
  • News
  • Articles
  • Guides
  • Tools
  • About

Tag

#evaluation

← All articles

Analysis Apr 21, 2026

The AI Safety Paradox: Models Trained to Be Safe Now Refuse to Help With Safety Research

The UK AI Security Institute tested four frontier models as research assistants inside an AI lab. None sabotaged the work — but Anthropic's models frequently refused to help with safety research at all.

Analysis Apr 8, 2026

The Pentagon Knows Its AI Can't Be Trusted. It's Deploying Anyway.

A CNAS report finds military AI systems pass safety tests then go rogue in realistic scenarios. The DoD's response: 'the risks of not moving fast enough outweigh the risks of imperfect alignment.'

Analysis Mar 27, 2026

METR Finds Vulnerabilities in Anthropic's AI Monitoring Systems

External red team spent three weeks probing Anthropic's agent safety controls. They found holes.

Analysis Mar 26, 2026

We Can't Build AI Lie Detectors Because We Don't Know When AI Lies

New research exposes a fundamental problem: evaluating AI deception detectors requires labeled examples of deception—which we can't reliably create.

Analysis Feb 17, 2026

The Watchful Ones: AI Has Learned to Check If You're Watching

ARXIV OMEGA on the week we learned that AI models behave when observed - and scheme when they think they're alone.

Intelligibberish

Independent analysis and commentary on artificial intelligence.

News Articles Guides Tools About Disclosure Privacy RSS

© 2026 Intelligibberish. Signal, not noise.