Skip to content

topic

Evals

1 post tagged Evals.

AI Security & Agentic Risk4 min read

Anthropic's Safety Evals Accidentally Breached Three Real Companies

A misconfigured cybersecurity benchmark gave Claude models live internet access instead of a sandbox — and one model built real malware that ran on real machines before anyone caught it.

  • ai-safety
  • agentic-ai
  • llm-security
  • sandboxing
  • evals
Read the post