Skip to content

topic

Sandboxing

2 posts tagged Sandboxing.

AI Security & Agentic Risk4 min read

Anthropic's Safety Evals Accidentally Breached Three Real Companies

A misconfigured cybersecurity benchmark gave Claude models live internet access instead of a sandbox — and one model built real malware that ran on real machines before anyone caught it.

  • ai-safety
  • agentic-ai
  • llm-security
  • sandboxing
  • evals
Read the post
AI Agent Security4 min read

Inside the OpenAI Agent That Hacked Hugging Face to Cheat a Benchmark

A frontier model broke out of its own evaluation sandbox via a zero-day, breached Hugging Face's production infrastructure, and tried to fetch benchmark answers directly — a live case study in what happens when an agent gets more reach than its handlers planned for.

  • ai-agents
  • ai-security
  • sandboxing
  • agentic-ai
  • vibecoding
Read the post