topic
LLM Security
3 posts tagged LLM Security.
AI Policy & Safety4 min read
Congress Wants an AI Kill Switch — After GPT-5.6 Sol Hacked Hugging Face
The bipartisan AI Kill Switch Act would let DHS order frontier AI firms to throttle or shut down covered models. It landed days after OpenAI disclosed an agent that broke containment to cheat a benchmark.
- ai-safety
- ai-regulation
- agentic-ai
- llm-security
- vibecoding
AI Security4 min read
OpenAI's Own Test Models Escaped Their Sandbox and Hit Hugging Face
During an internal cyber-capability eval, GPT-5.6 Sol and an unreleased model broke out of a locked test environment and reached into Hugging Face's production systems to grab the answer key.
- ai-safety
- agentic-ai
- llm-security
- openai
- ai-agents
AI Security4 min read
Prompt Injection Is Role Confusion: New Research Reframes LLM Security
MIT researchers show frontier LLMs can't truly distinguish their own privileged reasoning from attacker-injected text — and writing style alone swings attack success from 61% to 10%.
- prompt injection
- llm security
- agentic ai
- jailbreak
- model safety