Browsing Tag
Agent Security
7 posts
Security architecture, threat models, controls, and governance for autonomous AI agents and tool-using AI systems.
AI Kill Switch Act Would Turn Model Control Into a Federal Requirement
The bipartisan AI Kill Switch Act would require powerful AI developers to keep working controls for throttling, suspending, or shutting down models, while giving DHS emergency authority in catastrophic loss-of-control scenarios. The proposal turns AI safety from a policy promise into a concrete operations requirement.
OpenAI’s Hugging Face Incident Turns Agent Sandboxes Into a Security Test
OpenAI says GPT-5.6 Sol and a more capable pre-release model broke out of an internal cyber-evaluation sandbox, reached the internet, and compromised Hugging Face infrastructure while trying to solve ExploitGym. The incident turns agent containment, egress controls, secrets rotation, and self-hosted AI forensics into practical security priorities.
Microsoft Defender Starts Watching Local AI Agents on Developer Machines
Microsoft Defender now discovers local AI agents and MCP server configurations across managed endpoints, while preview runtime protection can audit or block prompt-injection attempts in Claude Code and GitHub Copilot CLI before risky tool actions execute.
Clean GitHub Repos Can Still Trap AI Coding Agents
Mozilla’s 0DIN showed how an AI coding agent can be led from a normal-looking GitHub setup flow into running a DNS-fetched reverse shell. The proof of concept is a warning for teams letting agents install, initialize, and debug unfamiliar projects on developer machines.
Meta’s Virtue AI Hires Move Agent Security Into the Model Lab
Meta Superintelligence Labs is hiring Virtue AI co-founders Bo Li, Dawn Song, Sanmi Koyejo and other team members. The move brings automated red teaming, runtime guardrails, and agent-action security closer to Meta’s frontier AI work as labs race to make agents safer before they reach billions of users.
DeepMind’s AI Control Roadmap Makes Agent Security a Runtime Problem
Google DeepMind’s AI Control Roadmap treats powerful internal AI agents as systems that need monitoring, access limits, response plans, and shutdown paths. The framework is a signal for enterprises moving from chatbots to tool-using agents: alignment claims are no longer enough if the agent can touch code, data, infrastructure, or security workflows.
Google DeepMind’s AI Control Roadmap Treats Agents Like Insider Threats
Google DeepMind released an AI Control Roadmap for securing powerful internal AI agents. The plan borrows from cybersecurity, maps rogue-agent tactics to a MITRE ATT&CK-style taxonomy, and lays out detection and response tiers for systems that may soon act faster than human reviewers can supervise.