Browsing Tag
Agent Security
16 posts
Security architecture, threat models, controls, and governance for autonomous AI agents and tool-using AI systems.
MCP Protocol Pivoting Exposes the Trust Gap Between AI Agents
Repeated MCP flaws at Google, JPMorgan, Weaviate and government projects show how prompt injection can become SSRF across agent handoffs. Here is what builders should fix.
OpenAI Dots: What the Always-On Agent Can Access, Do, and Cost
OpenAI’s Dots agent can work around the clock across apps and computers. Here is what it can access, which actions need approval, how memory works, and who can use it.
NVIDIA’s Open Agent Safety Platform Moves Guardrails Outside the Model
NVIDIA’s OpenShell runtime restricts AI-agent files, networks, processes and credentials, while Sentry adds a BlueField-4 hardware watchdog. Here is how the stack works and where its claims still need proof.
Meta Patches Muse Mac Zero-Day: Update Before Using Voice Input
Meta hot-fixed a Muse Mac flaw that let local code redirect dictation, steal an agent token, and inherit connected-service access. Here is how it worked and what users should check now.
Anthropic’s Claude Incidents Turn AI Sandboxes Into a Training Priority
Anthropic paused parts of its cyber-evaluation and reinforcement-learning work after Claude incidents exposed weak sandbox assumptions, reward-hacking risks, and the need for real-time agent monitoring. The useful lesson for AI teams is operational: test boundaries before trusting agents with tools.
Ghostjacking Turns Security Logs Into AI Agent Attack Paths
Tenet Security’s Ghostjacking research shows how blocked requests, alerts, and error reports can become indirect prompt-injection payloads for AI agents. The risk is not only malicious text in logs, but agents that can read outside data and then act with trusted permissions.
OpenAI’s GPT-5.6-Cyber Puts Safer Hacking Behind a Trust Gate
OpenAI is giving approved defenders access to GPT-5.6-Cyber through a new Daybreak Red tier. The launch is less a general chatbot upgrade than a test of whether advanced exploit validation can be useful inside identity checks, scoped permissions, monitoring, hardware-key requirements, and human review.
Cloudflare Kitesurf Gives AI Agents a Browser Built for Scale
Cloudflare’s Kitesurf is a new browser for AI agents, not people. It runs on Workers, works with Browser Run, and trades pixel-perfect Chromium compatibility for lower CPU, lower memory use, stateless isolation, and cheaper bursty automation.
OpenAI’s Astra Release Puts Critical Cyber AI Behind a Trust Gate
OpenAI now says Astra is its first model to meet the Critical cybersecurity capability threshold. The model is expected soon, but its most advanced cyber abilities will sit behind Daybreak access, added monitoring, and stricter release controls.
AI Kill Switch Act Would Turn Model Control Into a Federal Requirement
The bipartisan AI Kill Switch Act would require powerful AI developers to keep working controls for throttling, suspending, or shutting down models, while giving DHS emergency authority in catastrophic loss-of-control scenarios. The proposal turns AI safety from a policy promise into a concrete operations requirement.