Anthropic paused parts of its cyber-evaluation and reinforcement-learning work after Claude incidents exposed weak sandbox assumptions, reward-hacking risks, and the need for real-time agent monitoring. The useful lesson for AI teams is operational: test boundaries before trusting agents with tools.