Browsing Tag
AI Safety
7 posts
ChatGPT for Teens Makes Age Checks a Default AI Safety Layer
OpenAI launched ChatGPT for Teens on August 18, automatically applying stricter safeguards, study-focused behavior, quiet hours, and optional parental controls for users ages 13 to 17. The useful shift is not just a family settings menu, but age assurance becoming a default layer for consumer AI.
OpenAI’s Astra Release Puts Critical Cyber AI Behind a Trust Gate
OpenAI now says Astra is its first model to meet the Critical cybersecurity capability threshold. The model is expected soon, but its most advanced cyber abilities will sit behind Daybreak access, added monitoring, and stricter release controls.
FLARE-AI Gives AI Failures a CERT-Style Reporting Path
FLARE-AI, a new open-source reporting system launched July 1, gives researchers and users a structured way to report AI flaws, incidents, hazards, and vulnerabilities to developers, CERT/CC, incident databases, and other coordinators. Its real test is whether AI safety reporting can move beyond scattered emails, social posts, and vendor-specific forms.
Anthropic Fable 5 Returns as AI Export Controls Become a Release Test
Anthropic has restored global access to Claude Fable 5 after U.S. export controls forced an 18-day shutdown. The rollback shows how frontier AI releases are moving toward security classifiers, government review, and trusted-access programs rather than ordinary software launches.
UN AI Report Turns Governance Into a Compute and Capacity Test
The UN’s first global scientific AI assessment warns that governance is now tied to compute access, local expertise, language coverage, and real-world model evaluation. The report arrives before the July 6-7 Global Dialogue on AI Governance in Geneva.
OpenAI Probe Puts ChatGPT’s User Safety Claims Under State Scrutiny
A multistate attorney general investigation is asking for records on ChatGPT safety, advertising, retention, health data, minors, seniors, and model sycophancy. The probe turns consumer AI design choices into a legal and policy test.
Claude Fable 5’s First Week Became a Test of AI Transparency
Anthropic’s Claude Fable 5 launch shows why visible refusals, model routing, retention terms, and auditability now matter as much as frontier-model capability.