
Anthropic Details Fable 5 Cyber Safeguards and a Jailbreak Severity Framework
Anthropic published the technical detail behind Fable 5's cybersecurity guardrails: safety classifiers that sort requests into prohibited, high-risk dual use, low-risk dual use, and benign categories, plus a draft Cyber Jailbreak Severity scale and a HackerOne program for responsible disclosure.
Source: Read full story at Anthropic
In-house analysis
The framework is the story, not the classifiers. Whoever writes the shared severity language for AI jailbreaks sets the terms regulators later adopt, and Anthropic just published the first draft.

