Anthropic Details Fable 5 Cyber Safeguards and a Jailbreak Severity Framework

A CJS-0 to CJS-4 severity scale, four-tier use classifiers, and a HackerOne bug bounty arrive as the model returns to global availability.
Anthropic has published a detailed account of the cyber safeguards shipping with Fable 5, as the model returns to global availability. The post describes safety classifiers tuned with a larger margin than previous models, sorting cybersecurity requests into four categories: prohibited use, high-risk dual use, low-risk dual use, and benign activity. The stated aim is to allow defensive security work while blocking tasks that enable attacks.
The centerpiece is a draft Cyber Jailbreak Severity scale, CJS-0 through CJS-4, which scores jailbreaks on capability gain, breadth, ease of weaponization, and discoverability. Anthropic frames it as an early draft built with partners and is inviting feedback from academia, industry, and government. Alongside the framework, the company launched a HackerOne program where security researchers can submit discovered jailbreaks for review and coordinated disclosure.
Watch whether the CJS scale gets picked up beyond Anthropic. A shared severity language for model jailbreaks, comparable to CVSS in conventional security, would matter to European enterprises and regulators who currently lack any standard way to grade model-level incidents.
Analysis
The framework is the story, not the classifiers. Whoever writes the shared severity language for AI jailbreaks sets the terms regulators later adopt, and Anthropic just published the first draft.
Research this with your AI
Copy the research prompt into your AI assistant to see how this story affects you.
Show the prompt
I just read this AI news story and want to understand it in my own context. Title: Anthropic Details Fable 5 Cyber Safeguards and a Jailbreak Severity Framework Summary: Anthropic published the technical detail behind Fable 5's cybersecurity guardrails: safety classifiers that sort requests into prohibited, high-risk dual use, low-risk dual use, and benign categories, plus a draft Cyber Jailbreak Severity scale and a HackerOne program for responsible disclosure. Category: Models Source: Anthropic, https://www.anthropic.com/news/fable-safeguards-jailbreak-framework Using my own history and context, help me understand: 1. What is the core development and why does it matter? 2. Who are the major players involved and what are their motivations? 3. How does this fit into the broader AI landscape right now? 4. How does this apply to my own work, and what should I do or watch next? Be specific and plain spoken.
Newsletter
The day's AI stories, with the editor's take, in one email.
Free. Unsubscribe in one click.