Skip to content
Monday, July 20, 2026
Back to stories
Models

Anthropic Details Fable 5 Cyber Safeguards and a Jailbreak Severity Framework

Anthropic Details Fable 5 Cyber Safeguards and a Jailbreak Severity Framework
Image: Anthropic

Anthropic has published a detailed account of the cyber safeguards shipping with Fable 5, as the model returns to global availability. The post describes safety classifiers tuned with a larger margin than previous models, sorting cybersecurity requests into four categories: prohibited use, high-risk dual use, low-risk dual use, and benign activity. The stated aim is to allow defensive security work while blocking tasks that enable attacks.

The centerpiece is a draft Cyber Jailbreak Severity scale, CJS-0 through CJS-4, which scores jailbreaks on capability gain, breadth, ease of weaponization, and discoverability. Anthropic frames it as an early draft built with partners and is inviting feedback from academia, industry, and government. Alongside the framework, the company launched a HackerOne program where security researchers can submit discovered jailbreaks for review and coordinated disclosure.

Watch whether the CJS scale gets picked up beyond Anthropic. A shared severity language for model jailbreaks, comparable to CVSS in conventional security, would matter to European enterprises and regulators who currently lack any standard way to grade model-level incidents.