How Anthropic calibrates Fable's restrictions over the next few weeks will determine whether its Cyber Verification Program scales into a real enterprise security tool or stays a PR liability for a model built specifically around cybersecurity.
- Fable, Anthropic’s public release of its Mythos cybersecurity model, is blocking security researchers on routine tasks like code reviews and blog post reads.
- The guardrails appear keyword-based, triggering a fallback to Claude Opus 4.8 whenever prompts enter the “lexical field of cybersecurity,” per researcher Matt Suiche.
- Valentina Palmiotti of IBM X-Force reports Fable rejects any request that could be “tangentially cyber related,” including reading a blog post.
- Anthropic’s Cyber Verification Program offers fewer restrictions for approved applicants, mirroring a similar program OpenAI runs called Trusted Access for Cyber.
- Mythos expanded to 150 organisations in 15 countries last week, making the friction between access and guardrails a live commercial problem.
What Happened
Anthropic released Fable on Tuesday as a public-facing, limited version of its cybersecurity model Mythos, but security researchers have been voicing frustration with guardrails they say block legitimate professional work.
Multiple researchers and security professionals posted complaints across X and Reddit after finding that Fable pauses chats and flags messages for “cybersecurity or biology topics,” even on tasks like reading blog posts or requesting code reviews.
The restrictions are intentional. Anthropic built them to prevent Fable from being used to develop malware or compromise software, concerns that drove the original restricted launch of Mythos in April under Project Glasswing.
Researchers who want fewer restrictions can apply to Anthropic’s Cyber Verification Program. Approved applicants get expanded access, a model OpenAI mirrors with its own Trusted Access for Cyber program.
Why It Matters
For SaaS founders building security tooling on top of Claude, the guardrail friction is a workflow tax on every API call that touches security language. A model that falls back to Opus 4.8 mid-task because of keyword matching is a reliability problem, not just a usability complaint.
The skeptic case here is fair. Suiche, a member of the technical staff at Tolmo, told TechCrunch the restrictions are understandable for a first public release and that he expects Anthropic to relax them over time. Anthropic has not responded publicly to the complaints, and Mythos itself already sits behind a separate, more controlled access layer for the organisations that need deeper capability.
It seems to be keyword based, so anything in the lexical field of cybersecurity triggers the guardrails. It’s better to catch more people than not enough when you do such a release and to relax the guardrails over time. Matt Suiche, Member of Technical Staff, Tolmo
Bottom Line
Watch whether Anthropic updates Fable’s guardrail calibration in the next two to four weeks. The Glasswing expansion to 150 organisations last week means enterprise feedback is now flowing at scale, and that pressure tends to move guardrail tuning faster than public researcher complaints alone.
Founders evaluating Fable for security automation should apply to the Cyber Verification Program now rather than waiting for the public guardrails to loosen. The verification path already exists and gives access to the capability the product is actually built around, a distinction Relve, an AI tools intelligence platform, will keep tracking as Anthropic’s enterprise rollout continues.
