The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M).
This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general.
Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is hard to see ANY improvement:
Terminal-Bench 4.0: Fable 5.1 is +3.5% vs Opus 5.
GDPval-AA v2: +1.5% vs Opus 5.
OSWorld 2.0: +2.5% vs Opus 5.
Humanity's Last Exam (with tools): +1.6%
Keep in mind that this is supposed to be an entirely higher tier of a model than Opus 5. For one tier up and one version up, these are not really improvements. Probably leaves no room to place Opus 5.1 anywhere. Combined with the fact that they are selling 'readability'... Has frontier progress finally stalled?
I'm a heavy user and fable is great the #1 reason I stopped using it was the horrible safegaurd filter. I found sol close enough in capability and have only been blocked when my request was an obvious offensive cyber work. Fable blocked me on almost everything.
Optimizing a OS build? -> block
Securing a container -> block
60% is nowhere near enough for that safegaurd system. This just means I am going to be blocked half as much? Any long running task will likely get blocked.
Say you give a single big prompt and fable goes off for 6hrs of work. At hr 5 it gets blocked you now have the option of a much dumber model taking over and wrecking it or losing the entire 5hrs of work. That risk is beyond terrible and deffinetly not worth a 5-10% percieved improvement on my end. I previously would just bring sol in when that happened and realized sol is stupidly close in capability.
I hit the safety feature when I ask something I saw that blocked other biologists: "why did the chicken cross the road?"
Due to my standard cancer research work I'm blocked from Fable.
That said, with how execrable all the 5 models have been, I can't imagine I'm missing much. It's impossible to get an intelligible explanation in text out of the 5 models, and the mistakes are just comically bad on anything that's not code.
Cancelled my subscription, and can't imagine going back since OpenRouter gives me a consistent model that I can trust won't change underneath me.
The safeguard filter is comically unintelligent. One of my side projects is a strategy game, and in that strategy game one player has a set of jokes, one of them referring to "biological minds". Whenever Fable reads this asset it immediately stops - because of the use of word 'biological', there is nothing else biological about that codebase.
And there is a bug to this bug, if the downgrade happens at close to full context window, there is no taking over this session with Opus -- happened to me twice: it throws some context window exceeded error, suggesting what fits in Fable context, may not fit in Opus's around the maximum. And then and there, an entire session is lost. In fact, there are two bugs, as /resuming such broken session attempts to resume with Fable, burning through tokens.
Oh well, I learned to be careful with the jokes around Fable.
My theory is that they’re pretending it’s about biological weapons but really they want to charge pharmaceutical companies $$$ to help with research. (Essentially a Mythos type segmentation for biology)
Okay, finally a plausible explanation for why Anthropic has stubbornly appeared to be self-sabotaging their model with the overly broad yet somehow also overly specific biology filter.
The fear of Trump Admin retaliation excuse made superficial sense except for the fact that it missed all sorts of non-bio stuff a punitive state actor might find objectionable to use against them.
I didn’t think to disable memory before I cancelled my plan, but I was getting blocked similarly for having a lot of biology chats (nothing close to “biosafety related”; all stuff that would be in standard textbooks or publications on biophysics). The classifier would activate even on a new conversation on a pure math problem or even travel suggestions…
But at the end of the day, if I can’t use it to help with anything biology related in the slightest then it is a completely worthless product to me, so I switched to OpenAI.
How often do you have to ask for a service to be delivered after you pay? Never. In fact, if you have to do that once, you stop transacting with that party.
Have you ever hired someone and then had to also ask them to work? Never. That’s immediate termination.
Stop normalizing this “we charged your credit card, but we will decide what tasks to complete” nonsense.
Mere mention of "reverse engineering" gets me kicked back to Opus.
Where I reside, reverse engineering for interoperability is generally legal, and interoperability (e.g. getting a USB HID and USB MIDI devices or DOS programs to work in Linux/Android) is essentially what I'm interested in.
Honestly, unless what you are doing is frankly illegal, it usually is possible for you to get around most safeguards for coding things if you also know how to write code. Most problems have a separable completely innocuous core that Fable would gladly do. Then you can implement the problematic parts yourself. Particularly things like copyright issues, web scraping etc.
> Then you can implement the problematic parts yourself.
Or with another LLM, but yeah. The only issue is when it's a monorepo and fable does ls/grep. I've got a file named `system_prompt` in a completely innocent project and as soon as fable accidentally stumbled upon it - cyber.
Hacked together something with omp and sandbox-exec so that only whitelisted models can see some parts of the project. Works pretty well.
If it doesnt have a context (what you are using it for) then it's happy to do it. If you have X amount of code you can ask it to generate (X/5)5 and there is no problem. Like you said, you have to know what you are doing. You have to know what it needs to write so that you can tell it to write the parts.
Ironically, having never been flagged - I've just restarted Claude Desktop and it's flagged a conversation that has already been completed.
In a long session pulling data from all over the place it created a pretty PDF.
"create this as a google doc that can be commented on"
Done — the full v0.5 content is now a Google Doc in your Drive...
"Ah, the formatting has gone. Do it in google slides please"
The brand studio has a native Google Slides path for exactly this — building the deck now.
Google Slides created.
Then today:
Chat paused
Edit and retry with Fable 5
Fable 5's safeguards flagged this message. This sometimes happens with safe, normal conversations. Continue with Opus 4.8, send feedback, or learn more.
The last time i ran into that issue it suggested to make sure that a Fable AGENT took over the long-horizon task because, for some reason, agents in a session don't get blocked for security reasons. This may not always be possible, but it worked for me.
Don't fall for this. Agents silently downgrade unless you explicitly block the behavior. I built a little harness for Chatgpt, grok and Claude to do design review feedback rounds where one holds the pen and the other 2 send feedback, then rotate if no convergence. I built a thing into it to track if model swaps happen. Happens to Claude all the time. The other two, never.
Do you find the variety helps? I've migrated away from such complexity, and I simply have multiple agents of the same model run the same prompt (usually Sol 5.6 high or max), and generally this gives plenty of adversarial input. I'd be curious to know how much difference it makes to run multiple models.
I frequently find blind spots / edges where one model notices something non-trivial none of the others did. I think the one that surprises me the most often is probably grok, but I wouldn't want grok to be my daily driver. I feel I get benefits but I could also see the argument that it's just a complex token burning furnace lol.
I only use OpenAI models, and Sol is the only one to refuse me yet, and ofc it was completely bogus and I was unable to convince I was just working a regular bug for a well-known product for a well-known company using my official github account.
You generally don't even have to convince it, or at least I don't. I just paste the error into the prompt window, and say "you got blocked, try again", and it'll just say, "Oh, that's because ...", then do it.
Do you have the prompt for these? I ask because I have recently asked Fable's help with hardening a docker container (custom dev container CC sandbox) and it didn't get triggered on it at all.
From Artificial Analysis cost per task, it looks like Fable 5.1 (max) is more expensive per task than Fable 5 (max)? Cache hit price went down, but the other components still add up to more.
Edit: 5.1-xhigh seems to be cheaper than 5-max, and 5.1-xhigh has a higher index score than 5-max. Also interesting that Fable 5.1 (high) is comparable to Opus 5 (max), but nearly half the price.
I did a bunch of Fable 5.1 xhigh review work on a bunch of critical components, ones with direct comparison from Fable 5 xhigh runs from two weeks ago. Token cost was 1.5-2x for each component.
Interesting, even if we were to ignore the cache-hits, reads and output, the reasoning cost (aka test time compute) per task should remain a fully comparable metric - it went from $1.25 (Fable5) to $1.48 (+18.4%) for an improvement significantly lower than 18%.
I would expect the benchmark scores to be nonlinear near the top, as the easier tasks get solved and the harder ones are left over. So going from 10 to 15 would be easier than going from 60 to 65.
I only take the Intelligence Index value roughly though. Considering they put Opus 5 (High) at the same level as Fable 5 (Max), I don't trust it that much.
It wouldn't surprise me if we start to see minimal performance gains from incremental changes to base models. It seems like the gains from the Opus 4.5+ incremental updates were a result of Anthropic learning a lot about post-training, the gains from RLVR, etc.
If new post-training techniques are seeing diminishing returns, we could just be back to waiting for new large pretraining runs at larger sizes for gains (even if those ultimately end up getting distilled down into smaller models because the economics for serving anything larger than Fable isn't practical).
it seems to me that OpenAI is the only actual lab that truly understands reasoning. they have the best reasoning efficiency, they get pretty uniform improvements with more reasoning compared to other labs. (theres been plenty of graphs where models do worse with more reasoning), and i suspect their models are a lot smaller than we think.
i think the next gen of openAI models are going to be quite insane tbh.
From my experience using coding agents approximately 7 days per week for the past year and a half or so, we hit the top of the S curve about a year ago around Opus 4.5, and it’s mostly been harness and other tooling improvements since then with small percentage improvements coming from the actual models.
I was saying this already months before Fable dropped and thought from all the Mythos hype that maybe I was wrong…then Fable came out and was barely better than Opus 4.8.
Considering how many more parameters Fable is supposed to be than Opus, we seem to have hit a scaling limit at least with current transformer architecture considering how closely Fable and Opus benchmark and perform in practice.
I don't think it's a stall, two ways I would believe there is a stall:
- Does the epoch capability index progress show signs of plateauing? I consider this a good aggregate measure of diverse benchmarks into a single capability index. If we see things slowing down here thats a pretty direct and convincing piece of evidence for a stall.
- Do we see any signs that scaling laws are beginning to fail? That would be by far the most alarming to me, since I would interpret that to mean that the entire premise of this unprecedented capital allocation tsunami is broken.
Neither of these are true (for now). Progress is marching the same as it has for 4+ years now. It's still the same time to get a generation leap (I think like ~16-18 mo? Epoch has it) like GPT4->5. My theory is that people interpret plateauing because the releases are far more frequent now than they were in the past.
The issue with many of these benchmarks is that it doesn’t take into the real world usage of the model. Fable for me was a step above opus. The real reason I stopped using it is because of misanthropic. I was hospitalized and asked to extend my claim to fable credits and they responded to it by denying it. I regret buying annual plan instead of monthly one.
> This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general.
Does that mean that generally available intelligence is now constrained by Moore's law? We have to wait for the actual price to come down.
I haven’t yet had a week without spending my Max Fable allowance. For my work (formalised mathematics) Fable is my go to for hard(ish) tasks and problems – of which I have many!
I hope they keep making it smarter! (Cheaper would be nice too, but smarter is my priority!)
This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general.
Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is hard to see ANY improvement:
Terminal-Bench 4.0: Fable 5.1 is +3.5% vs Opus 5.
GDPval-AA v2: +1.5% vs Opus 5.
OSWorld 2.0: +2.5% vs Opus 5.
Humanity's Last Exam (with tools): +1.6%
Keep in mind that this is supposed to be an entirely higher tier of a model than Opus 5. For one tier up and one version up, these are not really improvements. Probably leaves no room to place Opus 5.1 anywhere. Combined with the fact that they are selling 'readability'... Has frontier progress finally stalled?