Fable and the end of the free lunch
dbreunig.com
dbreunig.com
I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster. Of course, both things continue to happen on a seemingly monthly basis
There's a lot of truth to this. I think we're starting to approach the point where increased intelligence has declining marginal returns, such that it might not even be worthwhile to improve models unless it can be done cheaply.
It was so amazing to get advices and reflect that it struck me : I could use this model forever - it’s clever enough to help me tons and do lot of work for me - even if ai would stop evolving I would love it
And this is inherent to how LLMs work.
Having a natural language interface where you need to go out of your way to specify that you want an accurate answer rather than just a plausible one defeats the entire purpose of it being a natural language interface for normal people. In certain professional contexts, it can be useful, but I don't buy it at all that it makes sense to ask everyone in their everyday lives to go out of their way to specify that they actually want correct answers to their questions.
You don't. It goes in the system prompt.
Take that away and you'll barely be able to make an app that display a pigeon riding a bicycle (or whatever you ppl are doing these days).
All current devices used to run AI are very far from an efficient solution to the problem. What you really want is a pure dataflow architecture, instead of a von Neumann machine. The reason people aren't really making them yet is that when you build one, even if you use SRAM for the weights, you are binding yourself to the dimensions of the model you target -- your chip is only ever going to run variants of that specific model. And SRAM is much more expensive than ROM, so if you want to make a cheap version, you need to design a specific model into silicon.
Once model improvements taper off, the next thing that will happen is everyone will chase speed. There is no physical reason why a mid-sized model could not run at >1 million tokens per second on leading edge silicon, if all computation that can be parallelized, is. No-one will go straight to that, even for a mid-sized model that's like 20 distinct reticle-limited chips. But something like the next version of Taalas HC1 (presumably called HC2?) will probably boost a ~30B parameter model to ten of thousand of tokens+ per second from a single stream within 12 months.
It's highly likely that over the next 10 years we find demand and loss of production further constrains supply.
Designing a specific model into silicon sounds like one of the worst possible ideas. No better way to freeze assumptions and limit growth. Software defined solutions dominate for a reason, because adaptability is key.
Models were never the answer. Eventually we'll get past the nonsense of observationally inefficient neural nets.
Well, I'm here to tell you that whatever is going on behind the scenes at Cursor with this Space-X acquisition in the works, the Auto setting is clearly routing all prompts through "Cursor Grok 4.6 High" right now.
This is a degree of subsidy that makes the Microsoft thing look quaint.
I reduced my $200/month subscription to the $20/month level and have proceeded to do what I would have paid about $1500 to do with Opus 4.7 or thereabouts, which is how Grok 4.6 High feels like it compares. I don't have anything remotely like hard evidence to back this estimate up beyond what I'm watching it do and I still somehow have ~10% of my monthly Auto capacity left on my account. It's completely nuts.
Can't say much more because I have more backlog to run before someone comes to their senses.
However, I've been using Cursor since it came out. I have a massive amount of institutional knowledge locked into their platform, and vastly prefer it to the other options available even if the switching costs were zero.
Going through what would amount to significant effort/time/cost to switch to a different coding tool as a sort of performative political rebuke is just not how I would recommend anyone protest DOGE.
I usually keep 4+ agents churning, many of them on tasks that take hours or day. I only played with Cursor a bit, but it seemed to want input from me every 10 minutes or so.
I am not saying that you're wrong, just very aware that we appear to use LLMs in radically different ways.
Only the most substantial requests run for ten minutes or more, and I would spend an hour or more writing the prompt for that action.
I am always the bottleneck for how quickly Cursor can do what I want, with the level of outcome that I get.
I would genuinely love to be a fly on the wall while you do what you do, because it just doesn't currently make sense to me.
Maybe Fable can do the same things better than other models, but having to tiptoe around to avoid tripping safeguards makes GPT 5.6 so much easier to work with that I don’t even bother with Fable (or Opus 5) now.
It happens to me all the time with things that have nothing to do with security, Fable spawns a subagent that then adversarially checks the code Fable just wrote and hits guardrails, with zero prompting from me.
Like middle school level genetics stuff from a guy who hasn’t been in school for decades.
They need to fix that. It’s just broken. Nobody is making bioweapons if they’re asking the dumb sort of questions I’m asking.
Also, it refused to identify an actor in a popular tv show from a photo. Apparently the policy is it won’t identify ANYONE from a photo, now. Even publicly listed cast members from a very popular show, from a photo of a scene in that show.
It claims that’s a fixed security policy. Nevermind how that makes absolutely no sense… argue about it enough and it terminates the chat.
I don’t know what the Anthropic clown car is even doing anymore, but I won’t be surprised when the others eat their lunch.
There can only be one fix: send Amodei packing and release unguardrailed models.
Having not asked a single security question it will write wildly vulnerable code, go back and fix it, and guardrail itself out of existence after charging me a large sum with no refunds for no output and having not fixed it because that might be secuirty adjacents.
And if it doesn't do this you end up with code that has such holes, store xss , no authz ... if it does not go back and notice it has written bad code.
Since they hide thinking and reasoning from the user (who is also paying for those tokens) it is a black box what is triggering it, has the LLM this time thought of "Oh, this has XSS" and used a bad dangerous word such as XSS, while the previous conversation did not ?
The tendency for absolute inefficiency is effectively unbounded until scarcity is imposed.
People say stuff like this a lot, but I have a different take.
The whole "such-and-such model is 90% as good as Fable at 1/10th the price" assumes that the value increase of intelligence is linear. But I think it's exponential: that last 10% makes a massive amount of difference. It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly (because top models like Fable also possess substantial non-software domain knowledge that help bridge the gap between user and software), or the depth and breadth of engineering expertise that helps avoid a nasty bug that would otherwise have cost you users and revenue.
Yes, it is totally possible to use Fable as the planner and delegate implementation to lesser models. I do that. But, my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by Opus 5, GPT 5.6 Sol, GLM, Qwen, Deepseek, etc. The reason I believe this is because I read the code Fable writes and compare it to code that any other model writes and the difference is night and day. It's not just 10% better. It's mid-level engineer vs. principal/staff-level engineer. And the thing is, even for rote tasks, a more senior engineer is going to be more likely to come up with a clean design than a mid-level engineer. They will also be much more likely to take a step back and ask important questions or propose different approaches.
So if you're using Fable and everyone else is using lesser models, sure they might be saving a lot of money, but there's a higher likelihood that your product will be higher quality, perhaps to a significant extent. And models that are released in the future will benefit from it as well.
In that light, I often go the other way: let Opus (and Haiku subagents) do most of the heavy lifting and then give Fable a shot at finding holes, especially if there are holes or unanswered questions or unearned assertions that I’ve caught on my own in Opus’ output. This, so far, seems like a clean tradeoff that doesn’t burn my Fable credits as hard and still gives solid results.
Really frustrating.
My brother, that's my job.
One of the primary reasons for this is that computers operate in a vast range of orders of magnitude. There’s several orders of magnitude between cache local cpu operation and dram, then several to disk, then several to network, then several to globally durable guarantees. When your code has literally thirteen orders of magnitude to optimize under, there’s never a free lunch. You always need to understand your stuff.
As 80% of enterprise software is CRUD with a bit of sprinkling of user authorization and tenant customisation. But subtly different for every business domain. It's mainly what properties the models and validations have that are different.
When you add a new module or whatever most of the code you have to write is rote code.
And sonnet can handle that crap just fine, you just point it at a similar example in the code, it picks up your userContext convention, how you're doing i18n, etc. and you're done.
I like saying that enterprise code is often shallow but wide. I must have written at least 4 purchase order systems in my career that are all completely different but almost exactly the same.
The Bitter Lesson says that eventually general approaches which leverage more data and more compute will outperform the handcrafted rules and heuristics that humans add in.
However, it does not say what to do today about the problems of today. We can’t just wait around for 10x faster compute and 10x more data.