And by benchmarks (unless they gamed them), seems to be at around Opus 4.7 level, which is what Elon mentioned in https://x.com/elonmusk/status/2074911038286295049.
I guess the Cursor data was very useful.
And by benchmarks (unless they gamed them), seems to be at around Opus 4.7 level, which is what Elon mentioned in https://x.com/elonmusk/status/2074911038286295049.
I guess the Cursor data was very useful.
Above that (max context is 500K) pricing doubles to $4/12.
No longer feels as inexpensive. Will likely just include this in the rolodex of <200k context tasks, like being one of my review agents.
I wish my company gave me more options than just using Claude to test these things out
As always it depends on what you are using them for, and how you are using them.
Anthropic have a fixed price regardless of context usage.
These per-token pricing schemes aren't directly comparable though since these models all use different numbers of tokens, even for input (Anthropic's recent tokenizer change generates 30% more tokens for exact same input), as well as for reasoning, and context/token usage also varies wildly by harness with Claude Code using 3x the context/tokens of Pi.
Does anyone know why they would charge more for higher context usage (other than they can)?
AI serving cost is apparently mostly hardware depreciation rather than operating cost (electricity etc), and if your large context request is occupying VRAM for some fraction of a second then you are paying for the depreciation that occurs in that time!
e. H100 costs $20-40K to buy, with a lifetime of maybe 3 years, and will only consume maybe $2K in electricity if run 24x7 for those 3 years.
To use an analogy: imagine your friend is the author of an unfinished book. They die with 19 chapters written, and on their death bed ask you to write the 20th chapter. Assuming you're up to the task, you can only do this well if you take the time to absorb the entirety of what's been written so far.
This is how LLMs work. Context caching is an optimization on top of this, but it has its limits.
The (pessimistic?) take is that they have loads of idle GPUs and want to get some revenue out of them rather than none. Compare this to OpenAI/Anthropic where every token used by a consumer has to compete with enterprise spenders, and there’s not enough to go around for everyone.
The competitor would have to port their training systems to your specific network architecture, system design, rdma Vs ethernet vs infiniband Vs nvlink etc.
Getting it running might not be too hard, but getting it running efficiently and making good use of all those flops will require considerable human effort and wall time.
Add that to the fact most frontier labs seem to have a single huge training run - and to my knowledge nobody has figured out how to distribute that training run between data centers effectively.
As their models get more competitive I'm sure prices will catch up.
really dont think they have a lot of idle power
His net worth is orders of magnitude bigger than the cumulative profits his companies have ever produced (even if you only count the profitable quarters)
My sister-in-law's mother drove one from Florida to the northeast without touching the steering wheel or pedal/brake, right down to the parking at each end.
I've been a humongous fsd sceptic for a while, but had to lay that aside after I went for a test drive (test ride?) in one of these things.
You will never get a company to accept responsibility for crashes with hardware owned external to the company. I mean there's all sorts of things you could maliciously do to break hardware that you personally own. This is especially the case with something like Tesla where the people who absolutely hate Elon already go to extreme lengths to create news stories to attack him with.
There is just no advantage to the SAE level nonsense. It doesn't even make sense from a technical perspective as vehicles do not function based on "ODDs" (Operational Design Domain) that the SAE levels rely on for their definitions. The SAE levels were created by people who don't understand automated driving vehicles.
> You will never get a company to accept responsibility for crashes with hardware owned external to the company.
This is a completely bullshit argument. For one thing, Mercedes-Benz has already done that. For another, they already have all kinds of liability regarding the proper functioning of the vehicles they manufacture that their customers own, and they manage it very comprehensively. If there is a failure due to a manufacturing defect, do you think they have no ability to determine or litigate the owner's maintenance or improper modifications that may have contributed?
The reason Tesla can't assume these liabilities is because the technology does not support them. It is not safe, and they have no shortage of Elmo stans who will use it anyway and suffer the consequences of their God Emperor's hubris.
Not bad to get a product that underdeliver 8 years late ?
If you listed it, how many features/LOC or vice-versa? Really hard to know if 200K LOC is good or bad, at the surface it sounds like too much, but I don't know what the application was either.
If I had to give advice I’d say to just do it. Work on some project that interests you and go for it
My time is more valuable that I will use a model that doesn’t f** up my code base.
I mentioned here (https://news.ycombinator.com/item?id=48766275) how poorly it handles my specific use cases. My coworkers in DevOps and frontend UI swear by its cost-effectiveness, whereas I strongly prefer the reasoning capabilities of Opus 4.8 and Fable 5.
Composer 2.5 seems to be SOTA for Helm charts and React/Vue, but, for my usecases it absolutely struggles spectacularly when tasked with rigid body dynamics or kinematic logic.
Could you support this statement with an official reference?
However, the fact that they finally have a strong post-training and RL setup bodes well for future releases. They certainly are not compute-constrained anymore.
I wonder how good their subscription discount is on both their subscription types.
Even so, I'm just not that impressed, I felt like I got more done by just using Opus.
If you're not explicit in the prompt or haven't configured your environment then the default behavior is to use subagents that match the host.
It's important to include the reason aka the why of your task [1] in your prompt. You'll get more mileage if you verbalize your thought process when prompting Fable. Anthropic say you should think of Fable as a "thought partner".
1: https://platform.claude.com/docs/en/build-with-claude/prompt...
2: You might find some of the example prompts listed here useful https://x.com/trq212/status/2073100352921215386
Some things require skill to use most effectively. It's fair enough to consider this a failure if the thing in question is "making a phone call", but when it's something like "getting an AI system to do a good job for you" this is not a reasonable thing to make fun of it for.
It's like...
"I wrote a program, and it segfaulted instead of printing out a list of prime numbers." "Yeah, look, you've got an off-by-one error here." "You mean I'm holding it wrong?"
"I'm trying to play the violin and it's making horrible noises." "You want to change your grip on the bow like this, and be more careful in where you put your fingers on the strings to get the right notes, and there's a whole art to how you adjust the speed and pressure and so forth to make it sound good." "You mean, I'm holding it wrong?"
"I'm managing a team, and one of the people on the team doesn't always do the things I tell her to." "Maybe you should sit down with her and see whether somehow your explanations of what you want aren't getting across, or whether she feels like you aren't treating her with the respect and dignity she deserves, or whether she's bored with the work, or etc. etc. etc." "You mean, I'm holding it wrong?"
Yes. In the second case you're literally holding it wrong. Some things don't work as well when you hold them wrong and it's worth some effort to learn to hold them right.
I hold no particular brief for Anthropic. I don't know whether Fable is really much better than Opus or whether the alleged improvements are all just pareidolia or something. But "getting the most out of this immensely complicated thing that's in some ways kinda like another human being can be tricky" doesn't seem to me like an implausible proposition, and if it's really doing something akin to human-like work[1] then it's not unreasonable if you have to approach working with it in something a bit like the ways you approach working with other people.
[1] If it isn't really doing something akin to human-like work, then why are you bothering with it at all?
Simple tasks are simply saturated just like simple benchmarks. There's a level of intelligence where you simply don't need more for some things.
I do wish the subscription had a separate weekly allocation for rare usage.
It may also depend on the workload. At work everything is very domain specific with barely (if any) public training data; both need thorough review and careful hand holding, meanwhile at home Fable is scared of libtorch and falls back to Opus even if it's not touching the ML parts.
Grok is stuck in a difficult place - not the best model at anything, and not the cheapest either. It's hard to make a case for using it on any dimension, even before you factor in the history (I'm not sure suggesting the company uses the model that refers to itself as "MechaHitler" is the way to a promotion).
Noam Brown (OpenAI) "Implications of Large-Scale Test-Time Compute" https://xcancel.com/i/article/2064210146558136827