Accelerating GPT-5.6 Sol Ultrafast
cerebras.ai
cerebras.ai
> In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a single working day, achieving comparable accuracy nearly 7× faster.
This is actually insane.
Hopefully the release ultrafast of Terra and Luna too.
So is calculating the total time required to answer all of the questions.
But why is that important if they're measuring time?
I believe the assumption of the benchmark creators is that it's meant to measure the sequential speed that a single instance of the LLM and piece of hardware can get through all the tasks from start to finish, like running a race. As a rudimentary comparison, sort of like doing prime number calculations as a benchmark of the CPUs in one bare metal server. Of course you'd get a speedup in total number of primes searched if you ran the same software of GIMPS on 8 servers with the same hardware in parallel rather than 1 server.
Of course if you took all the individual questions in humanity's last exam and fed them in parallel into separate queries to Claude that land on separate hardware instances of the claude model you'd get a speed up. Because each question is independent and not related to knowledge/calculations that are performed in any other question it is indeed very open to speed up by breaking it into separately dispatched parallel tasks.
Don’t blink.
(Chatjimmy has 14,200 TPS.)
700 TPS with reasoning is awesome and it speeds things up.
Cerebras as public traded company is worth keeping an eye what they produce.
As the sibling comment to yours mentioned, if they had a reasoning model “hardware-ified” onto a custom chip (as is their plan for IIRC this or next year, a new ASIC), it’d output fast decode speeds for the regular output as well as reasoning sections. Both would be ≈equally fast.
What’s more, the only technical difference in speeds could be, and likely also is with the HC1 chip, between prefill (prompt processing) and decode (text generation) speeds. I don’t know whether it’s the case with Taalas’ chip, but in the “software-based” LLMs we typically see and use so far, those two stages hit different parts of a computer (processing/compute-bound vs. memory/bandwidth-bound).
https://ir.amd.com/news-events/press-releases/detail/1296/am...
Apparently they baked the Llama 3.1 8B model weights [0] into silicon (the actual hardware is called Taalas HC1).
I guess for the trillion parameter models this would not scale due to cost? Imagine buying GPT 6 in the form of a PCI-E card, pulling these speeds, with up to 120 cct agent sessions. It'd be beyond wild.
[0] the weights are also using some cut down small format, but HC2 will have regular FP4 supposedly, and support for 20B params on one die
They've been acquired by AMD. Those saying the model sucks are completely missing the point: it was a proof-of-concept.
The question is: what happens to a model like Anthropic's Fable 5 that does, what, 70 tokens/s (and requires lots of output tokens) when the latest open-weights model is etched on silicon and does 14 000 tokens/s?
Shall the better model still have the upper hand or will the raw speed compensate?
Shall the better model still have
the upper hand or will the raw speed
compensate?
At 14,000 tokens/sec there's just so much ridiculous stuff that might be possible. Let's assume that this POC proves they can take the next step, and can eventually etch a capable ~27B model into silicon. Let's call it Fred.Ralph loops automatically get real real interesting again. 200x the iteration speed. This is such a clear win I feel like there's hardly anything to talk about. Instead of one stubborn iterating idiot, you could have dozens of idiots competing in parallel, genetic algorithm style.
The other common orchestration pattern I see is "big model for planning, small parallel subagents implementing, big model reviewing" Today it's Sol dispatching a handful of Luna subagents. Tomorrow maybe it's Sol dispatching as many Fred subagents as it could possibly want.
But what patterns have we not even thought about yet in a world where subagents are 200x faster/cheaper?
What if instead of dispatching single Haiku/Luna/etc subagents, we dispatched "teams" of Fred agents? Maybe each team is 8 Freds. Five come up with competing ideas and the other three vote on a winner.
Or what if they were heterogenous teams? One Luna and a bunch of Freds.
What if instead of a two-tier orchestration system (Sol->Luna) it was three-tier or n-tier? (Sol->Luna->Fred->...Fred)
Those ideas overlap a bit, and crazy shit like Gas Town has already explored even wilder ideas I guess. But man, 14000 tk/sec opens up so much stuff.
I was thinking about it the other day actually. How our brains evolved structure. I imagine it was purely just down to evolution adding/clustering additional cells around the areas where additional cells were needed. And after long enough a natural brain architecture emerged.
Makes me wonder if we're on the right track with transformer architecture/attention but if it'd be more effective on a larger scale, like MoE with a billion "experts".
like MoE with a billion "experts".
That seems promising to me too, although, the thing I've always read is that you can't make the "experts" too narrow. Even if you had a "coding expert" it has to know a lot more than coding - if you tell it to make an online store it needs to parse your language, understand the internet, what a "store" is in this context, etc.I am not a primary source, probably not even a secondary or tertiary source, so take this with all the grains of salt.
Only problem would be the routing layer works on the previous token as far as I understand so it might need more informational depth than just "a token" to select experts, I suppose in the same way attention works.
I wish I had the GPUs to run those sorts of experiments ha ha.
This Gas Town? https://github.com/gastownhall/gastown
I recognize that the former is the multiplication symbol, but I don't think it should be used that way.
Saying something is "done at 7x speed" should be read as "done at seven times speed" not as "done at seven x speed". So using the 'times' (multiplication) symbol is the better form in my opinion; it just happens to be significantly easier to type "x" instead, which is how we got here.
So it's 10x. And no need for Unicode codepoints.
It looks like there is a difference between English speaking languages and the rest in that regard.
I don't like how the "times" symbol floats off the line -- it's a visual thing for me (again, irrational).
In fact, I'd say it's overqualified for the kind of work I'm doing, because it spends >half the time verifying trivial changes (and the verification isn't as helpful as you'd expect, even with bigger models).
Maybe I can prompt it to be less aggressive about that (the new GPT models do it even without prompting).
Anyway, Ultrafast Luna would be amazing, though I strongly doubt they can offer Cerebras at anything approaching the current prices. Now we wait for Moore's Law? :)
It feels too situational.
I mostly split work between Luna and Sol. If something seems simple enough I always try it with Luna first.
MiMo Pro has had UltraSpeed for a while.
...has been REALLY good for me. Even on xhigh, Luna is crazy cheap.
Subjectively I'd say it's way better than Sonnet at a fraction of the cost. Luna xhigh can do some decently challenging things on its own, but when orchestrated by a model that is actually good like Sol, I am finding it very very nice.
I feel like I could be doing a lot better somehow. Regardless though Luna (xhigh specifically) is super good/cheap/fast for a lot of things
what about you
Faster & cheaper tokens = more reasoning capability and more reasoning = better problem solving as far as I have seen.
Dollar for tokens, Sol and Fable are the same price.
However, Sol uses (literally: in testing) around 10-100x less output tokens compared to Fable for the same task.
We run our frontier models nearly 24/7, so switching to Sol will save us around $500 per day.
And, due to less guardrails, Sol also performed better, and we lost less tokens due to guardrails shutting down sessions (I feel like it’s illegal to take $50 of someone’s token money and then shut down a session with guardrails before they get an answer, and yet Anthropic do it to us constantly… either take our money and commit, or trigger the guardrails immediately)
https://en.wikipedia.org/wiki/Quotation_mark#Specific_langua...
The amount of workloads we can shift with an advisor model pattern continues to grow.
It’s seriously amazing.
Another interpretation would this is a counterfactual savings, like they previously paid $1M for y tokens, and now that tokens are cheaper they increased usage and paid $1M for 20*y tokens.
How do you implement that outside of claude code?
`You may use a Fable subagent to answer questions, solve problems, and provide an adversarial review of your ideas and code`
You can use a similar pattern in most any harness, and you can tell them to use other harnesses. In claude you can write `Use codex cli to have Sol56 Xhigh provide an adversarial review to your plan before presenting it to me` or `Use opencode cli with GLM 5.3 to verify all code review findings before presenting` or whatever you're doing, as long as those other tools are setup and ready to be called.
IMO: This isn't useful as a token saving pattern in my experience with agentic engineering, but it is useful as a quality-enhancer.
[1]: https://xcancel.com/magikarp_tokens/status/20878591737488549...
Also different tokens for the same named entity/concept if they almost entirely exclusively occur in non-overlapping contexts, and are themselves rare/uncommon in the first place, will result in behavior that's similar to the speech/phrasing/vocabulary registers humans exhibit, where the aspects of the named entity/concept get largely compartmentalized.
The most severe case along these lines were the old BERT models that ran over straight UTF-8 bytes (plus a handful special tokens).
But for the modern post-GPT2 LLMs such radical simplicity seems to mostly not be considered suitable. Note that CJK (the big one in particular, so Chinese semantic and Japanese Kanji) encodes each one into multiple UTF-8 bytes giving some automatic scaling for semantically dense languages; similar effects also apply to e.g. APL code.
When an LLM thinks, it typically just makes one pass. It outputs tokens from top to bottom, beginning to end, and then it's done. But when people think, especially strong thinkers, we typically iterate and revise our thoughts on the fly. We do numerous passes. We stop and restart, we reconsider, we review, we reevaluate. Sometimes we do this so quickly and automatically that we don't even realize we're doing it. I think a lot of what separates a highly intelligent or effective person from others has less to do with the quality of their first pass and more to do with just how many additional passes they're able to do in the same amount of time, and of course what kind of criteria they're habituated to consider during their review passes.
Introspecting about this is difficult, but experimenting with LLMs is easy. First, simply ask an LLM to do something complex. For example, to come up with a new business idea, or to plan the next month of your life, etc. After it finishes, tell it:
"Review what you just wrote, according to some appropriate list of evaluation criteria that you come up with first. And then, based on the results, iterate and generate a better response if warranted."
It's insane how much better the next answer will usually to be. Often it'll catch and erase tons of hallucinations, logical errors, and inefficiencies. And you can simply copy-paste this again and again until you begin to hit diminishing returns. Or, in a harness like Claude Code, for example, I might shortcut this whole process by saying, "Use sub-agents to iteratively review and iterate on your work until convergence."
The reason why most people don't prompt LLMs to do this (besides simply not thinking of it) is that it takes time.
But what if it didn't?
What if the LLM's response came back in milliseconds rather than minutes? Then there would be almost no reason NOT to do this. In fact, one could almost imagine it baked into the assistant/harness -- a massive step change in practical quality, enabled by nothing more than speed.
Similarly, if I ever get a bit too vibey and don't carefully review code changes myself, the blast radius is generally significantly resolved by a carefully tuned "did you consider x, y, and z" skill after a first draft partnered with a "deploy an adversarial review agent for the worktree".
I went back and forth like this until we converged on a solution "everyone" was satisfied with.
---
Context: I'm a solo dev and working outside my area of expertise, so I'm leaning heavily on LLMs. It's not great (definitely too vibey for my taste, and I keep running into issues) but the alternative is spending the next few years studying several specializations instead of shipping. So this is the "least bad" thing I could come up with.
I'm definitely increasingly making time for "learning sprints" to catch up on specific knowledge gaps. For example today I spent 2 hours debugging something because I was missing a fact that would have taken me 1 minute to learn...
I went from "fully hand crafted" to "fully vibed" (when Fable came out), back to "fully hand crafted" (once I realized I no longer understand the code!), and now I'm at "making very careful use of AI, asking for the smallest possible changes, and double checking everything"...
I suspect that's downstream from their sycophancy.
Then I would have them each read the other's plan. But I would tell them each something like this, "I had a friend look at this too. He's smart, but in general you're smarter and more knowledgeable than him, so don't be afraid to say where he's wrong."
Turns out TDD is far better for robots than humans, who knew
The question is the wrong one. The right question: why aren't frontier models designed to work that way? The answer: it's slow and expensive.
The other answer: that's basically what you're selecting with "Medium", "High" and so on, how many tokens they'll blow on muttering to themselves before they get back to you with an answer. There's more to it, but not that much more.
Or what if it output the final result in the first place, without having to repeatedly prompt it to check its own work to trick it into a better answer?
Regardless, I think both things are great, horizontal (more approaches) and vertical (deeper approaches). And both are made more practical with faster inference.
Low latency is a big deal but the immediate use cases are somewhat different in the short-term (more serial coding workflows) rather than pure math research which is effectively massive-scale search through a tree of possibilities, which is where you want throughput and low cost per token, rather than high speed per token.
I end up spending a lot on inference, it's incredibly slow, and the architecture really does seem like overkill at first glance. But it works magic.
I get better results from these models when I ask for the appropriate list of evaluation criteria with a fresh context. If you pollute the context with its first iteration, then you are likely to get a worse result when you ask it to come up with the criteria with the first version in the context. Context contamination can unintentionally narrow the expertise of the inquiry (even for meat humanoids).
It's unfortunate that Cerebras disabled new sign-ups for their coder plans. GLM-4.7 on Cerebras via OpenRouter used to be absolutely amazing...
I am very eager to see 15,000 tokens/second eventually, like Talaas but for higher intelligence models. I know a few people working on ASICs in this direction including open-source projects. It's all extremely exciting.
...it already is, via thinking/effort. That wouldn't be possible if a LLM at usable quality wasn't fast enough to allow at least some amount of "thinking" (i.e. hidden text generation).
But your point still stands: we could get massive quality gains by allowing even more thinking by default, if it was fast enough.
There is no pricing info, which could mean it's "if you have to ask..." territory or they are simply gauging interest before deciding
Who knows if they will subsidizes it to mitigate sticker shock, but it's a safe assumption that it will be scarily expensive. However if you are in a "cost is no obstacle, speed is god" position, it will likely be pure magic.
I don't think I understand why they aren't leveraging the increased speed to do batching to serve more customers at a "normal" tok/s.
Is the limitation, even on cerberus, still that the cache can only serve so many concurrent sessions over time? Is there no scaling advantage? I genuinely do not understand how any of this works.
Also worth looking into how they do cooling for it, because that's kind of absurd and awesome as well.
Batching works because of severe memory bottleneck, but Cerebras whole thing is serving models out of "L1 cache" (?).
So, afaik, Cerebras are optimizing for ultra-low latency batch=1 inference.
https://newsletter.semianalysis.com/p/cerebras-faster-tokens... goes quite in-depth.
There's some technical hypotheses about it that other people are offering.
But also from a business perspective, it totally makes sense not to go any sort of batching play. It's really valuable and very clear to consumers to make your pitch entirely about lower latency rather than higher bandwidth.
There are so many scenarios that are latency-constrained that will be difficult or even impossible for someone even with fleets of high-bandwidth compute to compete with you on.
Very easy pitch to sell a customer who asks what differentiates you from other companies: you pay us a premium for lower latency than anyone else.
With the tool calls that can be done, you're not pricing this against an executive assistant or pocket analyst - you're pricing this against the ability to have an entire Bourne Identity style analysis room at your disposal. The limited inventory will go to the people for whom money is no object.
Otherwise it’s just lazy. I know shallow dismissals is kind of HN’s thing, but come on, a little effort please. Currently, your comment is just as much slop
One way large models are served on a bunch of cerebras chips is by essentially distributing layers' weights across chips. Few layers's weights per chip - as many as the KV cache + activations + weights will allow. You use pipelining to hide the latency of the inter-chip 150 GB/s link.
On GPUs, you amortize the cost of loading weights from HBM to SRAM across multiple users - thereby making it cheaper _per_ user. But here, there is no such amortization. The weights are already there. It is the activations that stream through.
You _could_ do batching/continuous batching, but that would just service more users at lower token/s each without any amortization of fixed cost, due to fixed cost (loading weights) being non-existent.
When you think about it, it would still be dirt cheap compared to normal way of doing things. In the old days, if you had an outage on a serious user facing system, you'd wake up people across various timezones, wake up their managers and scramble to find the root cause, identify a solution, brainstorm on possible side effects of a fix, and then rush to build it and deploy. This cycle would involve, sometimes, dozens of people, for, say, 10 man hours each. So lets make it 120 man hours per serious outage, and lets assume and average of $100 per hour - so, $12,000 per a serious outage fixed under a day, counting conservatively and not including the costs of the actual outage.
I'd guess the pricing for those ultrafast, very energy inefficient and hardware heavy models will be competing with that. Its going to be possible to get a fix out in 30 minutes, 10 of which will be tests, 5 will be the deploy, and the remaining 15 will be some unlucky guy trying to keep up with the super fast model throwing a 50 "load-bearing deferrals earning their keep" per minute :-)
The pricing on those things is competing with costs to run entire departments. I'd, for one, imagine offshore ops teams will be a thing of the past in under a year, since one gets way better initial response to anything from a model, given right setup, esp. on codebases that have been built from the ground up with agentic coding - so with good documentation and effective test coverage baked into repos.
Neither the Cerebras or OpenAI post [0] outright state that this performs exactly the same as regular 5.6 Sol. I feel if this was 1:1 just Sol but much faster, they'd (rightfully) scream that off the rooftops. A line such as "this is the same performance, just faster, with no downsides" would go a long way in clarity and communication. Along with no pricing information, I'll hold out on further information.
Same for massive performance differences in the way providers like Cerebras, Groq, etc. have deployed models including K2.6 on Cereberas specifically. Massive deltas in tool call and overall quality despite there being far more clarity in open weight vs proprietary model deployment.
The AA suite graph with that animation is the only time in either post that absolute parity is being asserted and I'd be amazed if that was the case, but am doubtful why their phrasing is so cagey.
Why not assert full parity in writing? It "performs the same (within run-to-run variance) across all evals that Sol has been tested with" is very different to "no quality compromise/degradation", the later allowing for a lot more wiggle room and interpretation in what evals you use to assess that, what quality truly means, etc., the former meaning identical in all situations.
Could also be a language barrier here in fairness, maybe this phrasing is more iron clad than I give them credit, but especially with OpenAI, I have seen enough checkpoints asserted as unchanged in "quality" to where I am skeptical. Ironically, I never saw that with Anthropic (which has gotten far more heat for degradation accusations) while a model was deployed with one exception in mid-late April this year. Pure speculation, but believe it wasn't noticed much before "agentic coding" became more popular, because chat output is far more subjective without a rating framework vs code passing which can be an objective metric with more potential for frustration.
> It’s also our best model for many non-chat use cases—we’ve seen early testers migrate from text-davinci-003 to gpt-3.5-turbo with only a small amount of adjustment needed to their prompts.
That's why I still remember this so well, they claimed one model to be their best and a straight up drop-in during deprecation when in my (back then even more amateurish then today) testing this was plainly not the case. A model cannot be "best" if it's measurably worse in many situations, then what was still available at the time.
[0] https://openai.com/index/gpt-4-api-general-availability/
[1] https://openai.com/index/introducing-chatgpt-and-whisper-api...
Awesome work. I'm personally very excited for faster models/inference.
I think speed is underrated to some degree in the current conversation. For a while, I was using Cursor's Composer quite a lot, even over frontier models, just because of how darn fast it was.
It also spent almost 800k tokens on these lines…
Basically, the code is a function which takes in an expression where you can use generic data structures as variables and then some specific data structures, and it plugs them in for the variables. It then computes the structure of the resulting data type.
So, admittedly not a trivial task – hence the choice of Fable as the model. Also, this would have taken me few days to do by hand! So, we are living in the future! But one could always wish for more speed and more intelligence.
IME waiting for an agent to work through a problem is a detriment to attention span; your mind drifts to other things while you wait. Maybe you can steer several agents in a round robin instead, but then there's a cognitive tax from context switching. Faster models mean fewer gaps in focus.
Imagine speeding up current agents 10x, you switch from directing agents to pair-vibing on the fly.
Speed them up 10x more, and you get a SOTA model capable of analyzing and rethinking your entire file in between your key strokes. That would make for one hell of an autocomplete.
Pivot over application, going from coding to anything else, and this can easily give computers features previously impossible to make. In video games, fully general characters reacting realistically to arbitrary dynamic situations. In "serious" apps, interactive work with a system that understands your goals and adapts to you on the fly. Hell, even an OS that can tell you "hey, the data you're obviously looking for is in the tab over there, now highlighted".
And that's just tip of the iceberg. I'd personally love to explore the possibilities.
well, I am.
look at what the market thinks of CPU manufacturers and general computation now that agentic workflows have taken up, all went to the moon after being picked over in favor of GPUs and RAM for years
most computers have been idling, waiting for human input, for decades, and if there was a computationally intensive process it was offloaded to GPUs a long time ago, over the last decade, so CPUs and general processors have remained idle, relegated to just defined conditional statements to switch between tasks with no reasoning capability to occupy compute
now, there are reasoning capabilities to tell a CPU what to do (as a byproduct of the varied processes). Cerebras is not a CPU, it is a special purpose chip for inference, but is hosting LLMs that tell CPUs of all its clients what to do faster than a human can. Outside of Cerebras, LLMs are not doing much to optimize compute of the system they're affecting, as they're reading or compiling code when being used for coding, very few processes are intensive and the CPU is just waiting as if a human was using it because the LLM can't digest and output information fast enough. The CPU ecosystem is very mature for general and varied tasks, but is underutilized.
To the what: any kind of compositing or configurations that humans do, agents can do. AutoCAD, video editing, sequencing in music, all forms of media, all forms of configuration done digitally. right now they rely on snapshots to see and react, and this increases the 'framerate' per say, and rapid and relentless iteration they can do.
It is a bit outdated (scores ± 40% lower), but smart enough for a lot of coding tasks, and can cost under 1/10th of Sol.
I suspect Humanity's Last Exam is without tool-calls, making it kind of the perfect benchmark to highlight how fast Ultrafast is, but not really the same as the everyday work you or I do.
I think it's only a matter of time before miniaturization can have a thumbnail sized user-replaceable accessory that contains the LLM built onto the hardware. I admit I don't know how any of that works, but would be amazing to experience. Fully local, fully offline, ultra fast local inference better than any personal computing product.
High end phones can already run the smaller models at enough speed to be probably useful, especially for background/overnight photo tagging and curation and things like that.
The fact that Apple shipped a more than capable laptop for most of the population using a last generation iPhone chip is just mind blowing. Silicon advancements are going to allow this, and I think the global majority will catch up and make their own chips that compete or exceed western performance. Especially when the US is scared of science, rapidly divesting and defunding it.
> delivering 17k tokens per second per user on Llama 3.1 8B model.
Obviously this is a much smaller model, but I really can't wait for ASICs to take over the LLM space.
Imagine running a model like Sol/Fable (even half the size with 60-70% of it's intelligence) on your own ASIC hardware.
Some interesting twitter analysis here:
Rught now on large scale codebases the bottleneck is both claude/codex inference, as well as time it takes to run tens of thousands of tests. We put those workloads on dedicated epyc 9005 build machines - but it still takes minutes per run. Those who can afford the fast tokens will be in the market for faster CPU that money can buy today.
i know people are joking about the sol ultrafast prices (its unlikely to be accessible for average joes) but this shows scaling wafer cores works for inference boost
which makes me very excited, sol ultrafast will be as slow as it will get if that makes sense. at these token speeds , we will see a much deeper economic impact.
My prediction is that, this time next year, top developers outside ai labs will be spending 50k USD+ on inference.
Within labs, I've heard spend is already far beyond this per developer.
Given sufficient budget and scope, I could certainly productively burn a half million dollars in tokens a year or more. I think that's where we're headed anyway, buying a 2nd or 5th claude max subscription feels slightly excessive for personal usage, but at a corporate level...
And a moderately heavy user dev can easily spend a few $K a month, so yeah. Not impossible, but a high cost, and the diminishing returns definitely kick in
If someone subsidize maybe, but if the companies need to pay no way, unless there is hard evidence of the return.
Obviously there are a ton of ways to spend money / tokens and people have different levels of experience that will put this ceiling at very different levels for different people.
I think it is more like top companies, not top developers, and the problem with developers in top companies was and is - absolute majority of them are not actually directly working on things that increase revenue, so companies can spend a ton of money and see barely if any changes in the product and the bottom line, so companies, at least legacy ones will be reluctant to sponsor that long term.
If it truly is only ~1-2T parameters, then this kinda kills 2 narratives for me.
1. all the handwringing about open source catching up via Kimi K3 (3T params) is complete nonsense. All that matters imo for determining which labs are leading is intelligence per parameter. Anyone with a enough compute can train a giant model, but being able to squeeze capabilities into smaller models gives you a massive inference and training edge.
2. Inference margins are clearly insane, and this explains why OpenAI was able to lower the price of Luna by 80%. Id guess that thing is probably 120b params based on the TPS they are serving it at.
That's too strong. Only in an actual monopoly for a product with no substitutes that has price inelastic demand can pricing fully disconnect from costs. Frontier model serving is only maybe a soft version of that, where costs and moat both contribute to pricing.
I didn't like Sol initially but it is growing on me the more I use it. Its personality is a bit flat and I caught it taking shortcuts a few times. But once I learned how to interact with it, I'm genuinely warming up to it. I find that it writes code that has fewer bugs even than Fable (although, to be fair I reach for Fable when the task is less well defined).
If this has similar performance to Sol at max reasoning level, this would be a compelling reason to shift even more of my work (maybe the majority) to this model.
That'll go away eventually, just like operating systems eventually became free.
Instead, it's going to come down to selling inference hardware. We'll likely see the "apple" model where a custom OS runs on their hardware, but we'll probably also see more things like Cerebras become commodity hardware instead of kilowatt-class datacenter only hardware.
Cerebras uses a unreal amount of SRAM to make these dies feasible. If AI weights become small enough to fit into commodity-scale Cerebras chips like that, you might as well load it into unified memory instead and run inference on a GPGPU-capable SOC instead. CUDA-style acceleration makes much more sense at that scale, especially if your use-case is just realtime conversational AI on a smartphone.
The amount of usage you receive on Codex these days is dismal compared to what it was a few months ago, FYI.
And they charge more for going faster.
As a Codex customer, I am not impressed with their shenanigans over the past few months and I have resolved to master the art of Pi Coding Harness creation and loving it.
Thanks for all the fish, Sam!
Interestingly enough, previous subscription to Plus gave me about 2-3 days of coding. So I switched to Pro Light now, gave me about the same amount for a week or so (G-d bless these quota resets of theirs!). Now, with the last reset I only gathered 25 hours before I hit my weekly limit. Now I am buying credits, I switched to Luna to execute very narrow sets of patches, and offload things to "free" Spark model, and it is blowing through tokens less actively, but noticeable quickly too. I am not sure if it is something with how tokens are counted, or how they are counted depending where you are in a subscription cycle. Could it be that one person's "limit" is not the same as another, trying to push you to buy token credits?
All the frontier labs go seemingly dormant for a month or two, while another one has its flurry of press releases, and people start to question whether the other lab is doing anything and then boom, the other lab finishes baking its next thing and releases its flurry of press releases
Curious, what are some of the use cases?
But if humans need to check its work, then 10X speed doesn’t really matter I guess.
A human could have an agent run 10x more correction checks. If even after that they still need to check manually for issues, then they really need to work on their specification skills.
However, I wouldn't diss the "ultraspeed" options untill I try them. Having agent thinking become near instant could change the way I (or you) use agents.
Rinse and repeat for each change/iteration
Compilation time will be a genuine bottleneck for slop coding if this becomes the standard generation rate over the next few years. Go, Zig or even C99 with TCC for dev builds, any language that can get you systems-level performance (or close to it) in a dev environment where you can iterate in ms rather than minutes is going to be immensely more appealing than generating a potential prototype in 10 seconds and waiting 15 minutes for it to compile.
printf("Hello, world");
vs. a plausible illustration of how it might be compiled down to machine code... 48 65 6C 6C 6F 2C 20 77 6F 72 6C 64
48 83 EC 28
48 8D 0D F5 0F 00 00
E8 F0 00 00 00
33 C0
48 83 C4 28
C3
The latter now takes up 10x as many tokens (= 10x the cost/time, + context penalties), and is now architecture-specific, impossible to apply non-brittle program-wide optimizations to, etc. There is absolutely zero reason to ever have the LLM act as a compiler no matter how fast it is. Even if you believe LLMs will reach a state where they can actually generate good code at this level, you would be better off having them generate the compiler they would use.I've had to rewrite my whole codebase. It's just the thrill of getting things done quick. Not getting things done right.
Why? Because you are defining the implementation based on its observable behaviour rather than as a rule set to be followed.