Guess they don't care about regular devs atm and are focused only on hardware sales.
Guess they don't care about regular devs atm and are focused only on hardware sales.
OpenAI's Sol ultrafast (powered by Cerebras) is still in preview, presumably because they're overall capacity bound.
Because they already have / had an okay coding subscription product for a bit and it gives them visibility and mindshare (in regards to their hardware, even if they don't compete with other providers that much). They could do what Kimi did - make a good subscription with good models, once you get enough customers to get some good PR and such, pause the signups so you don't have to spend more on running the service than you want/can. Do enough of that and people will talk about your offerings organically, make yourselves known to even devs as "That one company with their own hardware and the super fast subscription." experiencing which would do more than any marketing.
Cerebras is a B2B hardware company. It feels like a distraction: think of the opportunity cost, and resources/headcount not working on other things that would drive more impact.
Should NVIDIA do a coding subscription too? I'm sure they can make money off it, but I think it would be -EV.
Yes, obviously! Well maybe not a subscription but definitely an inference service.
https://resources.nvidia.com/en-us-inference-infrastructure/...
https://www.nvidia.com/en-us/data-center/dgx-cloud-lepton/
In their case not to gain mindshare or money or whatever, they're already a market leader, but to run something that validates the use case of their own hardware (across a bunch of 3rd party models) on a practical level and gain whatever insights or details might be relevant to pass on to other hardware and software teams.
They sort of do? They offer free access to various versions of nemotron via multiple routing services.
But it's not fully open to just anyone, I wasted time signing up to find out that I couldn't even sign up for it to test it out.
They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks.
Edit: I don’t know if they actually have a proper cache. This could just be a billing artifact.
They do not seem to discount cached input for the self-serve Developer tier. Maybe they do for enterprise rate cards?
https://inference-docs.cerebras.ai/capabilities/prompt-cachi...
Why would they want to target regular devs right now? If they sold to regular devs instead of enterprises, the complaint wouldn't be about model choice, it'd be about how expensive it.
I still think that was a really great model that got overlooked. It was really great in terms of latency/throughput while still being fairly intelligent.
I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.
Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs and you have to fix a non-insignificant amount of it all manually: https://blog.kronis.dev/blog/i-blew-through-24-million-token...
Admittedly that post was before agentic development truly took off and that 3k EUR figure when paying per API tokens would nowadays be closer to like 6k EUR for the volume of work I do, but still.
It's the same how Qwen 2.5 was pretty problematic for anything remotely serious, same with Qwen 3 Coder Next (80B), and at least the most recent versions are getting better but still not quite good enough in real world use cases outside of benchmarks. They've come a long way, regardless!
Oh yeah, I'm still amazed how good the current iteration of models are for coding (I have a fear it's too good to be true - so will get taken away..). Exactly a year ago I switched from GPT 5 to Gemini just because the coding with R language was terrible; and even with Python it kept forgetting and mixing basic stuff. Gemini at the time had much longer context window and was miles ahead on R syntax.
Current experience of just leaving a Codex Agent chug until a stable solution is completed is still mind blowing to me.
> Guess they don't care about regular devs atm and are focused only on hardware sales
They aren't trying to make a few bucks off tokenmaxxers. They're trying to be the underpinning of compute for all AI. They're going to beat Nvidia.
Because they generated some buzz and are near-SOTA and would be a great benchmark for a PoC subscription that doesn't necessarily aim to compete with other vendors at a similar scale (since their main business is the hardware). Mistral is conceptually cool but is lagging behind. I guess Muse Spark and Laguna would also be okay, just not as recognizable. Meanwhile both Kimi K3 and GLM 5.3 are near-SOTA in performance and considerable in size, a great choice for proving the platform!
As for the 2nd part of your question - that wasn't a relevant concern or consideration here, unless the models would be tainted to a degree to prevent them from having a good coding subscription that gets more developer mindshare towards what their chips can achieve and generate some good PR.
You pick the vendor who is ahead, create an agreement with them to get access ahead of public release , and bake that model into hardware , because that’s how you make money .