According to this analysis, at 70 tokens per second for Kimi K3, you could expect to run >800 parallel streams on that setup:
1,058 karma · joined May 25, 2026
According to this analysis, at 70 tokens per second for Kimi K3, you could expect to run >800 parallel streams on that setup:
The chart shows GPT-5.6 Sol and a surprisingly large drop in performance when the switched it over to GPT-6 Sol.
Cost per GPU hour versus API price of generated tokens assuming 100% utilization.
This could be a margin around 98.3% for 5.6 Sol.
If the utilization of the GPU was 25%, it would drop to 93.1%.
Revenue sharing or training expenses are not considered here in this "inference margin".
I wonder if margins on GPT-6.1 Sol and Opus 5.5 are now 75% or 90%.
> So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.
That inference wasn't profitable is a widespread myth.
Analysis based on Kimi K3 suggests that OpenAI and Anthropic have margins well north of 95%: https://inferencex.semianalysis.com/run/kimi-k3-on-b200
Over the last months I have seen news that OpenAI made breakthroughs in inference efficiency multiple times.
I have no reason to believe that the leading US labs don't have their own optimizations, or that they learned of this particular optimization from DeepSeek.
The results are less buggy, animations are much better.
It can work autonomously for hours and the result is decent most of the time.
That wasn't usually the case with 5.5, which needed more feedback and iterations to get things right.
Not sure what point you're trying to make here.
So, I have the following thoughts:
How much more productive can AI for knowledge work realistically make the global economy? 5%, 10%, 20%?
What if it also does robotics soon and can be deployed in manufacturing and construction? Can we get 50% more productive, or 200%?
Then there's science, medicine, and so on. Who is to say what happens if AI solves also that?
While I do not know if AI will be capable for these use cases, I see no reason to believe it unlikely that AI could be applicable to eg. robotics and speed up production by factor 3x.
I'm pretty confident I can release it in October, after some 500 hours of work.
Usually in these discussions I get "so it has zero users then".
One time someone on here told me that for AI to prove its worth for software engineering, my application would need to be in use already for at least ten years.
However, I don't know how future larger models such as the cancelled 6.1 Astra will be priced.
If the price stays high, this would indeed be quite bad for the $200 subscription..
So yes, it is clearly cheap in comparison.
Tibo said that the existing $200 subscriptions keep the 20x factor for a while.
Ultrafast would have been nice with the temporary "Pro 400" plan.
The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.
I am very skeptical of claims that old models weren't much worse. Compare this to February's GPT-5.3.
With the 80% price cut, this is competitive with Opus 5.5 despite the subscription downgrade.
Additionally, it was said that existing 20x subscriptions retain the higher limits for some time.
I have seen you make these immature accusations that users here are OpenAI employees multiple times today.
It seems weird to me that just using a ton of output tokens manages to produce a decent result in the end.
It seems to work well though. Sometimes I fear that it might be more likely to eg. run an incorrect, destructive command, but maybe that concern is not justified.
It doesn't suffer the same verbosity issue and the writing is somewhat clearer.
The vision capabilities of Astra are a bit better I believe.
That said, Opus is 3x cheaper and the obvious winner when usage is considered.
There were rumors of a model called "Astra Minor".
I expect they will release this today.
If it performs slightly better than Astra and, considering the subscription cut, is 1/4th the price, it could be competitive with Opus 5.5.
Opus 5.5 appears to be around 1/3rd the price of Astra with a subscription.
This is very unlikely to change without regulatory intervention.
Competition is what stops the exploitation.
Epic Games has always offered lower fees and their CEO has been advocating for small developers for years.
There is no reason why Google and Apple get to bundle their stores, and these other big names have to collect the scraps from users who care to look up alternative stores and go through scare screens.
Of course I understand that they are not going to spend money on improving their review process for our benefit until regulators force them to.
In fact, Google is making alternative stores even less convenient with their latest 'side-loading' hurdles.
Regulators need to wake up, issue record fines, mandate equal access and a 'Choose your app stores' screen on device setup.
Otherwise competitive third party stores and a fair market are going to remain a pipe dream.
Here I expect factual discussion rather than misleading statements that suggest there was no way for the Muse assistant to follow an instruction never to do something again.
Vague arguments that such instructions are not deterministic are uninteresting, because it is obvious.
A bit hypocritical there, no?
Considering what the author says about people who believe AI produces code that is good enough.
And in my view it clearly does, particularly when you care to iterate in order to iron out issues you find in manual testing.
That seems ironic, considering they ship like 20k of context in Claude Code's system prompts etc.
I think it's unreasonable to pretend that this couldn't possibly work, that you understand to what degree it does work, and that the only thing we should be discussing here was that it isn't deterministic, which every reader already knows.
However, it is untrue that it doesn't have the capability to memorize an instruction and diffuse it to new sessions.
Simply saying "do not ever do this again" can result in the behavior not reoccurring with any likelihood.
You'd have to benchmark whether with the instruction in place it would violate it, and in how many cases, so that you can understand the risk better.