Kimi K3-256k
kimi.com
kimi.com
Having a lot of active context increases the per-token cost (flops issued and bytes read per token out) so it makes sense to pass that cost on to users. I'm actually surprised it's implemented as a hard cutoff instead of a smooth gradient.
Fewer nodes dedicated to prefill per instance, and fewer nodes in total since you don't need to support a higher KV cache.
Disaggregated inference also means they can tune the balance of compute dedicated to prefill seperately from decode
One would think compute-constrained actors like Anthropic would have done so, unless prefill isn’t really a bottleneck compared to decode?
If you create 3 buckets of inference pods, say, 256k, 512k, and 1M, then you have to worry about filling/dynamically-scaling all of them.
And my guess is there's probably not a huge amount of customers that want somewhere in between: if you're willing to pay the long context surcharge; you're probably semi-price-insensitive anyway to just use 1M.
That seems like it is not an awfully hard problem, because the AI buildout scale is so massive. You simply don't provision too many of the smaller buckets, and when they are fully loaded you route the traffic to the larger buckets. Now you don't have a utilisation problem.
Now of course you can simply organize it that way internally, and present a single price to the customer. Perhaps thats the smarter thing to do.
The cost may be smoothly variable, but likely there's a bimodal distribution of users who barely use any context and users who push it to the max. Average price across both extremes fits nobody, but averages per kind of workload can be close enough.
Having it as a separate model makes it easier to load-balance the traffic.
--
Jesting aside, I cannot think of a single session in recent memory that used less than ~70k tokens, so I presume you are joking.
Apparently it was taken out of context or misattributed or something. Similar to "Al Gore said he invented the internet"
I've played with explorer agents giving exploration summaries to help the implementer agents use more of their context for implementation, but it doesn't work as well. There's always something lost in the handoff.
Edit: I was wrong, thanks to longwave for pointing this out. It's absolutely possible to start out on the 256k model and then switch to the 1 million model when you get close to the context limit without invalidating the cache:
"When switching from k3-256k to k3 (1M), if k3-256k is close to the 256k limit and you don't want compact to lose information, you can switch directly to 1M. The current version switching from 256k to 1M does not affect the cache."
So I'd guess it's API level.
Kimi K3 supports a context window of up to 1 million tokens. We achieve this through extending
the context window progressively as training proceeds, following a four-stage curriculum. The
window grows from 8K to 64K tokens during pre-training, and from 256K to 1M tokens during the
cooldown phase.
https://arxiv.org/pdf/2607.24653 page 12With a lot of architectures, you technically don't need to retrain to extend the context window either; e.g. RoPE scaling; but performance is typically crap.
I usually keep my context in chats below 256k anyways so this would be tremendous honestly.
Who cares? China has always copied and then copied the means of production and then out produced. See also Tesla and now all the Chinese cars eating their lunch.
As long as I get really solid AI models for cheap that do what I need I don’t care if they’re Chinese or otherwise.
I’ll still never use Grok from SpaceX AI cuz eww no, I have principles. ;-)
Cheaper?
This seems like a hot take not really informed by experience.
Please read my essay on context window saturation. I am also the author of deepcq (intent-aware memory querying & context mgmt for LLM agents)
https://platform.claude.com/docs/en/about-claude/models/what...
This.
When working on hard problems (not "vibecode me a script to show an alert box", but e.g. "let's see what this three-level LUT-state-machine obfuscated binary does"), hitting the 1M (!) context window with Claude Code feels like you were talking to Claude Claudewski when his shift just abruptly ends, he packs his things, throws the office keys at Claude Claudeson in-between the front door frame while handovering like "Hi! Nice to see you, good luck." and now here we go again, you are working with someone who just experienced an acute amnesia. It tries everything it already tried, everything it was told in the initial prompt to not do, everything it was told in follow-up prompts not to do. "You were right, this approach does not work and we don't have 20 TB RAM on this machine for full symbolic execution, let me try..."
In Codex, it's so seamless that I sometimes just notice "wait, the context was 20 % remaining, it is 70 % now, wow, when did this happen", while it seamlessly works on the task. Basically never had an issue with context on Codex, be it coding features, cracking hard crack-me ciphers, or researching basically anything.
(For full disclosure, my last experience with Claude was a few weeks ago when I cancelled the subscription, maybe they fully reworked the traumatic "Summarizing" - "Oh, hi! Where are we? Who am I? What we are doing? This is taking too long, let me take a shortcut..." lobotomy they were doing in the meantime.)
256k is enough when the harness uses it properly and the model is not stupid. And also when the tokenizer is not tuned to invoice as many tokens as possible...
Claude might as well not even do it in my experience.
See Kimi Vendor Verifier.
A nice thing about open router is you can specify filters, like "US hosting without data retention"
They did exceed the expecations of pretty much everyone! I've also blogged about using the model, it's a bit on the slow side but pretty good!
> so were totally surprised and overwhelmed by the requests ... or do they just don't have the hardware being in china?
Yes, this is mostly the case: https://x.com/Kimi_Moonshot/status/2078855608565207130
As a user, I much prefer that to service disruptions or severely degraded or secretly quantized performance. However if I didn't have an account, I'd be pretty pissed off about not being able to give them money and become a user.
I am now rather pretty pissed towards antrophic for stopping my flow and forcing me to search for alternatives.
It's not ideal that people can't join but as a user I am happy that they focused on serving existing customers.
Kimi had become that popular. I was a subscriber of Kimi back when latest version was Kimi K2. Later I unsubscribed because I jumped over to GLM subscription (they had amazing deal). Now when I wanted to try out Kimi K3 to find out what the fuzz was all about, I couldn’t subscribe to them.
I remember reading a post from Moonshot team about this, they are doing this because they are almost at peak capacity and want to reserve it to keep the quality for their current customers.
We are actually witnessing an open-weight model catching up at catching mainstream users attention. And instead of behaving like Anthropic, they actually care about their users experience.
When they first announced, I created an account but didn't subscribe, so I am stuck on a waitlist now.
I was able to press "Join waitlist" and then within 48 hours got accepted. The limits aren't very high, no where near the endless subsided+resets given on ChatGPT/Claude. I recommend ChatGPT for good value output!
Others here mentioned "the providers could be quantizing it!" but some of the providers on OpenRouter have partnered with Moonshoot and OpenRouter shows the int when you expand on the provider.
It should say "mxfp4" but providers like Baseten report FP8.
I’m primarily using gpt-5.6 and then opus for reviews.
I forgot why.
But I can't use an AI cli without that feature.
Is this the exact same model just with less VRAM allocated for context window?
Doubt these are related, but it made me laugh a little.
Of course - Anthropic and OpenAI have an advantage in the amount they can subsidize the usage, but I think those days are waning.
Codex and Claude require editing a .json file, but most other harnesses have direct connections via a /provider or /login command.
For pressure at the country level leading to this kind of thing I think it's very unlikely. Here in Sweden it wouldn't just require a vote in the Swedish parliament and before this there'd have to be förarbeten and you can't just brazenly push things through with insane arguments, Swedish social convention goes against it-- and there's just no way to get it through.
It also might not even be legal. "We aren't at war with China and I'm a communist, and the US LLMs are so aligned with values inimical to my political ideology that this is interference with opinion formation" might be an actual legal argument that the ECHR or CJEU might actually have to accept.
I guess it sucks if one wants to expert broadly, but if you're big on vertical integration and the Japanese won't sell I guess you take what you can get.