HN: https://news.ycombinator.com/item?id=47934437
Reddit: https://www.reddit.com/r/worldnews/comments/1sxzzop/comment/...
Full: https://reddit.com/r/kimi/comments/1v9aqsi/comment/p0cbe0n/
Mirror: https://redlib.us.catsarch.com/r/kimi/comments/1v9aqsi/k3256...
First you spend like 2-3 hours working on a plan, once you have that you just tell the model to go and implement it, do adversarial sub-agent review loops before each commit and also make sure that all tooling and tests pass (including coverage requirements). You do need to poke it in a slightly different direction every few hours, though. Not even any novel work, just some refactoring and SSE notification hardening, bug fixes, alongside environment tuning and getting rid of some bottlenecks (also migrated from Oracle to PostgreSQL but that's mostly done).
That said, Kimi somehow manages to use less context in the main thread than Anthropic's models (even when you use sub-agents and also dynamic workflows in Claude Code), might have something to do with either how the model is tuned or their Kimi Code harness - because even in most of the longer form sessions it doesn't seem to fill up quite as quickly (note: because the kimi vis tool doesn't have a full summary view across all agents, these are the main long running agent stats across some sessions, not sub-agents):
total tokens cache hit rate wall time peak context
283M 98% 3963m 466k
258M 97% 2724m 467k
98M 94% 1353m 393k
67M 97% 614m 434k
75M 98% 1447m 498k
53M 99% 191m 375k
6M 96% 139m 124k
7M 98% 86m 118k
11M 99% 61m 147k
I could see 256k context being sufficient for all sorts of work, even if intermediate progress/plan tracking files and docs might have to be used along the way, in addition to whatever plan support the harness has (for example, if you document something that will be relevant for load testing you might need that in 10 turns but not during the ones before then).If I feel like the model and I explored a lot of options I won't want to keep context as it might be confusing.
I think the more you use it the better judge you are of whether you should purge, compress, or just keep the context before executing the plan.
I use handoff skill to ask model to write a prompt for itself.
I've found Opus 5 far better as a subagent with very limited context window use, which could suggest that it might not have good long context performance unlike its predecessors. (I was one of many tearing my hair out trying to work with Opus 5 for the past week.)
Yes, pretty much - if there’s a lot of noise and jumping around and wrong conclusions and corrections, compress and only leave the correct stuff (maybe make some plan file briefly mention what NOT to look at/do). But if it’s all fairly straightforward then can just proceed with the execution.
Most of the time the planning stage ends up short of 200k tokens anyways, it mostly takes hours just cause I’m slow and need to explore the various options - still cheaper than building the wholly wrong thing and having to redo everything.
Compressing the context can also drop important information so it might be better to only do that when you need to / use the plan mechanism/files / do it after completing some large stage of the plan and so on - so a judgement call.
A greenfield project would probably allow at least 2x fewer tokens to be used in most of those long tasks, but I was mostly after consistency and bug fixes along the way as needed.
The model is then given a tool that allows the model to decide to roll back to a checkpoint + a message containing any additional useful information.
It's specifically instructed to use that tool[1] in cases like when it has inadvertendly read a large file where most of the content is not relevant to the task, or after a web search where it's found what it's looking for but most of the content isn't needed, or when it's written code that didn't work as expected, or similar.
It basically lets the model backtrack and "forget" irrelevant details at the end of the context but give itself hints on how it should continue from the checkpoint.
Though, interestingly they seem to be abandoning it in their new CLI (kimi-code), unless it's been folded into other functionality. Not sure if they just feel it's not needed any more with their newer models or if it just didn't work as well as they expected.
[1] named "D-Mail", or "DeLorean Mail" in a reference to Steins;Gate, which again references Back To The Future. See https://github.com/MoonshotAI/kimi-cli/blob/main/src/kimi_cl... and https://steins-gate.fandom.com/wiki/D-Mail
One way or the other, a prefix is saved. Only the additional info how to continue is new.
Haven't gone back to it, have been using Claude Code on a private project I'm still architecting.
But yeah, I have no idea about anything about software because you made an assumption off very little to go by.