I guess this smaller context but faster llm engine could be good to develop your harness on to get faster results and faster iteration.
I guess this smaller context but faster llm engine could be good to develop your harness on to get faster results and faster iteration.
128k is pretty much only good for chat/conversational/question asking (including tool calls for searching things and spitting back/parsing a set of results) or human interactive agent purposes.
I wonder what the approximate context window of a human programmer is... less than 12k lines I'm sure.
But that's before the prompt.
Then after the prompt, every tool call the agent makes adds to the context. Longer turns can easily consume 50k-100k tokens between the agent and various tool calls (reading the filesystem, reading files, reading compiler output, reading memories).
Then each "turn" with the agent stays in context and is fed into the next turn. Two or three turns and you're up near 250k.
You have to have a pretty inefficient use case for a subagent to think that 128k is not enough.