> We evaluate 12 popular LLMs that claim to support contexts of at least 128K tokens. While they perform well in short contexts (<1K), performance degrades significantly as context length increases. At 32K, for instance, 10 models drop below 50% of their strong short-length baselines. Even GPT-4o, one of the top-performing exceptions, experiences a reduction from an almost-perfect baseline of 99.3% to 69.7%. [0]
I copied all the code that I could from TFA and pasted it into OpenAI's tokenizer. It counted ~15k tokens. Many other tokens were generated in TFA's chat, some of which are not visible to the user. I think it's fair to assume that the entire chat was at least 25k tokens, right? Therefore, I believe that by the end of that chat, 4o's performance was significantly degraded.
I think a major skill to develop for LLM supported coding is to compress a chat after just few thousand tokens into something like a step1.md file. Then start a new chat with "read step1.md" as the first prompt, and so on..
Is my logic sound here?