Big Post About Big Context
gonzoml.substack.com
gonzoml.substack.com
With GPT-4 Turbo, it is definitely NOT a good idea to just throw lots of extraneous code into the 128k content window. That distracts and confuses GPT, makes it code worse and less able to follow complex system prompts. You get much better results if you curate a smaller portion of the code that is relevant to the task at hand.
I am really interested to find out if Gemini has these same tendencies. If so, quality RAG is going to remain valuable for curating the context. If not, then Gemini would have a huge advantage over GPT-4. It would be really valuable to be able to naively harness such a large context window.
It is unable to find edits: https://twitter.com/mengdi_en/status/1763338245696274469
It hallucinates changes badly: https://twitter.com/thomasahle/status/1763408041960231010
To be fair, finding diffs is an odd way to use LLMs, but does show the highly variable abilities of these models over bigger contexts.
In other models, reasoning performance declines with more tokens (no data yet here for Gemini 1.5 Pro): https://arxiv.org/html/2402.14848v1
Note in that last chart that significant declines in reasoning are apparent over just a few thousand tokens.
Original: "The rest, that are within the note of expectation, alreadie are i'th' Court"
New: "It looks like everyone is here already"
Then I gave the entire text to Gemini, ~400000 tokens, and asked it to analyze the play's diction and style, and to identify any lines that didn't match the prevailing dialect.
It found that one line I'd changed.
My test demonstrates that even with a large context, Gemini 1.5 finds the needle in the haystack.
And it's not even searching for specific terms or meanings, the only clues are subtle differences in style and diction.
It also seems to line up with random paragraph location doing the worst in that paper.
I don't miss sweating the details on what exactly to fit into 4096, but if I just use GPT-4, I'm left with the distinct impression the current context extension stuff is hacky, and something needs to be done in training to actually make use of it.
Google miraculously getting 1M context with perfect recall is more fuel for my "don't tell me, show me" fire when it comes to Google. They're batting roughly 0 for 2 years now
In terms of the actual output it can be annoying if the convo goes on too long and it hits the window, but for technical things this hardly matters as you can just restart it - more of an RP issue as i understand.
The status quo is that you have employees ingest and internalize this material over years of employment, then have those experienced employees manage large projects by telling the project if they're in compliance with the organization's standards.
A reliable recall system like this isn't just a drop in replacement for new employees, its a drop in replacement for a lot of what makes the expensive employees important: the backwards-and-forwards knowledge of the standards/practices.
> 6) A very important and at the same time difficult class of solutions is model result validation
Thanks.