Increase GPT4's context size by asking it to compress your prompts
twitter.com
twitter.com
A) people have had a hard time reproducing it, and
B) more damning, the "compressed" version uses more tokens than the original (https://gist.github.com/VictorTaelin/d293328f75291b23e203e9d...)
> In practice, plain, well-designed summaries should be optimal to fit larger documents in the context.
> This concept has potential, though; building lookup tables seems to outperform long text summarization.
It's a clever idea, and I agree that lookup tables and external storage of memory is likely going to be important at some point, but I suspect that's going to come out of giving LLM more ability to externally reference "long-term" memory rather than compressing everything into immediate context.
Sessions are getting treated as more valuable than necessary.
This should look more like functions.
Needs lower prices and greater availability for that.
But until then… smashing duplos together.
How likely is it that context size will greatly increase in the coming year or two? Are there fundamental limits, or could we reasonably expect greatly increased context size in the future?
Or is context likely to stay fixed around 32k for the foreseeable future?
That is the question.
GPT4 seems to prefer long-winded replies, i.e., when I ask for what amounts to a one-line fix, it repeats the entire block of code with the one line correction. In contrast, Replit's Ghostwriter often gives a concise reply showing only the one-line fix, and I have to ask for the entire block when it isn't clear where the fix is applied.
Obvious ideas being held up as something more than that.
If you think this is incredible, give yourself some time to go play with LLMs. Try things. Clever things. Stupid things.
It helps set a more reasonable scale for these sensational sounding snippets.