For 2 or 3 newspapers it works; my idea was to use it as grounding to discover relationships between people, companies and jobs.
As for the "everyone's life", I have always assumed that there would be a graph system to point to "forgotten" documents.
Gemini said my idea was amazing and new in its implementation, even if not in spirit, but I'm assuming it was being sycophantic as usual.
My sense is that this is sort of accurate, but more likely it's a result of two things:
1. LLMs are still next-token predictors, and they are trained on texts of humans, which mostly collaborate. Staying on topic is more likely than diverging into a new idea.
2. LLMs are trained via RLHF which involves human feedback. Humans probably do prefer agreeable LLMs, which causes reinforcement at this stage.
So yes, kinda. But I'm not sure it's as clear-cut as "the researchers found humans prefer agreeableness and programmed it in."
* With Claude's 1-million context window I have been doing some slightly longer range tasks — ~1-3 days of work — with RPI/QRSPI frameworks(see last few days of comments else where on HN) in one context window. They involve a grill-me session with 20-60 sometimes more questions for tasks to get alignment which produces the design and the plan in one window.
My experience with this has been that it front-loads a lot of the LLM interactions, which can be exhausting without a reward (i.e. output.) And then, when I get the output, it's so large as to be hard to review/grok.
In other words, it feels a bit like when my coworker delivers me a month's worth of work in a single PR.
We don't, no. But wouldn't it be great if we did? I'd sure love to be able to hold the entirety of the code of my organisations monolith in my head at once. It would make everything so much easier. It would definitely also cut down on the bugs I write!
Similar if I could recall all of my organisations confluence pages. Id probably be a lot better at my job. Same with all the slack history. All the hr documents, press releases, meeting transcripts. Theres practically no end to useful context even just in text form, and even if much of it is not relevant to any one task, having all of it in working memory would be fantastic, if only it were possible. I could probably make incredible cross organisational efficiencies and probably be far wealthier if I were some savant that could hold all of this in my head at once.
I get that we have agent harnesses to try and fetch only the relevant information. But most of the problems result in either failures in this process, or previous things falling out of context. I very rarely see failures where the agent forgets stuff already in context. The harnesses are making up for this exact limitation!
That sounds like the beginning of a sci-fi story where the conclusion is forgetting is not such a bad thing.
It seems far more likely that it would all get baked-in to the LLM during training, but maybe it will turn out to be really useful to train up a "generic robot controller LLM" and pass in a huge number of tokens to better optimize it.
I do not think it is the direction for everything.
Generally, we need consolidation of experiences and memories to just remember the important conclusions, ideas, and concepts, and then the ability to remember the full details if they are relevant (which they usually are not.)
But for some applications I am sure a billion token context would be useful.
It is likely most people need a 10 core CPU or whatever for most tasks, but for some applications you want a supercomputer with 1M cores.
So we need a taxonomy, we need memory layers, we need summary/details. If there is one thing I have learned about how these LLMs work, if you give them a few flexible tools they can work the shit out of them to achieve objectives. We just need to right tools and right structure for context.
We simply don’t know how to incorporate new information without losing old capabilities reliably. Pans handle this through extensive evaluation, heuristics, and experience.
What we do know is that models can adapt to their context, and extending the context window is an infrastructure and capex problem first. A billion useful tokens would obviate the need for any out of band memory structures.