Clearly hardware vendors love LLMs. But it's just a highly inefficient approach for a lot of problems.
Clearly hardware vendors love LLMs. But it's just a highly inefficient approach for a lot of problems.
They would still be valuable for prototyping, where fast iteration makes it possible to learn more about the problem you are solving and whether it is even worth solving. They are also valuable for iteration-and-distillation approaches, where you can use the data generated from an expensive model to train a cheaper model.
So it looks like a VERY good idea. Who gets the gist of what I'm saying?
You'd get a repeat of Microsoft Tay, the short-lived 2016 chatbot that had to be taken down after not even a full day because 4chan managed to turn it into a full-blown hate spreader in hours [1].
Before something like this can be reasonably released to the Internet at large, we need to find out how to teach the "ingestion" part how to judge the input that it's being presented with... basically, AI school. The same way we teach our children that it's not OK to steal other people's stuff, to not be the one first throwing punches or to not discriminate against other people, or that there are reasonably trustworthy media and absolutely untrustworthy, we need to teach AIs. And we'd also need to figure out ways to teach an AI basic, hard truths: the Holocaust happened, the Earth is a globe not a pizza, the Earth is not hollow, and the moon landings were real.
At the moment, we're half-ass attempting that by annotated training data (and this is the true moat of OpenAI, not the weights or prompts!), but even as we invest literally millions of hours of training compute time, an average high schooler that has been trained for 18 years can reasonably pass the above-mentioned criteria to be a productive member of society.
As someone who's run into problems related to short context windows quite often, can you explain what "worth it" means? Also, when you say "not cheap", what alternative do you have in mind?
Of course it is true for real long context but it's not clear if its going to work with the sparse context hacks intended to keep memory size low.
Having the ability to keep the same coding session forever without having the LLM forgetting all the time and making the same mistakes over and over would be a game changer.
People like RAG-based solutions because you can include references in your answer very easily (eg, Perplexity, or see the DAnswer "internal search" product launched today on HN). That is extremely hard to make work reliably from a fine-tuned model.
They often want to see the reference - for example the LLM is constructing an answer about some company policy question it is good to include references to the actual company policy documents by URL. This is easy using RAG, but very hard to do reliably just using a fine tuned LLM.
See Perplexity for a great example of this at web scale.
But nonetheless, sending 2M tokens to a LLM isn't an efficient way to solve most problems.
I think there is sufficient evidence to think it works. A bunch of people who do have access have put it through a bunch of real-world exercises to test this, and while GPT4 128K and Claude 200K both fail Gemini is passing.
I agree sending 2M tokens to a LLM is costly at the moment and that RAG has a bunch of advantages in some circumstances.
Figuring out new algorithms is monumentally more difficult than more hardware, and even then when more efficient algorithms are found, we throw more hardware at it and get a million times more done.
However, caching might be a sweet spot for these multi-modal and large context LLMs. Take a bunch of documents and perform reasoning tasks to distill the knowledge down into something like a knowledge graph, to be used in RAG.