Infini-Gram: Scaling unbounded n-gram language models to a trillion tokens
arxiv.org
arxiv.org
From the paper, having a better n-gram index is giving as much perplexity improvement as a 10x increase in parameters. I wonder if it would make sense to train new foundation models with more attention heads and use infinigram in lieu of the feed-forward layer.
These are non-scientific concepts. You are basically saying "humans are doing something more, but we can't really explain it".
That assumption is getting weaker by the day. Our entire existence is a single, linear, time sequence data set. Am I "extrapolating from my experience" when I decide to scratch my head? No, I got a sequential data point of an "itch" and my reward programming has learned to output "scratch".
There are known faculties humans have that LLMs especially do not, such as actual memory, the ability to simulate the world independently via the imagination and structured thought, as well as facilities we don’t really understand but AIs definitely don’t have which are the source of our fundamental agency. We are absolutely able to create thought and reasoning without direct stimulus or as a response to something in the environment - and it’s frankly bizarre a human being can believe they’ve never done something as a reaction to their internal state rather than extrinsic.
LLMs literally can not “do” anything that isn’t predicated on their training set. This means, more or less, they can only interpolate within their populated vector space. The emergent properties are astounding and they absolutely demonstrate what appears to be some form of pseudo abductive reasoning which is powerful. I think it’s probably the most important advance of computing in the last 30 years. But people have confused a remarkable capability for a human like capability, and have simultaneously missed the importance of the advance as well as inexplicably diminished the remarkable capabilities of the human mind. It’s possible with more research we will bridge the gaps, and I’m not appealing to magic of the soul here.
But the human mind has a remarkable ability to reason, synthesize, extrapolate beyond their experience, and those are all things LLMs fundamentally - from a rigorous mathematical basis - can not do and will never do alone. Any thing that bridges that will need an ensemble of AI and classical computing techniques - and maybe LLMs will be a core part of a part of something even more amazing. But we aren’t there yet and I’ve not seen a roadmap that takes us there.
We want AI to hoop jump something we don't even understand ourselves. The only empirical evidence we have is as we increase compute, the results get better.
You could use the same framework to generate an internal dialog for a bot.
A lot of people don't think before they speak. If you tell me you have a small conversation with yourself before each thing you say out loud during a conversation, I will have doubts. Quick wit and fast paced conversation do not leave time for any real internal narration, just "stream of consciousness".
There is a time for carefully choosing and reflecting on your words, surely, but there are many times staying in tune with a real time conversation takes precedence.
> You could use the same framework to generate an internal dialog for a bot.
We can, for sure. But will it works? Given my (admittedly limited) experience with feeding LLM-generated stuff back in the LLM, I'd suspect it may actually lower the output quality. But maybe fine-tuning for this specific work-case could be a solution to this problem, as I suspect the instruction-tuning to be a culprit in the poor behavior I've witnessed (the bots have been instruction-tuned to believe the human, and apologize if you tell them they've made mistakes for instance, even if they were right in the first place, so this blind trust is likely polluting the results).
If you pay attention, you can catch that it is all just an illusion.
4 + 5 = 9
or 1729 is a taxicab number
if those phrases are in the database but not 4 + 05 = 9
5 + 4 = 9
11 + 3 = 14
48988659276962496 is a taxicab number
if those are not in the database.You can ask LLMs to generate poems on any topic in the style of specific authors. That's a rudimentary version of what you're describing.
Back then neural LMs were just beginning to emerge and we briefly experimented with using one of these unlimited N-gram models for pre-training but never got any results.
David Wheeler reportedly had the idea in the early 80s, but he rejected it as impractical. Then he tried publishing it with Michael Burrows in the early 90s, but it was rejected from the Data Compression Conference. Algorithms researchers found the BWT more interesting than data compression researchers, especially after the discovery of the FM-index in 2000. There was a lot of theoretical work done in the 2000s, and the paper mentioned by GP cites some of that.
The primary application imagined for the FM-index was usually information retrieval, mostly due to the people involved. But some people considered using it for DNA sequences and started focusing on efficient construction and approximate pattern matching instead of better compression. And when sequencing technology advanced to the point where BWT-based aligners became relevant, a few of them were published almost simultaneously.
And if you ask Heng Li, he gives credit to Tak-Wah Lam: https://lh3.github.io/2024/04/12/where-did-bwa-come-from
[1] https://snats.xyz/pages/articles/from_bigram_to_infinigram.h...
1) this reminds me of how I was thinking about “what if you tried to modify an n-gram model to incorporate a poor-man’s approximation of the copying heads found in transformer models? (one which is also based just on counting)”
2) What if you trained a model to, given a sequence of tokens, predict how many times that sequence of tokens appeared in the training set? And/or trained a model to, given a sequence of tokens, predict the length of longest suffix of it which appears in the training set. Could this be done in a way that gave a good approximation to the actual counts, while being substantially smaller? (Though, of course, this would fail to have the attribution property that they mentioned.)
3) It surprised me that they just always use the largest n such that there is a sample in the training set. My thought would have been to combine the different choices of n with some weighting.
Like, if you have one example in the training set of “a b c d”, and a million examples of “b c e”, and only one or two other examples of “b c d” other than the single instance of “a b c d”, does increasing the number of samples of “b c e” (which aren’t preceded by “a”) really give no weight towards “e” rather than “d” for the next token?
1) If we view N gram counts as observation counts, we can generalize Norman Megill's estimation of Bernouilli probabilities to manysided dice (as many sides as token alphabet):
(O+1)/(T+S) where O is the number of Occurences, T is the number of Trials and S is the Size of the alphabet or number of Sides.
Having 0 observations does not result in probability of 0.
2) Since the method allows for collecting occurences in the corpus, this could be used to provide alternative language models contexts in the corpus.
[1] https://www.researchgate.net/publication/2473004_Unbounded_L...
That is from 2003 and actually quite interesting. It shows, for example, that these models were applied to rather small texts, because of the presence of context with length of -1 (uniform distribution). When text is large the need for such context vanishes, but for small texts it can give an advantage.
EDIT: Oh so this thing takes 10TB disk space to keep an index while a LLM takes… 175GB (assuming GPT-3 in fp16). A huge resource requirement that cannot be ignored.
You can cluster data using unsupervised random forests and then use these cluster indices as features.
For example doing a sum of random numbers. If the token you are trying to predict is not in the training data, even if similar patterns exist, this model defaults to the Neural Model.
I guess then it is an aide to the neural model on filling the easy patterns.
Weighted Finite State Transducers are where they really shine; you tend to want to do massive parallel operations, and you want to do the without being bottle necked on memory access. I think these challenges make this framework challenging to adopt.
[1]: https://github.com/alphacep/vosk-api/issues/55
[2]: https://github.com/outlines-dev/outlines?tab=readme-ov-file#...
It's interesting as speech recognition has become more popular than ever through services like Alexa, and other iot devices support for OS speech recognition has very little development. Don't get me started with accessibility apis either...
Unfortunately most implementations (especially those that are iot focused) don't have very important features for robust speech recognition.
1. Ability to enable and disable a grammar
2. Modify grammars while the engine is loaded
3. Scoped grammars that are context-specific
4. Recognition callbacks
5. Multiple grammars active simultaneously.
Unfortunately I don't think vosk api will ever support those features. I know there's a few PRs that address a few of those points but have not been merged for years.
Given the criteria above there's very little open source that allows for complex grammars that's easy to run for an end user locally.