Self-Retrieval: Building an information retrieval system with one LLM
arxiv.org
arxiv.org
> To accurately generate the exact passages in the given corpus, we employ a trie-based constrained decoding algorithm (Chen et al., 2020; Cao et al., 2021; Lu et al., 2021) in which the generated tokens can be constrained in the dynamic vocabulary. Specifically, instead of generating a token from the entire target vocabulary at each step, we use a prefix tree (trie) to constraint the target vocabulary and ensure that the generated content is within the corpus. During the construction of trie, we remove stop words from the initial token to improve semantic representation of the trie.
I might not have a clear idea how the inference loop works but it sounds like a whole host of solutions could present themselves if it was easy to plug in various types of logic at inference.
Llama.cpp and its downstream users like ollama have long supported using BNF grammars to constrain the output that way.
Use trie constraint for quotes.
Use BNF for grammar languages (json, python etc).
Projects like llama.cpp/ollama should make it automatic/dynamic, just rely on triple quote sections where you enter those constraint modes automatically.
Ie. every time you enter into section starting with "```json" you automatically switch to JSON BNF.
Every time you enter into "```json:Foo" you enter JSON BNF + JSON-SCHEMA for Foo object definition.
"```python" for python grammar etc.
"```quote:documentRef" you enter trie based constraints.
"```llm:otherllm" you enter other llm.
"```whatever:whatever" you enter whatever you want.
If you want just json output you start output with "```json" and that's it.
As its all inference time it could be plugin based, ie:
1. character based - given input (from the start of opening "```foo") it returns allowed next characters, or
2. token based - same as above but returns allowed native tokens (not sure how performance would behave here, would it be acceptable?)
IMHO also very interesting area would be exploring stable AST representations for programming languages (a'la darklang I believe?) – where variable names are detached from AST itself, ie. differently named functions that otherwise have the same structure have precisely the same AST representation. This would dramatically reduce space to navigate around.
ps. [0] in case somebody else is interested
Llama.cpp could have upstream support for it but for them to consider it will have to be ironed out idea.
In llama.cpp support may have some different, more low level kind of support, I don't think it makes sense to attach state logic to text input there.
I think there are a lot of methods from these older Markovian setups that can be employed in the outputs samplers of modern models, as well as the inclusion of structured searches and so on. Parts of deep learning have always focused on structured output search, but historically the LLM style generative setting has not employed these approaches (though I find beam search for generative settings needs tweaking, it usually works pretty well in smaller scale problems for me).
[0] Avoiding Plagiarism in Markov Sequence Generation, Papadopoulos et. al. https://axon.cs.byu.edu/Dan/673/papers/papadopoulos.pdf
This is clever, I haven't seen an effective way to train an LLM to search a document yet and I can imagine this being very effective. I suppose this relies on the over fitting you get when fine tuning on a very small dataset.
What's funny about them is that it's a fairly involved procedure that turns your language model into an actual stochastic parrot, both showing that such a model useful and demonstrating that the original parrot concept was rather ill conceived.
Perhaps such a custom LLM will be available on HF, Ollama, etc. before long?
They train the LLM directly on the corpus so that the documents are embedded in its weights.
There are some ESL issues in the paper but it's not too hard to understand.
they do not outright say that in the paper as far as i could tell. i only got it from reading hn comments. just very confused why they use a nonstandard term like "internalize" which just pisses me off because ML is hard enough without inventing your own terms