Let's say I have a book and I want to ask multiple questions about it. Every query will pay the price of the book's text. It would be awesome if I could "index" the book once, i.e. pay for the context once, and then ask multiple questions.
Let's say I have a book and I want to ask multiple questions about it. Every query will pay the price of the book's text. It would be awesome if I could "index" the book once, i.e. pay for the context once, and then ask multiple questions.
Or anything where the question answer isn't 'close' to the words used in the question?
How well does this work vs giving it the whole thing as a prompt?
I assume worse but I'm not sure how this approach compares to giving it the full thing in the prompt or splitting it into N sections and running on each and then summarizing.
The problem is that humans have continuous information retrieval and storage where the current crop of embedding systems are static and mostly one shot.
This weird leaky memory has advantages and disadvantages. Forgetting is useful, it removes garbage.
Machine models could vary the balance of temporal types, drop out Etc. We may get some weird behavior.
I would guess we will see many innovations in how memory is stored in systems like these.
Background: https://summarity.com/hyde
Demo: https://youtu.be/elNrRU12xRc?t=1550 (or try it on findsight.ai and compare results of the "answer" vs the "state" filter)
For even deeper retrieval consider late interaction models such as ColBERT
Does the embedding structure somehow expose the themes? And if so, is it more the embeddings that are answering the question by how it groups things?
Just a few manufacturers hold the effective cartel monopoly on LLM acceleration and you best bet they will charge out the ass for it.
Barring a Chinese invasion of Taiwan, these APIs will halve in price over the next year.
On the other hand, for all peoples worrying about China, they are pretty restrained given the enormous turmoil the invasion would cause. If they wanted to break the USA, now would probably be the time to do it?
https://www.theregister.com/2023/03/14/us_china_tsmc_taiwan/
Ironically, that's an example I like to list as "pure sci-fi fantasy, divorced from economic reality."
The total cost of the iron ore that goes into making a new a car is about $200-$300 dollars, depending on various factors (size of the car, ore spot price, etc...).
Even if -- magically -- asteroid mining made not just "iron ore", but specifically the steel alloy used for car bodies literally free, new cars costing $30,000 would now cost... $29,700.
You can save more by skipping the optional coffee cup warmer, or whatever.
In reality: 90% of iron and steel is recycled, and asteroid mining is not magic.
It's stuff like platinum and germanium that makes asteroid mining potentially interesting.
On Earth, geological processes concentrate elements into ores, primarily through volcanic and hydrological means. Neither are available in small, cold asteroids devoid of liquid water. Hence asteroids are generally undifferentiated mineralogically, making mining them much less economically viable.
You often see total quantities listed as an amazing thing, glossing over the fact that the Earth has more of everything and in usefully concentrated lumps.
It makes the economics slightly trickier.
Perhaps it's because under the hood there's additional safety analysis/candidate generate that is resource intensive?
[1] I'm not sure if these huge context lengths are achieved the same way (i.e. a single input vector of length N) but given the cost is constant for input I would assume the resource usage is too.
edit - This occurred to me after the fact but I wonder if the difference is that the use case I work with is processing batches of many different embedding requests (but computed in one batch), therefore it has to process `min(longest embedding, N)` tokens so any individual request in theory has no difference. This would also be the case for Anthropic however.
I could imagine something where encoders pad up to the context length because causal masking doesn't apply and the self attention has learned to look across the whole context-window.
[1] Original Google Paper - https://arxiv.org/abs/1706.03762
[2] Original GPT Paper - https://s3-us-west-2.amazonaws.com/openai-assets/research-co...
If you do the math of how much memory bandwidth is required by a forward pass vs. how much compute, you'll see that inference is entirely limited by memory bandwidth and will use compute resources very inefficiently. In contrast, input processing is able to fully use the available compute.
Of course, there are ways to mitigate this problem, like processing multiple token streams in parallel, but the fundamental problem remains.
Otherwise, it might make sense to have a separate routine which compresses the context as efficiently as possible. Auto encoder?