So a context size of 100k requires 100x more compute than a prompt size of 1k.
For which applications is that worth it?
Note you could reduce the cost to less than linear by using a retrieval method, but I don’t think that is what is being proposed.
So a context size of 100k requires 100x more compute than a prompt size of 1k.
For which applications is that worth it?
Note you could reduce the cost to less than linear by using a retrieval method, but I don’t think that is what is being proposed.
Programming, primarily. Code takes many more tokens per kilobyte of text than written English. So even quite short blocks of code eat up a lot of tokens.
The current AIs can do trivial, generic things using popular libraries. None can really help make changes in a large proprietary codebase where the prerequisite knowledge is the structure, design, and APIs of the private code.
With 100K token windows, a model could be given entire database schemas, or reams of interface definitions, Rest API schemas, or whatever, and then make edits based on that context.
It wouldn't even matter if it was slower than human, as long as it was cheaper.
Look at it this way: An 8-GPU NVIDIA DGX server is what, $400K to purchase at retail pricing? That would be "good enough" to run really beefy LLMs. If you use that server for about 3 years, then even factoring in all ancillary costs, that's about $13/hour to run. Or about 30 cents per minute.
So even if it takes some huge 100K token super-smart model a full minute to run through a prompt like "given the following reams of context, find the bugs in the given code below", then that's almost certainly worth it to most dev shops. Bugs can cost thousands of dollars to find and fix.
Merely finding half of the bugs for mere cents per function could yield staggering savings.
I'l also curious how adding 100x of compute into a longer prompt compares with using the compute for something else.
I'm sure there is a design space exploration paper out there or waiting to be written comparing the recent long prompt models against other uses of 100x more compute than an LLM.
For example, is it better to have 100x longer prompts, or 100x bigger models?
Ultimately we need both. There's no point having a superintelligent code assistant which doesn't have enough working memory to understand what your program is doing. And there's no point having 100x longer prompts if the system isn't smart enough to contribute code changes.
I think we can have both, but we'll need to do more work on our language models first. I mean, humans have extremely limited working memory, but we can work on arbitrarily large programs. We do it by paging context in and out of our minds. As such, I don't need to think about the entire google chrome codebase to make a change to one small part of it.
I'm interested in the approach of the LongMem paper (from Microsoft Research). As I understand it, their approach does something like humans where the system learns to page parts of the input in and out of working memory as needed. (I haven't read the paper in detail yet).
> > For which applications is that worth it?
> Programming, primarily
...and law and medicine where a doctor/lawyer may want to prompt with lots of information on their patients/clients and possibly related cases.Certain scientific experiments (like the LHC) deal in enormous quantities of data that may be desirable to include in a prompt.
That's just a few examples off the top of my head. I'm sure there are plenty of others. Programming is far from the only application.
> Conditional computation avoids applying all model parameters to all tokens from the input sequence. [CoLT5] applies heavy computations only to the most important tokens and processes the rest of the tokens with a lighter version of layers. It will speed up both training and inference.
[CoLT5]: https://arxiv.org/abs/2303.094752
This seems like part of the conversation where advances in software are expected to deliver bigger gains than hardware for a while.
>At inference time, thanks to ALiBi, MPT-7B-StoryWriter-65k+ can extrapolate even beyond 65k tokens. We demonstrate generations as long as 84k tokens on a single node of 8 A100-80GB
https://huggingface.co/mosaicml/mpt-7b-storywriter
Edit: here is a snip of the interview - https://share.snipd.com/snip/93d96e73-9841-4361-bd82-ae086e1...