Accelerating Generative AI with PyTorch II: GPT, Fast
pytorch.org
pytorch.org
Code can also be found here: https://github.com/pytorch-labs/gpt-fast
And a twitter thread summary here: https://twitter.com/cHHillee/status/1730293330213531844
Is this faster than HuggingFace's Text Generation inference container?
One note is that this release is optimized for latency, while I think HF TGI might be more optimized for throughput.
For example, there's an AMD backend for Triton (and it's also integrated into torch.compile), which is why we can mostly do the same optimizations on Nvidia and AMD GPUs.
Ideally, there'd be an Apple Silicon backend for Triton, and then this repo would mostly work out of the box :)
Even ignoring that, most of the development is running experiments. You're gonna be hesitant to run lots of experiments if they each cost money whereas when you pay upfront for the hardware, you're gonna have the incentive to fully utilize it with lots of experiments.
I'd go with rtx 4090 and deal with memory limitation through software tricks. It's an underrated card that's as performant as cards that are magnitude pricier. It's great way to get started with that budget.
z790 chipset w/ mobo that supports x8/x8 bifurcation
96gb ddr5 @5600mhz
any recommendations on how to buy one? e.g. 24GB model, any particular model to run LLMs? what is the biggest baddest LLM you can run on a single card?
have been thinking about it but was sticking with cloud/colab for experiments so far.
I'm happy with my 4090 though. Dealing with splitting between GPUs sounds like a chore and also I like the gaming abilities of the 4090.
If you also want a Mac consider an m3 MacBook with maxed ram.
Is it just a difference in speed, or are there some new theory details to learn?
Thanks for sharing your work.
How did they unlock this key ? In retrospect it seems so simple, but without the KV-cache this possibility would not have emerged at all. Hats off !
> While these projects are performant, they often come with tradeoffs in ease of use, such as requiring model conversion to specific formats or building and shipping new dependencies.
I think it should be acknowledged that (at least IMO) pytorch model formats are not very portable and this is a big part of the problem. It would be nice to see industry move towards a better format (gguf?) that can easily be ported between frameworks and not leave you stuck using torch to load it. Likewise, pytorch is a massive dependency to include with a project, especially for simple inference, so while other projects have new dependencies, they can often be a lot lighter than for a pytorch model, again particularly for inference code.
However, I do think these model conversions are often a significant pain for users.
So, in some sense, the goal here is to show that the performance component and the "convert your model for deployment" component can be disentangled.
We also have work on allowing you to "export" an AOT-compiled version of your model with torch.compile, and that should allow you to deploy your models to run in other settings.
Also, I liked the part of the article about torch.compile producing faster matrix-vector multiplication than cublas. I've seen the same thing on CPU, that it's way faster to just write and manually optimize a loop over a bunch of dot products than it is to use BLAS routines because of how simple the "matmul" actually is. I don't know how widely known that is.
Also… where are all the people like us that work on applications on top of HF models congregate?
it's just the pytorch profiler + chrome profiler (chrome://tracing)
Is that possible with "just" pytorch? Could it be added to gpt-fast?
long-term, bodes well as both methodologies should be combineable, just curious for current point-in-time wrt being relevant for production use
What's a good use case for an order of magnitude decrease in price per token? Web scale "analysis" or cleaning of unstructured data?
Most use cases outside of classic chat.
For example, I made an on-demand educational video project, and the slowest part was by far the content generation. RAG, TTS, Image generation, text rendering, and video processing were all a drop in the bucket, in comparison.
It would be an even wider gap now, and TTS is super-realtime, and image generation can be single step.
- Hook LLM to VMs
- Ask for code that [counts to 10]
- Run code on VM
- Ask different LLM to Evaluate Results.
- Repeat for sufficient volume.
- Train.
The faster it can generate results the faster those results can be tested against the real world, e.g. a VM, users on X, other models with known accuracies.
* reading: If you want it to do inference over a lot of context, you'll need to do multiple inferences. If each inference is faster, you can 'read' more in the same time on the same hardware
* thinking: a lot of analytical approaches essentially use writing as both memory & thinking. Imagine iterative summarization, or automatically iteratively refining code until it's right
For louie.ai sessions, that's meant a fascinating trade-off here when doing the above:
* We can use smarter models like gpt-4 to do fewer iterations...
* ... or a faster but dumber model to get more iterations in the same amount of time
It's entirely not obvious. For example, the humaneval leaderboard has gpt4 for code being beat by gpt 3.5 for code when run by a LATS agent: https://paperswithcode.com/sota/code-generation-on-humaneval . This highlights that the agent framework is the one really responsible for final result quality, so their ability to run many iterations in the same time window matters.
So, in practice, a full "text completion request" can often take on the order of seconds, which dwarfs the client <-> server roundtrip.
I mean, practically speaking, completions from say, ChatGPT or Claude take seconds to finish :)
Also how does that interact with MoE models? Do you have a mini version of the MoE, with smaller experts?
Anecdotally, folks often seem to use say, 70B base + 7B as verifier. But I think there's a lot of room for experimentation and improvement here.
You could... say, take a 70B model and maybe just chop off the last 90% of layers and then fine-tune. Or perhaps you could use a model that's trained to generate 8 tokens at once. Or perhaps you could just use statistical "n-gram" predictor.