How many users would realistically be able to use it at the same time when running on such a device? I am interested in its scalability.
How many users would realistically be able to use it at the same time when running on such a device? I am interested in its scalability.
- memory bandwidth
- interconnect bandwidth between the CPU and GPU
- interconnect bandwidth between GPUs
- thermals and power if you're doing a good job of optimizing the rest
I don't see how a batching mechanism would improve on any of those, superficially it looks as though that would make matters worse rather than better. Can you explain where the advantage comes from?
https://groq.com/wp-content/uploads/2020/05/GROQP002_V2.2.pd... the "batching" section of https://docs.nvidia.com/deeplearning/tensorrt/archives/tenso... https://le.qun.ch/en/blog/2023/05/13/transformer-batching/
If you're doing inference on multiple prompts at the same time by doing batching, you don't take more time in streaming. But each streamed weights gets used for, say, 32 calculations instead of 1, making better use of the GPU's compute resources.
You could run it using OpenVINO on IntelCPUs, but the performance would probably take a hit. It would be a lot easier though since you can just use ggml.
Otherwise ~1.5 tokens/s is definitely the minimum you'd want streaming tokens to a single person.
You'll get a log_2-based scaling efficiency with nearly any batchsize increase, pending some limitations (memory, etc).
That should be enough at least to roughly sketch it out.