Continuous batching to increase LLM inference throughput and reduce p50 latency
anyscale.com
anyscale.com
[1] https://github.com/huggingface/text-generation-inference/iss...
> For example, suppose you prompt with a sentence "What is the capital of California: ", it would take ten forward pass iterations to get back the full response of ["S", "a", "c", "r", “a”, "m", "e", "n", "t", "o"]
From the content I have been reading trying to understand LLMs, I thought that the output was a token and not a string of chars. What am I missing here?
It's an extremization that still is true for character-based models.
LLMs use a vocabulary of statistically chosen tokens. GPT 3 vocab, for instance, splits Sacramento into three tokens-
Sac - 38318
rament - 15141
o - 78
There is a rule of thumb that about every four letters in English text becomes a token but that's just the average.
California is a single token (25284). As is Canada (17940). And so on.
In the case of GPT-3 the vocabulary has a size of 50,257 tokens. GPT-4 increases that past 100k (see https://gist.github.com/s-macke/ae83f6afb89794350f8d9a1ad8a0...).
It's very similar to compression algos, really. Find recurring sets of characters.
Kinda hard to believe, but I have no intuition about this.
The speed people are just creating wrappers or minor changes and using words like "disruptive", "game-changing", "democratising <something>" just feel so inflated and boring at this point.
I hope this gets very soon into the next phase of the hype cycle[1], so the we can talk about something else. [1] https://en.wikipedia.org/wiki/Gartner_hype_cycle
To be fair, the overhead of running LLMs at scale -- the cost per interaction -- is the limiting factor for commercial deployments. Efficiency is an enormous advantage.