Presumably there's a host of applications that just need to tokenize, though, and this would be great for those!
Presumably there's a host of applications that just need to tokenize, though, and this would be great for those!
sglang_speed [huggingface]: mean=10.31ms median=6.48ms p99=45.98ms rps=96.8
sglang_speed [gigatoken]: mean=10.13ms median=6.54ms p99=45.16ms rps=98.4
input_len= 2048: TTFT mean 30.74 -> 29.05 ms (+5.5% reduction) | median 31.00 -> 28.80 (+7.1%) | p99 33.02 -> 32.02 (+3.0%)
input_len= 8192: TTFT mean 105.20 -> 96.36 ms (+8.4% reduction) | median 103.87 -> 95.49 (+8.1%) | p99 126.88 -> 113.84 (+10.3%)
input_len= 32768: TTFT mean 687.05 -> 633.66 ms (+7.8% reduction) | median 708.14 -> 657.35 (+7.2%) | p99 728.95 -> 678.79 (+6.9%)
These are preliminary numbers, so I will need to do some more testing before including this in the README.If I have the same number of CPU cores and they all can do their work in half the time they can double the number of requests now
Heck, even latency alone you can just reduce the standard deviation and get smoother flows. Little's law is a great callout here too, one of my favorite computer science principles.
Tips & Tricks on parameters/settings?
What happens at peak? Do people have to wait now? Increase of latency?
Source: https://www.gartner.com/en/newsroom/press-releases/2026-07-2...
Latency can be just as important as overall throughput, especially for inference providers like Groq and Cerebras.