Time to first token, especially for smaller models, can be sharply reduced.
Latency can be just as important as overall throughput, especially for inference providers like Groq and Cerebras.
Latency can be just as important as overall throughput, especially for inference providers like Groq and Cerebras.