- Gemma3 12B: ~100 t/s on prompt eval; 15 t/s on eval
- MistralSmall3 24B: ~500 t/s on prompt eval; 10 t/s on eval
Do you know what different in architecture could make the prompt eval (prefill) so much slower on the 2x smaller Gemma3 model?
845 karma · joined June 6, 2015
- Gemma3 12B: ~100 t/s on prompt eval; 15 t/s on eval
- MistralSmall3 24B: ~500 t/s on prompt eval; 10 t/s on eval
Do you know what different in architecture could make the prompt eval (prefill) so much slower on the 2x smaller Gemma3 model?
They are great for reference if you already know the subject, but if you want to learn something, you'd be better off using introductory textbooks for the various fields you're interested in. I'm happy to make recommendations, and you can also look up syllabi for math undergrad at Stanford, Princeton, etc.
Two blog posts by a professor at U. Chicago, qualifying it of intellectual fraud:
https://www.galoisrepresentations.com/2019/07/17/the-ramanuj...
https://www.galoisrepresentations.com/2019/07/07/en-passant-...
https://www.reddit.com/r/MachineLearning/comments/8zm4kl/d_l...
No, it doesn't mean something different in French.
I also don't see how Mathematica's notation is any better; I view it as way worse.
https://www.npr.org/2019/12/16/788587668/the-efficient-chris...
Original Hollywood reporter article: https://www.hollywoodreporter.com/news/universal-notifies-th...