HN
Hacker News
Top
New
Best
Ask
Show
Jobs
Comment by goldenarm | Hacker News Reader
Parent
Full thread
goldenarm
·
How many tokens per second?
View on HN
LuxBennu
·
Roughly 8-12 token/s on generation depending on context length. Prompt processing is faster obviously. Haven't benchmarked it super carefully though, just eyeballing the llama.cpp output.
Reply on news.ycombinator.com