HN
Hacker News
Top
New
Best
Ask
Show
Jobs
Comment by nik736 | Hacker News Reader
Full thread
nik736
·
Which models will this be able to run at an acceptable token/s rate?
View on HN
simlevesque
·
gpt-oss:120b
https://til.simonwillison.net/llms/codex-spark-gpt-oss
hamdingers
·
Am I missing it or is there no information about performance? Looking for a tokens/sec
simlevesque
·
He didn't give that info but the transcript linked at the end shows how much time was spent for each query.
aseipp
·
Right now I get 59 tok/sec on GPT-OSS 120B using Unsloth's dynamic 4-bit quants, via llama.cpp
https://news.ycombinator.com/item?id=45881049
Reply on news.ycombinator.com