HN
Hacker News
Top
New
Best
Ask
Show
Jobs
Comment by ancientworldnow | Hacker News Reader
Parent
Full thread
ancientworldnow
·
This was trained to be run at FP8 with no quality loss.
View on HN
hislaziness
·
The model description on huggingface says - Model size - 12.2B params, Tensor type - BF16. Is the Tensor type different from the training param size?
Reply on news.ycombinator.com