HN
Hacker News
Top
New
Best
Ask
Show
Jobs
Comment by djoldman | Hacker News Reader
Parent
Full thread
djoldman
·
How much VRAM does it use during inference?
View on HN
stellaathena
·
~40 GB with standard optimization. I suspect you can shrink it down more with some work, but it would require significant innovation to cram it into the next largest common chip size (24 GB, unless I’m misremembering)
komuher
·
Is 40GB already on float16?
stellaathena
·
Yes
Reply on news.ycombinator.com