HN
Hacker News
Top
New
Best
Ask
Show
Jobs
Comment by RobMurray | Hacker News Reader
Parent
Full thread
RobMurray
·
why? it's mostly reads. the weights are static.
View on HN
bigyabai
·
llama-cpp's process is, but macOS itself will swap hard when 10-14gb of memory is paged for LLM inference. Dense models especially would thrash zram.
Reply on news.ycombinator.com