HNHacker News
TopNewBestAskShowJobs

gitpusher42

459 karma · joined November 17, 2018

submissionscomments
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
The full route changes almost every token. The cache works through partial reuse, about 40% of experts repeat on the next token and 57% within two tokens, cutting I/O from 166 to 88 ms/token on M2 Mac.

The longest exact repeat we found was only two tokens. Coding tasks may have higher reuse if code related experts are selected repeatedly

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
AFAIK it should not because it is only reading
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
The process stays at around 2 GB with 16 slots and a 4K context on both the M5 and M2. But yeah, Apple might be doing some magic under the hood
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Uh, I’m afraid it is Apple only. It is written using Apple’s GPU language, Metal, and heavily relies on the Apples’s shared memory architecture

Windows PCs would require a completely different approach

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thank you! Under good conditions it achieves approx a 67% cache hit rate with 16 expert slots
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
I haven't tried it but it should work! You can try it and share your results, it would be really appreciated

I tried it on my wife's M1 MacBook Air 512GB and it gets 4–5 tok/s

Also, it must be easy to adjust for iPhones and iPads in theory

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
It depends on the use case.

I measured this exact model with a 4k context on the mlx engine. It runs at 75 tok/s on my M5 Mac Pro and using 14 GB of RAM. For my engine the same model uses 2 GB of RAM and produces 31–35 tok/s.

The project is still experimental so performance may vary as it continues to improve. If you want to save around 12 GB of RAM for other tasks and you are ok with 35 tok/s (afaik it is roughly comparable to ChatGPT’s speed for basic responses) my engine may be a good fit.

If you need maximum speed and flexibility just use MLX

← PreviousPage 3 of 3