HNHacker News
TopNewBestAskShowJobs

gitpusher42

459 karma · joined November 17, 2018

submissionscomments
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
There was an ai winter for very long time. The math for NNs was already here, but not enough compute/data

I saw a pretty cool project to run an llm on an esp32 device https://github.com/slvDev/esp32-ai

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thank you!

afaik ollama relies on llama.cpp and mmap. mmap loads pages on demand and doesn't use the same explicit cache or parallel reads like my engine. Most likely ollama/llama.cpp will be way slower in this case

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
not with this engine. Kimi is a very different model. You can try to check HN later, I believe someone will build engine for this model for edge devices
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thank you! Let me know how it goes and share your tok/s results
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
I tried it. madvise didn't make mmap better than pread. I also tested F_RDADVISE, it helped on short decodes, but somehow got worse on longer decodes. Not very clear why, most likely problem somewhere at APFS and it is closed source and not much docs for it
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
m5 device also uses the same approach. the same 16 cache slots. and experts are evicted from memory as needed. And activity monitor shows 2gb usage for m5 pro (24gb btw)
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Uh, not really, unfortunately. Asahi uses Vulkan for gpu, but kernels for this project are metal.
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Yeah, correct! You can set up this engine to get more expert cache slots (e.g 32 instead of 16) to get a better hit rate and better tok/s. it will be 3.5gb instead of 2gb.
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
I double checked m2 logs. Cache hit rate is about 59-69%. 250-320MB went through `pread` per generated token. It is 3gb/s during this i/o phase.
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thank you very much!

I think it will throttle quite soon, but I haven't tried runs longer than 30minutes with this engine.

However, there is no constant load on ssd or gpu. i/o and gpu work are alternating and there is a brief idle periods for each i/o and gpu during inference (because gpu waits for i/o and after that i/o waits for gpu)

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
For Mac I would start from MLX engine. For exact model choice it is better to check bench results, and select model based on your need. A lot of good feedback about Qwen3.6, but I haven't used it in my tasks
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thank you! Text is not the main part of this repo. The main part is the technology and the list of experiments (and some useful knowledge I got from this project, haha). I have always been bad at writing or editing text (in both my native language and English), but without text it is impossible to share this project online.

Text was the last and most difficult part for me. It is not perfect (and this project is not perfect as well), but I believe it does the job of communicating my ideas

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
I am not native, my English is far away from perfect. I am using LLMs for checking my texts or grammar. I always trying to edit it properly, but sometimes I missing parts like that because I don't really have this "language feeling" as natives. Apologies for this
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
I think it depends on usage pattern. You trade speed for lower memory usage. Maybe engine specialisation and faster SSDs is the future for local inference, who knows
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thanks! I wanted to add hugging face token field to speed up model download, but then I realised that people might not trust to give their tokens And yeah, local models are better for security, at least your conversation stay on the machine
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
This approach will only work for MoE models. There is a Qwen 35b-a3b. You just need to do GPU stop after router and read the requested experts to ram. And it is possible to build similar engine for this model (or feel free to adopt my engine) Not sure about generic approach for now, but coding with ai agents is relatively cheap now, you can try it
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thank you! Thankfully this is purely software engineering problem. And as usual there is no free lunch. You trade speed for lower memory usage
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Size reduction is mostly based on Experts size. And it is limited by SSD speed. Check for Colibri and Flash-Moe, they are doing similar things with bigger models, but tok/s is not high
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Switching from mmap to parallel pread. From 0.5tok/sec to almost 4tok/sec. Running GPU work while reading missed experts also helped a lot, 4.4 -> 4.7
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thanks for sharing! SSD read speed is the biggest limiting factor here, unfortunately
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
I am relying quite heavily on system caching and pread. And yeah, M5 is a way faster and I can guess Mac can cache something, even if process stays under 2gb.

It was 83ms read per token for M2 and 12ms on M5 pro. Total is 163ms/tok vs 30ms/tok for M5. So yeah, there is a faster read and faster gpu processing

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
My friend just compared this project to running Cyberpunk on a very old machine at 16 FPS, huh
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thank you! Share your tok/s results later
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thank you! That’s useful. I might try lowering the minimum version later. The 2.4x prefill improvement will only work on the apple10 GPU family. The M1 uses apple7 as I remember
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
You’re right, Gemma isn’t the best model for coding (afaik more "everyday tasks" related). My first idea was to use Qwen, but its architecture was much more complex to implement in this stack. I chose Gemma so I wouldn’t spend all my time debugging custom kernels and could actually move the project forward with simpler approach
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Yeah! Check the Colibri and Flash-MoE projects. They’re already doing that.

https://github.com/danveloper/flash-moe https://github.com/JustVugg/colibri

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
It is super cool! Diffusion Gemma was released around the middle of my project, and I seriously considered switching to it. But I decided to finish the project as it was.

I believe it would be a perfect match!

Feel free to use any parts of my project or drop me a message. There’s my LinkedIn link at the end of the readme. Or I will drop you a message later!

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
I think I first saw Flash-MoE (https://github.com/danveloper/flash-moe) in April. Huge respect to them, it was a big inspiration for this project!
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
My first version used plain `mmap`. On the 8 GB M2, a cold 3.36 MB expert took 10 ms with mmap and 2.8 ms with `pread`. The full simulation was 0.50tok/s for `mmap` vs 4 tok/s for `pread`

With `mmap`, OS loads pages reactively as the model touches them. It doesn’t know which experts were selected or when their reads could overlap with GPU work

And common weights still use mmap for simplicity

So, I believe llama.cpp might run it under 2gb, but I assume it will be slower

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
What exact specs do you have? It might be because it's the 256 GB version. afaik, those versions have much slower memory bandwidth than the 512 GB models

My friend tried it on an M4 MacBook Pro and got 25–27 tok/s

← PreviousPage 2 of 3Next →