HNHacker News
TopNewBestAskShowJobs

gitpusher42

459 karma · joined November 17, 2018

submissionscomments
gitpusher42··on Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
Yeah, original TurboFieldfare supports Gemma https://github.com/drumih/turbo-fieldfare
gitpusher42··on Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
Thank you for using TurboFieldfare as a starting point for this project and thank you for mentioning it at the README. I am glad it inspired more people to explore area of on-device AI further!
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
It is only a wild guess, but if you have 256gb version and a lot of apps running it can be pretty slow.
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thank you for testing and sharing results!

I think I might understand your use case. You ssh the Mac and want something like `ollama run` with an interactive chat in terminal. Am I right?

There is already experimental OpenAI-compatible server in this repo:

``` swift build -c release --product TurboFieldfareServer .build/release/TurboFieldfareServer \ --model scratch/gemma4.gturbo ```

After that a small terminal client can run inside the same ssh session and talk to `/v1/chat/completions`

The client needs to keep a messages array, add each user message, send the full array with `stream:true`, print SSE chunks until `[DONE]`, then add the response back to the array. `/reset` can clear it

There is a python example in the server docs. (https://github.com/drumih/turbo-fieldfare/blob/main/docs/OPE...)

It is non-streaming, but can be used as a starting point.

The server is still experimental and I am fixing some problems currently. But you can try to vibecode a simple terminal client around it.

If not, create an issue on Github and describe desired behaviour

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
I think there is a limit based on MoE number of active parameters and quantisation, bytes count for active experts. But I believe we will see more project like this for different models.
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Uh, maybe, who knows. I am pretty bad solving leetcode, btw
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Yeah, gpt-oss-120b is also MoE, so the same ssd-streaming and caching ideas should work. Feel free to fork and try implementing it!
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Yeah, it must be exactly the same. The same weights are used, nothing skipped or pruned. But it might have differ to MLX for greedy decode because of small floating-point nums difference
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
yeah, looks like a page cache matters a lot I tested on mine m5 pro with 8gb memory pressure, got 27t/s instead of 35t/s Someone tested on m4 max. In regular state it was 48tok/s, but 32-42 with memory pressure
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Apple does something similar with their latest foundation model. They process input prompt and based on results they preload required experts. Quite neat solution for the edge devices
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Yeah, sure! You can select different options in the app settings at the right panel, it shows how much memory it will use

For CLI and Server, use --max-context

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
uh, tried most of this

mmap benchmark did basically page touch experiment and cold reads were much slower, unfortunately (10ms vs 3ms)

I tried MADV_WILLNEED, F_RDADVISE and preadv. preadv reduced parallelism because requested experts are rarely adjacent in the file.

pread is still the fastest. And I think Flash-Moe got the same result too

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
afaik there is some research at this area. Also the new apple foundation model uses related idea. they process the whole prompt and based on prompt load required experts and use only these experts for generation. It doesn't require fitting full model into memory or per token ssd streaming
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Yeah, must be possible. Not fast, but possible if you have enough ram. I think you can search online for projects, I think I saw something related
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
uh, don't worry. Just install the latest Xcode from the App Store. It includes everything you need to run this project
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
It heavily relies on M-series Mac unified memory architecture. And shaders are written using Metal, Apple's own gpu programming technology. It cannot be ported directly to classic architecture (ram+vram)
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Not sure it will be really usable. Check for Flash-Moe and Colibri repos A lot of request for qwen3.6 moe, it might worth exploring
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Yeah, Gemma is not the best for coding I guess. qwen must be better
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Yeah, the same ideas should work for qwen. You can try porting this engine to use Owen.

Owen 3.6-35b-a3b was my initial idea, but I switched to Gemma because of its simpler architecture and kernels

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thank you for testing and sharing, it is useful info!
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thank you very much for sharing! Great results and useful info!
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thanks! Will try!
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
uh, it's a bit difficult to discuss the classical approach with vRAM and regular RAM. Not really familiar with optimisations and hacks, I always worked with apple platforms and shared memory. But description sounds cool, good luck with your project!
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Check for colibri, dwarf star and flash-moe. they do similar things with bigger models

https://github.com/JustVugg/colibri https://github.com/antirez/ds4 https://github.com/danveloper/flash-moe

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Yeah, I tried both rearranging experts on disk and predicting the next expert using statistical approach. Reordering helped on the test prompt, but failed on another prompt. Markov and cross layer prediction didn't work either
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
It is Apple platform only implementation because of Metal (and Swift). Other platforms would require CUDA or Vulkan and a complete rework
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
hm. just open repo, copy commands into your terminal and you will get app installed (if you have swift toolchain installed)

after that download 14gb of weights and enjoy offline inference (and a bit of Gemma4 intelligence) for your everyday tasks

multi turn chat is coming!

gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Thank you! If you can use it for your tasks I would be happy!
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
uh, I don't think it is possible to compare them. DwarfStar4 is for high end macs and a lot of ram. this project is more targeted to low end devices and "general use" Gemma4 model
gitpusher42··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Yeah, I checked it. One expert is about a 3.36mb block. If a cache miss happens I read whole block with one pread.

And there is some reuse. ~41% selected again for the next token, ~57% within two. Each layer has its own experts, so no reuse between these layers.

Page 1 of 3Next →