459 karma · joined November 17, 2018
I think I might understand your use case. You ssh the Mac and want something like `ollama run` with an interactive chat in terminal. Am I right?
There is already experimental OpenAI-compatible server in this repo:
``` swift build -c release --product TurboFieldfareServer .build/release/TurboFieldfareServer \ --model scratch/gemma4.gturbo ```
After that a small terminal client can run inside the same ssh session and talk to `/v1/chat/completions`
The client needs to keep a messages array, add each user message, send the full array with `stream:true`, print SSE chunks until `[DONE]`, then add the response back to the array. `/reset` can clear it
There is a python example in the server docs. (https://github.com/drumih/turbo-fieldfare/blob/main/docs/OPE...)
It is non-streaming, but can be used as a starting point.
The server is still experimental and I am fixing some problems currently. But you can try to vibecode a simple terminal client around it.
If not, create an issue on Github and describe desired behaviour
For CLI and Server, use --max-context
mmap benchmark did basically page touch experiment and cold reads were much slower, unfortunately (10ms vs 3ms)
I tried MADV_WILLNEED, F_RDADVISE and preadv. preadv reduced parallelism because requested experts are rarely adjacent in the file.
pread is still the fastest. And I think Flash-Moe got the same result too
Owen 3.6-35b-a3b was my initial idea, but I switched to Gemma because of its simpler architecture and kernels
https://github.com/JustVugg/colibri https://github.com/antirez/ds4 https://github.com/danveloper/flash-moe
after that download 14gb of weights and enjoy offline inference (and a bit of Gemma4 intelligence) for your everyday tasks
multi turn chat is coming!
And there is some reuse. ~41% selected again for the next token, ~57% within two. Each layer has its own experts, so no reuse between these layers.