I think I might understand your use case. You ssh the Mac and want something like `ollama run` with an interactive chat in terminal. Am I right?
There is already experimental OpenAI-compatible server in this repo:
``` swift build -c release --product TurboFieldfareServer .build/release/TurboFieldfareServer \ --model scratch/gemma4.gturbo ```
After that a small terminal client can run inside the same ssh session and talk to `/v1/chat/completions`
The client needs to keep a messages array, add each user message, send the full array with `stream:true`, print SSE chunks until `[DONE]`, then add the response back to the array. `/reset` can clear it
There is a python example in the server docs. (https://github.com/drumih/turbo-fieldfare/blob/main/docs/OPE...)
It is non-streaming, but can be used as a starting point.
The server is still experimental and I am fixing some problems currently. But you can try to vibecode a simple terminal client around it.
If not, create an issue on Github and describe desired behaviour