51 karma · joined April 14, 2012
If you want to keep using the same model, these settings worked for me.
llama-server -ngl 99 -c 262144 -fa on --cache-type-k q4_0 --cache-type-v q4_0 --host 0.0.0.0 --sleep-idle-seconds 300 -m Qwen3.5-35B-A3B-UD-Q4_K_XL.gguf
For the harness, I use pi (https://pi.dev/). And sometimes, I use the Roo Code plugin for VS Code. (https://roocode.com/)
I prefer simplicity in my tooling, so I can understand them easier. But you might have better luck with other harnesses.
I use llama-server that comes with llama.cpp instead of using ollama. Here are the exact settings I use.
llama-server -ngl 99 -c 192072 -fa on --cache-type-k q4_0 --cache-type-v q4_0 --host 0.0.0.0 --sleep-idle-seconds 300 -m Qwen3.5-27B-Q4_K_M.gguf
Links obviously VERY NSFW. [1] http://naughtyamericavr.com/ [2] https://virtualrealporn.com/
[1] https://blogs.msdn.microsoft.com/visualstudio/2014/02/17/int...
Before
-------
CPU model: Intel(R) Xeon(R) CPU L5520 @ 2.27GHz
Number of cores: 8
CPU frequency: 2266.788 MHz
Total amount of RAM: 988 MB
Total amount of swap: 255 MB
System uptime: 8 days, 12:03,
I/O speed: 69.9 MB/s
Bzip 25MB: 8.96s
Download 100MB file: 47.2MB/s
After------
CPU model: Intel(R) Xeon(R) CPU E5-2680 v2 @ 2.80GHz
Number of cores: 2
CPU frequency: 2800.086 MHz
Total amount of RAM: 1993 MB
Total amount of swap: 255 MB
System uptime: 2 min,
I/O speed: 638 MB/s
Bzip 25MB: 5.10s
Download 100MB file: 146MB/s
Test: https://github.com/mgutz/vpsbench