What was even more impressive is the 0.6B model which made the sub 1B actually useful for non-trivial tasks.
Overall very impressed. I am evaluating how it can integrate with my current setup and will probably report somewhere about that.
What was even more impressive is the 0.6B model which made the sub 1B actually useful for non-trivial tasks.
Overall very impressed. I am evaluating how it can integrate with my current setup and will probably report somewhere about that.
Which I find even more impressive, considering the 3060 is the most used GPU (on Steam) and that M4 Air and future SoCs are/will be commonplace too.
(Q4_K_M with filesize=18GB)
Conversely, the 4B model actually seemed to work really well and gave results comparable to Gemini 2.0 Flash (at least in my simple tests).
I haven't evaled these tasks so YMMV. I'm exploring other possibilities as well. I suspect it might be decent at autocomplete, and it's small enough one could consider finetuning it on a codebase.
The /think and /no_think commands are very convenient.
Here’s the LM Studio docs on it: https://lmstudio.ai/docs/app/advanced/speculative-decoding
I'm running Q4 and it's taking 17.94 GB VRAM with 4k context window, 20GB with 32k tokens.
This part, yes. I assume the setting a complete environment is a little more involved than the 4 commands sibling is also refers to.
As a Windows/MacOS/Linux dweller kinto is a godsend so I can have macos keyboard (but you could have linux or windows by default) on all OSes https://kinto.sh/
E.g.: I go a little bit overboard for the average macOS user:
- custom system- and app-specific keyboard mappings (ultra-modifier on caps-lock; custom tabbing-key-modifier) via Karabiner Elements
- custom trackpad mappings via BetterTouchTool
- custom Time Machine schedule and backup logic; you can vibe-code your install script once and re-use it in the future; just make it idempotent
- custom quake-like Terminal via iTerm
- shell customizations
- custom Alfred workflows
- etc.
If all you need is just a sensible package manager and the terminal to get started, just set up Time Machine with default settings, Homebrew, your shell, and optionally iTerm2, and you're good to go. Other noteworthy power-user tools:
- Hammerspoon
- Syncthing / Resilio Sync
- Arq. Naturally, the usual backup tools also run on macOS: Borg, Kopia, etc.
- Affinity suite for image processing
- Keyshape for animations on web, mobile, etc.
As a Python person I've found uv + MLX to be pretty painless on a Mac too.
The latter is super easy. Just download the model (thru the GUI) and go.
Soon some AMD Ryzen AI Max PCs will be available, with unified memory as well. For example the Framework Desktop with up to 128 GB, shared with the iGPU:
- Product: https://frame.work/us/en/desktop?tab=overview
- Video, discussing 70B LLMs at around 3m:50s : https://youtu.be/zI6ZQls54Ms
edit: ok.. i am excited.
The quality of the output is decent, just keep in mind it is only a 30B model. It also translates really well from french to german and vice versa, much better than Google translate.
Edit: for comparision, Qwen2.5-coder 32B q4 is around 12-14t/s on this M1 which is too slow for me. I usually used the Qwen2.5-coder 17B at around 30t/s for simple tasks. Qwen3 30B is imho better and faster.
[1] parameters for Qwen3: https://huggingface.co/Qwen/Qwen3-30B-A3B
[2] unsloth quant: https://huggingface.co/unsloth/Qwen3-30B-A3B-GGUF
[3] llama.cpp: https://github.com/ggml-org/llama.cpp
It's using 20GB of memory according to ollama.