inferMLX: Simple Llama Model LLM Inference in macOS with MLX
github.com
github.com
For example, it lets you make DeepSeek Llama 8B 'think' for longer by suppressing its thought closing tokens, or you can limit its thinking process to a certain number of tokens in just a few lines of code. This helps the smaller model solve things it gets wrong out of the box.