It simply did not want to use XML tools for some reason something that even qwen coder does not struggle with: https://discuss.samsaffron.com/discourse-ai/ai-bot/shared-ai...
I have not seen any model including sonnet that is able to 1 shot a working 9x9 go board
For ref gpt-4o which is still quite bad https://discuss.samsaffron.com/discourse-ai/ai-bot/shared-ai...
I've been asking it to perform relatively complex integrals and it either manages them (with step by step instructions) or is very close with small errors that can be rectified by following the steps manually.
There's a reason companies are setting up clusters of A100s, not MacBooks.
17.6 tokens/s on an M4 Max 40 core GPU
See also Hugging Face's MLX community:
https://huggingface.co/mlx-community
QwQ 32B is featured:
https://huggingface.co/collections/mlx-community/qwq-32b-pre...
If you want a traditional GUI, LM Studio beta 0.3.x is iterating on MLX: https://lmstudio.ai/beta-releases
Why pay Apple silly money for their ram when you could take that same money, get a MB Air and build a desktop with a 4090 in it (hell, if you already have the desktop you could buy TWO 4090s for that money). Then just set up the server to use remotely.
offline. anywhere.
separately, if you purchase anything from Dell to Framework, you quickly learn there's no "apple tax", there's a cost for a type of engineering to a given spec (for oranges to oranges "TCO" you should spec all the things, including screen, energy, and resale value), and it's comparable from any manufacture, with the Apple resale value dropping their TCO lower than the others.
I've been off Mac's for ten years since OSX started driving me crazy, but I've been strongly considering picking up the latest Mac Mini as a poor man's version of what you're talking about. For €1k you can get an M4 with 32GiB of unified ram, of an M4 pro with 64GiB for €2k which is a bit more affordable.
If you shucked the cheap ones into your rack you could have a very hefty little Beowulf cluster for the price of that MBP.
Given how unreasonable that is I thought this model did very well, especially compared to others that I've tried: https://github.com/simonw/pelican-bicycle?tab=readme-ov-file...