Are there any tutorials you can recommend for somebody interested in getting something running locally?
Are there any tutorials you can recommend for somebody interested in getting something running locally?
You can, almost, convert the number of nodes to gb of memory needed. For example, Deepseek-r1:7b needs about 7gb of memory to run locally.
Context window matters, the more context you need, the more memory you'll need.
If you are looking for AI devices at $2500, you'll probably want something like this [1]. A unified memory architecture (which will mean LPDDR5) will give you the most memory for the least amount of money to play with AI models.
[1] https://frame.work/products/desktop-diy-amd-aimax300/configu...
When local models don’t cut it, I like Gemini 2.5 flash/pro and gemini-cli.
There are a lot of good options for commercial APIs and for running local models. I suggest choosing a good local and a good commercial API, and spend more time building things than frequently trying to evaluate all the options.
It's been a while since I checked out Mini prices. Today, $2400 buys an M4 Pro with all the cores, 64GB RAM, and 1TB storage. That's pleasantly surprising...
there are no out the box solutions to run a fleet of models simultaneously or containerized either
so the closed source solutions in the cloud are light years ahead and its been this way for 15 months now, no signs of stopping
would probably need multiple models running in distinct containers, with another process coordinating them
Opensource solution to run fleet of models in containers