And Linux runs better than ever on them; I'm running debian 13 with almost no driver issues.
For $2k you can get 32 GB DDR5 RAM and 16 GB fast VRAM. Bump the RAM to 64 GB and you're still below $3k.
86 karma · joined March 19, 2026
And Linux runs better than ever on them; I'm running debian 13 with almost no driver issues.
For $2k you can get 32 GB DDR5 RAM and 16 GB fast VRAM. Bump the RAM to 64 GB and you're still below $3k.
The second server is 2x Radeon RX 7900 XTX (48 GB VRAM combined). It's a fairly recent gaming PC that's being repurposed. Idea is to power limit those cards too and run some overnight stuff w small/medium sized models.
Intel just released some 32 GB VRAM cards, but sounds like support across AI tooling is a bit rough atm.
During each task I context switch to some other work, emails, chores etc.
Important to take breaks and assess before starting a new session.
The analytics of thousands of accounts sending tokens to new accounts. Better use a VPN a migrate on an unusual hour in your time zone :D
There's SWE-bench Multilingual for example, but translating a problem into multiple natural languages before passing it to the LLM has not been benchmarked afaik.
If there's some residual of the natural language left when the middle layers execute, that would in part validate Sapir-Whorf.
Easy to way to start hacking LLMs; there's much of value there and a fun way to get into it before tackling heavy math / CS topics.
For browsers and general apps, devs have blown up memory usage like crazy the past two decades.. there's so much low hanging fruit in optimizing for reduced RAM usage.
Like many things it was cheaper to just use more memory, now it may become worth it to spend some time thinking really hard how to get your Electron message app using a few GB less..
Decided as a constraint to exclusively use local AI! This was fun in that the first step became assembling the first server able to run a small local model, that would then assist with everything else.
After I got the first one running it was used for almost everything, except it could not assemble the 42U steel server rack.. (shoulders hurt a bit now, probably good exercise!)
The first thing I tried on the new servers after first boot of debian was feeding the entire Linux dmesg log with one simple instruction: "Check all dmesg entries and provide recommendations for any errors, issues or other considerations".
This was very helpful even with smaller local models, as a complement to just searching for various errors (drivers etc). Learned a lot of new things like BMC network configs.
Home lab networking in general was incredible to work through using local AI. Being a bit rusty on various things like firewalls, local DNS etc it was refreshing asking questions so dumb that one might not want them in the logs of hosted AI providers given a history as a SWE...lol
And more complex things like how packets flow in mikrotik RouterOS.
Some general findings:
* The latest generation of local AI models are _way_ better than even just 6 months ago. In particular dense models 7B+ are surprisingly useful for anything Linux, network configs, small to medium sized scripts.
* Latest gen open models from small AI labs generally beat last gen models of the same size from larger labs.
* Don't trust recommendations for any specific model - try it for real stuff and get messy with it - feed it system/app logs, mad half-spelled ramblings late at night along with more clear and well written instructions the next day...
* Larger open models of decent quant (Q5 and up) are now so good enough that the bottleneck for many use cases is no longer the model, but your workflow.
* Simpler workflows beat complex prompts, skills, AGENT.md etc. I run most things with the pi-mono coding agent with no extensions.
* Have the same model verify a finding/claim in a fresh context. This drastically reduces false positives and improves correctness of findings. Going further, run a third verification with a different model.
* If you grew up with the sounds of floppy disks, 56k modems etc, you might just like the coil whine of local GPUs... it's oddly comforting and different models sound different when working on the same tasks.
Since many exploits consists of several vulnerabilities used in a chain, if a LLM finds one in the middle and it's fixed, that can change a zero day to something of more moderate severity?
E.g. someone finds a zero day that's using three vulns through different layers. The first and third are super hard to find, but the second is of moderate difficulty.
Automated checks by not even SOTA models could very well find the moderate difficulty vuln in the middle, breaking the chain.
Has some interesting github links.
Though beware that the increased score on math and EQ could lead to other areas scoring less well; would love to see how these models score on all open benchmarks.
Maybe in the future circuits will become modular and composable like models are today?