Also if your thing doesn't work with `pi` out of the box, then low effort
458 karma · joined May 3, 2017
Also if your thing doesn't work with `pi` out of the box, then low effort
it’s more akin to 3d printing to me, i get the design all setup and let the machine do 3) and then get to play with it in 4)
these LLMs are great are generating arguments but they don't ask questions, we will need mathematicians to shepherd them into more discoveries
i really want to see open weight models crack some breakthroughs
i think as the novelty wears off on the average person, it’ll probably be seen akin to stick figure drawings in the future
we're literally looking at insane margins over compute, as energy gets cheaper, margins get wider - china focusing on cheap solar is probably going to be a key reason why their AI is so much cheaper
Sonnet 5 today was incredibly slow for example
it's my theory at least, the Hindenburg Research of AI
can i build a mini pc myself? probably but meh
it happens to all models…when the internet is increasingly generated, things happen
I'm on a 48gb M5 Pro right now and it's been okay, a lot of my rough experiences have been with MLX and I'm finding that GGUFs are okay now
local models do involve some context engineering to get it okay, but it's not that rough
you get a macbook for work, you run the macbook
they're not going to start giving GPUs to employees to run local models
You have dense models (qwen 27b, gemma 31b) who are pretty smart, but pretty slow
You have MoE models (gemma 26b, qwen 35b, north mini code 30b) who are pretty fast, but make a lot of mistakes
You need a lot of memory to run these well, quantization makes tool calling weaker, so most run at 4 bit quants and are wondering why it kinda sucks and that's because you've essentially lobotomized the model (I recommend unsloth quants, i recommend 6bit for MoEs and 5bit for dense)
So you need a lot of compute to make the pre-fill fast, you need bandwidth to make the decode fast, you need a lot of memory to hold everything - lot of ifs
On top of that, your laptop becomes a loud hot churning machine, it's uncomfortable to work with.
So are they good? not really. Do they work? yes
edit: just wanna clarify - i think open models are the future, i think they're super important, i'm contributing constantly to the ecosystem - i think people should play around with these models, i think people should use `pi` and learn how it all works - but don't download a model expecting it to be good out of the box, you will have to tune and configure a lot of stuff to replace a "coding agent" that most people are using models for
China may not care about open source, but they know they will personally fund AI through government investments while US relies on private investments, best way to scare private investments is a free capable alternative for everyone
Add on the fact that they actually invested in energy infrastructure and can offer AI very cheap to their citizens and you can get a population well versed in AI to reduce menial tasks and focus on more productive things (if we're to believe the claims of the technology)
omlx + gemma 12b 6 bit + pi
it’s feasible for sure
MoEs for speed (qwen 35b, cohere 30b, gemma 26b)
Dense for more methodical work (qwen 27b [reigning champ], gemma 31b, gemma 12b)
MoE i recommend 5bit+
Dense i think 4 bit is okay
Play with your context size, you don’t really need that much, have lazy loading for tools and mcps
my pi extensions for anyone looking for a skinny quick setup, i have use `--no-skills` right now too:
"npm:pi-codex-goal",
"npm:pi-simplify",
"npm:pi-mcp-adapter",
"git:github.com/elpapi42/pi-minimal-subagent",
"npm:@wierdbytes/pi-statusline",
"npm:@aliou/pi-guardrails",
"npm:pi-lens",
"npm:@juicesharp/rpiv-todo",
"npm:pi-hashline-readmap",
"npm:@mrclrchtr/supi-review",
"npm:pi-cmux",
"npm:@mrclrchtr/supi-context",
"npm:pi-tool-search"
think of local models as "zero sugar" models and that's where we're at right now. I think it's crazy how good these models are compared to last year's frontier modelsbut overtime if you adjust your verification rubric, it’s not too bad, gets pretty good, if you do make it do TDD, it gets kinda crazy and you’ll have 2000-3000 tests after awhile, or on my common case, 6000-7000 lines of code in single files (i usually have a cron to audit files for decomposition and create tickets)
i wouldn’t use it at my job yet, but it’s been fun to use for personal projects - it’s like modded minecraft automation or factorio
but AI dev workflows get complicated fast
you start with claude code or codex and it's cute, but then you realize - hmm configuration is cheap, the AI can do it!
then you start looking into MCPs and skills, fuck it, oh-my-pi looks awesome!
wait a second? I can just have AI make my own personal AI harness! Next thing you know, you're writing the 5th version of "little-coder" or similar using the Pi library
ahh shit, you just read an article that `tools` are actually crazy important for AIs, using `sed` is dumb when `hashline` + ASTs are way better, lets just start writing our own tools!!
...anyway I just use Zed, simple agent on the left, code on the right
i have some pretty complicated automated workflows that use `linear` + a orchestrator -> implementer -> reviewer -> releaser workflow, but it's less a dev stack and an AI factory