A tiny, smaller than 1b parameter model, fine-tuned, can kick ass for constrained work. I do not have a lot of budget, I fine-tune only on a 16GB M4 Mac Mini. But that also tells me the potential is wild. Progress has been slow since I moonlight on this.
I have been trying to build a set of models + agents for full-stack development, where each model does only a small piece, like take user prompt and break into backend/frontend tasks. Then a Rust+Diesel model, a Rust+Auxum model, a Solid+Router model and so on. I know this is wild but this is just theory - can 5 or 6 Qwen 3.5 0.8b models do full-stack web development? My hunch says they can, better than what most people expect. Heck, with a good harness, it might beat all the cheaper models for the specific task, like Haiku or Luna.
That being said it’s much easier at the moment to continue to use the frontier providers for most general tasks, that is the argument I’ve heard.
For creating these types of fine tuned local models, on constrained hardware for inference, I do think this is the way to go for specific tasks too!
Have you actually fine-tuned yourself? Email categorization comes to mind and there are tons of non-LLM approaches even that will give fantastic results. How did spam filters work before LLM?
I think LLMs just made us think that is the only way. It is not.
Or are the subagents generating your training data using a closed/paid model?
For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.
The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.
Think of it as distillation, but focused on a specific task.
I think OP's point remains, if you generate 140k pairs, your local model would need to run that many to offset having just used the generator (SOTA or not) model to begin with.
I wonder if another approach if latency is a concern is just to do a two shot pass with Jev (perhaps given small context you'd want one to match command, then one to match args of given command) would be an extremely fast, and cheap way to do it - rather than training your own.
Not quite. One, because you save on the initial query being sent multiple times. Two, because the reasoning will be very similar at the beginning; "user asked me to", "let's check what's in this project already", etc. you'll get similar actual output cost, but input and reasoning will be shorter.
This particular example is maybe a niche, but 1400 people can use a few hundred queries in a reasonable amount of time.
Which 5$ is a pretty easy sell if its useful in any way, It's pretty easy to justify a purchase if its yours forever and doesn't use much CPU so is easy to run I mean people were spending 1000$+ on mac mini setups to run local llms or run remote agents.
They also used Astra for the coding, which they can't get on a $10 subscription. And then there's the actual training cost.
I ask cause would this be a kind of model distillation?
I have a small model I'm looking to train on some data, and I have some real live data but I'd love to be able to extend it.
edit: others have asked any you have replied "soon (tm)", looking forward for that day
I have a single command that fires up llama.cpp on cpu only using gemma4 e2b, answers a single question from the command line and exits. This takes about 3 seconds to load from an SSD, and is smart enough to solve exactly these "remind me of the syntax" scenarios if you dont wanna switch to a browser.
When I do something "heavy", my 9800x3D will kick on the fans and start making lot of heat. This is fine when I'm intentionally doing say a transcode, but if I'm just querying syntax it could get pretty annoying. Do you "feel" it? I know that will be pretty machine dependent.
Not my project
I am building a natural language to CSV/Excel commands for a "wrangler" type desktop app. Same issues. The MVP is being built with parsers of sorts, entirely code generated. Then I want to fine-tune a tiny model at some point.
I’m really pushing my recall, but I want to say it was written in Ruby and stored pre-configured commands that it just did traditional search over.
I vaguely recall it working okay because 99.99% of the questions people asked were the same (“tar command to gzip a directory and strip the prefix” is something I google like once a month).
https://huggingface.co/vwdubb/Qwen3.8-27B-Fable-Distill-NVFP...
side quest, are fable distillations only wrong when it's another country?
we dont know what the result is and how its impressive.
Aside on the aside, I welcome this new era of really personal software. Not Ai's being sycophants, rather being able to easily and quickly change, adapt, or extend software I am not familiar with.
I had to intervene a few times. For instance, as smart as the models are said to be (Astra), it would copy the full training run, train on the server, pull every checkpoint to the local machine, then run tests, update. So, the bandwidth bill was as high as training bill for the first 6 hours. It could have simply tested each checkpoint on the server, saved time and money, didn't occur to it until I said.
I wouldn't put my house on it. Brave.
This is the part where the narrator looks at the camera and says "Don't try this at home, kids!"
Are you confusing this with an OAuth token or something?
https://docs.cloud.google.com/billing/docs/how-to/budgets-sp...
[Search: Can I refund Google cloud?]
It looks like we’re not able to ask for a refund since we did actually use all of that compute intentionally.
Would you like me to write you a pleading email to send to the support team?
Can you (or your agent) please write a tutorial or share some good links.
(Namely the fine tuning part.)
I'd like to learn how to do this as well.
Edit: will do as soon as possible
It's good to hear you're enjoying yourself, but I suggest retiring that expression. It's really beginning to grate.