Using an open model feels surprisingly good
matthewsaltz.com
matthewsaltz.com
The smaller/open models are less good at that. But that’s not how software development is done. You don’t prompt a whole app and be done with it. If you’re using it as an aid to traditional software dev, iterating on small, targeted functions, GLM works as good if not better than Claude. Anthropic expects low quality prompts. If you rubber duck GLM, you get absolutely pristine output in most cases.
I think the issue is that no one is content with incremental progress. We all know one shots are mostly possible, so the age of the personal project is kind of over. There’s no motive to invest dozens of hours getting a working prototype when Claude can give you something right now. So you can’t make small changes until you have that codebase in place already. You’re forced to make sweeping changes if you use AI from the beginning. And it’s not like it matters, there’s no personal attachment to any one part of the code, it’s not even seen!
IMO because they speed-ran it so much (raising orders of magnitudes more, pushing out new stuff orders of magnitudes faster) it will also lead to the rest of the world catching up faster and shrinking the pool of people internal to them who continually benefit compared to what actually went on with IBM. There's only so much upside when you join a company with hundreds-of-billions valuation.
I guess the other part of their business plan bet was "become literal god of AGI". Maybe it'll still happen!
That only lasts until the AGI decides their god would be more useful if turned into paperclips.
Going open weights is an easy way to gain mindshare in this sort of environment, which is exactly why Meta also went open weight. The difference is that in China you had leading models going open weight which puts a lot of downward pressure on other orgs to do the exact same. I doubt China's grand vision extends beyond achieving the best system possible, and fostering a highly competitive environment is exactly how you achieve that.
The fact that open weight frontier models may also cause the investment bubble in the US to burst is likely incidental. AI funding is a bubble, and it's going to burst. The exact final cause is again mostly just incidental. The economic damage this will cause in the US will also cause substantial downstream damage to China as well, and they generally aren't so big on the whole punch yourself in the face because you don't like another country, that has become trendy in the West. So the idea of a secret economic attack doesn't even make much sense. In the status quo China wins, so they don't even have any motivation to cause chaos.
I think the people in the Chinese labs have a different set of motivations, beliefs, and values that make them more amenable to sharing AI as a common good rather than some economical advantage. I've heard that workers in the Chinese labs share far more knowledge between themselves, and the narrative around AI in China is different from what we see in the USA. So it is possible, I believe, that the labs are to some extent releasing open models "for the good of humanity", or at least they believe it is the right thing to do.
Parallels to the US auto industry not paying attention to its buyers and the right strategies when they lost technical dominance to other countries a few decades back.
Claude feels more like gambling, that's what it is. Built for the vibe coder.
It doesn’t have any subscription or limits. You just load up api credits.
I honestly have a better experience using DeepSeek, in part because since the cost is lower I feel more free to experiment with it. But in part because the model is excellent, and can handle everything I throw at it in a collaborative manner.
I am really excited for when V4 is out of preview and they are done with all the post-training bit.
But the situations where you need that are narrowing every release. The "Composer 1" era of low-end models is pretty far away now.
What kind of fields are we talking about here? Compsci? EEng? BioChem?
This has lowered the barrier of entry for many non-SWEs. I'm primarily a data scientist myself in the environmental sector with limited front end development.
My partner proposed an idea to help manage her horses and over the course of several weeks we fleshed out an android app that would enable/assist her with horse care.
I spent a few days in planning mode pointing claude to the services we want to utilise and the functionality.
We now have a fully developed app that we are testing but found several features missing. For example, Claude implement X but without a level of verbosity it failed to implement the edit/deletion of X.
The app is highly tailored to her use case and I had already built the main parts as a POC but she had no user-friendly method to interact with the information she required. Claude made this possible in such a quick time frame that just seems insane. As someone time poor (as most horse people are), it would have taken me a year+ to produce something unpolished when compared to what Claude produced.
To me it is great that LLMs are allowing more people to use computers and software the way they were meant to. But the people that are using them this way, in my opinion, wouldn't have commissioned anybody to do it anyway. They'd just live with whatever process/pain they have. So jumping from "I can now produce a prototype in a weekend" to "software development is dead" has always felt strange to me. Just something I've been thinking about lately.
I understand what you're saying. My opinion is I see agentic workflow similar to the star trek universe where they ask the computer questions and get a response while they continue with their work. But that doesn't mean not learning how to do things from first principles. This is what I tell juniors in my field when I pass jobs to them. They can use LLM but to ensure they know and understand what's going on and most do from their university degree. I think this is where we need to pivot towards when discussing LLM.
Whether you can't describe the analog circuits that correspond to the opcodes of the major mobile architectures running the platform of the app you're designing a UI for, or you couldn't have built any of the services you're hosting on from scratch, no one is going to be disappointed.
There are more people making more money off whatever we're currently calling making computers do things than ever before in history. The rest is largely business as usual.
Not in a deep deep level like "this is the optimal way to manage low level memory access in an Android application" but on a higher level of "It's possible to do a custom Android app in a week that's of decent quality for a small userbase".
Every piece of software doesn't need to be targeted for billion user hockey stick growth and an IPO as a target. It's perfectly fine to build something that's just right for 10 or 20 people. Or just for your immediate family.
For example I have a quiz tool I built with LLM assistance for Christmas that mine and my siblings kids use every December, they get clues appropriate for their level and rewards when they solve the daily task.
It's fine, has custom handling for every user etc. But also something I'd never bothered to build by hand from scratch nor something I'd pay someone to build for me. It was a fun "hmm, I wonder if..." project I started a few years ago on a late evening in November.
Like a quote I've read on here multiple times about replacing e.g. excel or whatever, people often mention that users only use 5% of a software's functionality, but it's a different 5% for each person. The argument being that to replace excel you'd then need to mimic every esoteric thing it does, but when users can now achieve that 5% in a highly customized way with no real knowledge needed, it's going to have a major impact on big software.
The people who were twisting Excel+VBA into massive "applications" in-house should be getting Claude licenses and guidance how to do that with a custom application instead.
We've done that internally already with a few very successful use-cases and a bunch of "this is nice" -level things. I think one case saved us 4 figures a month when we built a bespoke service that does just the things we need and could replace a SaaS that had 420 features - of which we used 2.
The "pay someone to build it" is definitely the one I wouldn't touch, or even consider.
If it's a good idea, they'll already have the code and will run away and make their own business with it. If it's a bad idea, I've wasted a bunch of money.
And in both cases I'd need to spend an inordinate amount of time explaining the terminology and use-cases to the contractor, book meetings with them during office hours to talk about requirements and get demos etc.
With an LLM I can be walking the dog late at night, get an idea, and within 2 minutes I'll have Claude working on it via my phone.
Then I come back home and maybe test whether it worked while watching TV.
If I had fuck you money, I might have a contractor on call that would agree to workflows like that, but I'm just a middle class software dev so I don't :D
This is my experience with using AI to generate tailored apps too. The first 80% makes me really excited but then it either only kind of works, has bugs, or has missing features. Even when I do eventually get it working I feel dirty using it because I just know it's badly implemented. Of the dozens I've created I don't think I still use any of them. I have a few glue scripts still in use.
So not fully developed then.
After open models will also be able to do that, I expect everyone to forget "how software development is done".
I see a lot of people saying codex is the best cli agent harness. But I haven't used either so can't compare.
If you have a pair of 3090s and run Qwen 27B, or an old Threadripper with heaps of system RAM and Deepseek or MiniMax or Kimi, no it won't be as fast as Claude.
Most local LLM nerds are not running locally for superior speed, we're doing it for sovereignty and/or privacy, or maybe just because it's fun which accidentally became useful this year.
I was stunned that Claude simply started spinning its wheels in an attempt to actually build something.
What would be useful is a cost metric. I'm curious how much I'd be willing to spend as a premium to not have those companies piping my conversations directly to the NSA. Maybe only some conversations? Claude and OpenAI are heavily subsidized, by all accounts, so Kimi K3 on a private endpoint might end up costing more or less - that's what I want to know.
Just the GDPR/CCPA fines alone WILL bankrupt both companies.
The TTFT and tok/s are much higher on the small model so that makes it competitive for a bunch of things. It feels like what old Sonnet used to by the end of last year which is honestly damned good.
I can do things like take a photo of a doctor's note among add the appointment to my calendar, send a PDF to my archive tagged, OCR'd etc, search info from internet, look data from Google Maps.
I can even integrate this to home assistant and talk to my agent with my open source Alexa-like system.
It was truly remarkably easy to build and package it exactly to my whims, in my case as a single docker container with a process reaper that runs llama, my Go code, tts, chat harness, browser in xvfb, and even a mailer daemon. With Gemma, it even runs on a RPi 5.
What’s absolutely wild to me is that over a couple hours, I could probably have it import parts of Home Assistant for my devices directly, and do other wacky stuff in what is essentially software for one.
Paid 20 euros in total.
In my case, I was foolish enough to run it all on my hardware which is pretty damned fast but has a duty cycle of 5% and runs idle most of the time. It's definitely better done via API.
Tell me more about the Alexa-like system! I have mine at home on a custom OpenWakeWord model trained on the word 'Aurora'. And I, too, have it hooked up to Home Assistant. I have a bunch of Eufy E21 baby cameras mounted on the walls[1] so it has vision through the house. The vision model uses GPT-5.5 on the subscription because I haven't yet set up a Qwen or Gemma multimodal that can see.
With Frigate on my home server I can even watch for events like my daughter waking up! And at night the agent sends out the vacuum if we've cleared the floor of baby toys. I feel this close to the dream of sci-fi AI. Because I auto-forward a bunch of my email to it etc. it knows about what's going on with me.
My wife will sometimes ask the agent information about me etc. and it's way easier to get a fast response rather than waiting for me to see the message etc.[2]
0: https://news.ycombinator.com/item?id=47538158
1: https://wiki.roshangeorge.dev/w/Blog/2025-12-01/Grounding_Yo...
I can't wait until good enough models can run on affordable hardware.
EDIT: Oh, home assistant. I use Home Assistant https://www.home-assistant.io/
https://www.home-assistant.io/voice-pe/
Then you can set the background AI to be any OpenAI compatible API. So I just created one to my local Rust Agent, and connected it to home assistant. Now I can yell from the couch to create me a new Proxmox container with the next free static IP address etc. :D
It’s funny how we converged on a similar system. My agent is also written in Rust.
What a time to be alive.
I know the author has a long history here and I'm sure the post is a genuine reflection. But how is posting "I used my own product and it felt really good" anything other than self promotion?
It doesn’t mean the post doesn’t accurately reflect the author’s views - it probably does. And the conflict of interest is well disclosed in the brief post. In fact, it’s quite natural for somebody who likes open models to work somewhere associated with them.
So personally I don’t find any impropriety going on here - but yes, it’s an ad.
We have seen self-criticism and self back-pats here. There's no need to be so critical. Others can test and have their own opinions, too.
If someone explains really well why they find their own product great, that's different. I don't mind that at all.
I see this as a post from someone who understood the value of "open". Now they're open to understand the value of "free" (as in unencumbered). In the age of AI, we need to understand the value and embrace more open and free, both as in models and software.
It's good food. It's worth thinking about.
Advertisement is not sharing, it's manipulation.
> ...personal account. I work at Modal, and today we just launched Kimi K3 on managed endpoints, and I know Kimi K3 is supposed to be pretty solid, so instead of upgrading my Claude plan, I wanted to give it a try. (I didn't directly contribute to this feature, so I haven't gotten to play with it yet.)
Emphasis mine.
We can't trust for him to give honest opinion. The information value is near zero.
I read about owning the data, knowing the path, being able to control what you have and the lightness it brings. That feeling came through his employer, which is something we can debate all year long, but I see nothing about advertising Kimi-K3 or managed endpoint or whatnot.
Maybe because I'm in all this for a long time, and having an AI capable server near is not something alien to me.
Anyway, in short, there's no AI related content for me in that post. I'm more on the experience side.
Now I can't wait for someone to distill K3 into a Qwen 3.6 27b or Poolside S 2.1 sized models for a proper fast local Composer 2.5 replacement.
For example, if you modify things at the level of small functions, open models seem to perform just as wel
Harness-aside, I get this feeling sometimes when I swap from a big frontier model to something more nimble like Composer.
I can get into a better thinking and q&a loop with fast models, similar to how I can flow through a file more easily with vim.
Making it about the hardware costs is not the play, they’re not an actual barrier.
but why?
But if you just wanna outsource understanding for 3 bucks in tokens, that’s cool too.
I'm not coding with Claude cause I enjoy the inherent novelty of LLMs. I'm doing it cause it's enabling me to quickly solve problems without getting stuck on the hitches that have deterred me from bothering with dozens of side projects my entire life.
(FWIW, I have the same attitude/preference as you do. Until I can run SOTA models on my own hardware without taking out a second mortgage, I'll pay our AI overlords for the privilege.)
If it sounds good it must be good. The code this new model writes is incredible! It told me so!
Really just feels like endless FOMO with how fast the iteration cycle is for harness and model development.
I don't think it's very good though. As an example I can use one CC instance to delegate to several to achieve complicated/open-ended goals.
That just isn't possible with other harnesses, and it's definitely not benchmarked.
With that said I think Codex/Kimi code are all behind CC as well, so maybe it's a question of effort and not the model.
I've tried using local models, but they run slow on MPB M5 (base) 16GB, and I am not upgrading my laptop any time soon.
Edit: I see, had typo last word of sentence was meant to be “blog”
Also, if you don't like a submission, flag it and move on. Commenting that you don't like it doesn't add anything to the discussion.
People need to learn to just post the prompt rather than the LLM output which just fluffs the prompt.