21,526 karma · joined August 19, 2009
My recent books can be read for free online on my web site or optionally you can pay for them at https://leanpub.com/u/markwatson
Twitter: mark_l_watson and Mastodon: @mark_watson@mastodon.social
I have had my 32G mac mini for 2 1/2 years and I have enjoyed watching one technology advance after another improve the quality of work I can do locally. I bet that what I will be able to do in one year on my old hardware will be even more awesome.
It is like a ping-pong game: the advantage flips back and forth between providers.
1. it would put us (USA) at a competitive disadvantage, and cooler heads will prevail in this fight
2. there are good US open models. I have the latest gemma4:27b with better tool support functioning at a high level in the pi coding harness. Thinking Machines seems to be on a good path, we will see what they and other US companies can do.
That said, and this is off topic: Siri on the newest iOS beta is surprisingly good now. I asked it what model it was using yesterday and it said Gemini for difficult problems, then secure Apple model in cloud, and local Apple model.
> Right now unless you're paying by the token, there's no cost based reason to use the open weight models for daily coding work because the monthly coding plans from Anthropic and OpenAI are a better deal.
Maybe. I am on a $20/month Anthropic subscription this month but I also use Claude Code frequently with Deepseek v4 flash and pro, GML5.2. For simple work Deepseek v4 flash is so nice because it is fast.
What you say is true however, the US hyper-scalers are still (desperately?) subsidizing subscriptions for market share to boost there valuations.
I really want to see AI inference costs approach zero, and I think I just need to wait a few years to see that.
EDIT: their business model is interesting, aiming for supporting organizations with privacy and security concerns.
I enjoyed turning off web search for Inkling to experiment with what innate knowledge is encoded in the model weights. A fun thing I do is check what innate knowledge very large models contain about me, as an individual. Inkling has an interesting concise shadow of what I do. (I have written a lot of books, so I am in training data.)
Now I write to learn new things myself and my mode of operation has changed radically: I still manually research and write code but I also ‘flesh out’ my writing and coding experiments using agentic tools. My approach is different because I don’t care at all about work efficiency so I spend a lot of time reading and studying material I generate with AI, and iterating in ways that wouldn’t make sense in a business environment. I still write open content books on whatever currently interests me but I am thinking, after 35 years of being an author, of winding down this activity because I think my readers might be better off doing their own research and coding, and with agentic design, research, and coding tools this gets easier.
BTW, I miss the old way of doing things, and some of what I do now is attempting to make AI less noticeable to me as I work, and using AI to work on my own work environment.
Everyone has different interests and abilities, and if what we might call the current AI hype cycle has any lasting value (open question) it is in letting us as humans ‘go our own way’ and customize tech and our life style to match what we want.
re: this article: excellent! I always do about 90% of my Swift dev on the command line and I learned new tricks from the article. Bookmarked.
I haven’t let coding harnesses run REPLs. When I do let a coding harness run tests I specify in, for examp,e, my Common Lisp skills file to run ‘sbcl -load …” so bash test commands are one liners.
It would be interesting to work on skills and a harness to use REPLs - nothing bad about that idea, I just haven’t tried it.
That said, Siri seems a little bit better now - my subjective opinion. It is a little bit less frustrating.
[1] https://open.substack.com/pub/marklwatson/p/running-opencode...
I accept that for you and your work this is true.
I have a different experience: for a month I paid big money for Opus and got a lot done. Now I am gorging on GLM 5.2 running on Fireworks.ai and I am also getting a lot done for about 15% of the money.
Everyone should do their own evals on their own work.
There are so many models, and I personally ignore benchmarks so it takes some time to try different models on my use cases. Fortunately, it is ‘good enough’ to do the work to find a few models that work for me, and just use them for a month or two before re-investing time for my own evals to possibly change models.
People should evaluate what works for them and ignore other people and benchmarks. (Apologies if that sounds snarky.)
All small-scale stuff. For large integrated projects I am finding DeepSeek v4 Pro commercial API to be very inexpensive and helps me produce good results.
I find that for local coding, I need to spend a lot of time building concise SKILLs for specific things I work on and try to only enable one or two skills per coding session.
To the author of the linked article nice job, and if you feel like adding to it, please add details on your setup.
That said, I expect much of the AI bubble to pop. Google Gemini with Antigravity is a good product, as is a Claude Code subscription but I have switched to using DeepSeek v4 Pro with the Claude Code harness and DeepSeek v4 Flash with the OpenCode harness (when I am not using local models with little-coder/pi) and at least for the foreseeable future I don’t think I am going back. Fast APIs at low cost trumps having to spend a little more time to get the same quality of results.
I encourage people to at least once a month to do a quick evaluation with their own problems and workflows. Estimate cost as both what inference tokens cost for a task and also how much human effort it takes to get required results.
I disregard benchmarks.
One fix for this problem: Allow US companies to buy memory chips from China. I saw an article about a month ago, that if my memory is correct in this, said that China is ramping up high-end memory manufacturing.
Fix number two: my country (USA) should cease and desist with the craziness that is data center buildouts for AI.
Clearly ‘BIG MONEY’ always needs a new thing (cloud -> crypto -> AI) and the powerful get what they want.
If the US Congress acted to benefit regular people rather than special interests (both party's are corrupt, disbelieve that if you want to live in a fantasy land) then anti-dumping laws would be passed.
If all companies and individuals paid the real price for tokens, then we collectively would work more efficiently. As is, the filthy rich get even filthier, and regular people will get screwed.