I might be wrong about this, but obviously Google would like to provide inferencing at the lowest cost to themselves, so perhaps their slow ‘pro’ releases and rapid ‘flash’ releases is an attempt to guide people to use more profitable models?
21,520 karma · joined August 19, 2009
My recent books can be read for free online on my web site or optionally you can pay for them at https://leanpub.com/u/markwatson
Twitter: mark_l_watson and Mastodon: @mark_watson@mastodon.social
I might be wrong about this, but obviously Google would like to provide inferencing at the lowest cost to themselves, so perhaps their slow ‘pro’ releases and rapid ‘flash’ releases is an attempt to guide people to use more profitable models?
Venders coupling coding harnesses with their own models is usually a good thing. Poolside.ai has a combined harness with their own models that works well locally, and the DeepSeek harness with their models is very interesting.
Ollama Cloud offers the same thing: they supply a web search tool bundled with cloud API inference services.
I have spent two months experimenting with a wide range of US and Chinese models, and I had a lot of fun doing that, but I am in the process of switching to just using local models, using Gemini on an API if I need it, and once or twice a month when I really need help on something difficult, I use something top-tier like Kimi K3.
I like to start by prompting with “examine this code base for problems and improvements and write to IMPROVEMENTS.MD” and then carefully look over the suggestions, and either fix myself or let the model+coding harness try.
I am a huge enthusiast of running local models, but when multiple quality USA vendors provide models like GLM 5.3-flash, I run locally just for the fun of it.
For the purposes of comparing to Fable 5.1, I would mention GLM 5.3 that is about 1/12 the cost.
Does Fable 5.1 really provide much benefit over models like Kimi K3 that are 1/3 the cost? Or GLM-3 that are 1/12 the cost?
If you can talk about your work, what kind of tasks do you work on where the higher cost is very much worth it?
Ed Zitron mostly covers the costs of data centers, circular spending, and predictions of large the market for AI has to be to justify the data center expenditures.
I was disappointed this article didn’t really cover Zitron’s main arguments.
Maybe a paid article placement? I don’t know, but I was dissapointed: I read Zitron’s material and I wanted to see good counter arguments to his rants about costs of data centers, circular spending, and predictions of large the market for AI has to be to justify the data center expenditures arguments.
I have always been a Lisp devotee, but a few years ago when I started using uv, I then started seeing Python as a language I could really enjoy using so I put effort into making my Python dev setup nearly frictionless.
re: Ridley Scott making another Alien franchise movie: if he is doing it because he really wants to do the movie then that is great; if he is doing it for business reasons to make money, that is not so great
Not to go off topic but I am pleased to see open model support from US companies like Poolside.ai, NVIDIA, IBM, Google, etc.
The financial aspects don’t work however: I can learn and experiment with what I have for local models, and I pay as I go on FireWorks.ai for open model inferencing and no matter how much I use this service my monthly bill is between $10 and $40 and much faster than any reasonable home rig.
Hybrid ‘small local’ and buying inference is the way I choose.
re: data centers: pump and dump. Wealthy investors will have made their money and walked away, and the corrupt democrat and republican politicians in Washington will, as usual, protect the interests of the ultra wealthy and leave the general public to pay for poor decisions. There will be a government bailout.
Anyway, on a positive note, I am all in for small local models that are augmented by strong hosted models for specific tasks. Use technology to help people, not make billionaires even more money.
Now I might try Ornith-1.5-35B first even though I don’t see an official MLX version.
I usually use small local models, and the work to set up very concise skills and efficient tooling is a big part of the fun. I have also adopted the practice of writing my own custom coding harnesses (these can be less than 2000 lines of code, not the huge project you might expect.)
A weird thing: whichever Lisp language I am currently using for work (or a side project) is my favorite.