21,534 karma · joined August 19, 2009
My recent books can be read for free online on my web site or optionally you can pay for them at https://leanpub.com/u/markwatson
Twitter: mark_l_watson and Mastodon: @mark_watson@mastodon.social
I can recommend my own layered approach, using the lowest capability models that get stuff done:
1. I maximally use local models like gemma4:26b-a4b-it-qat for everything that works with this free option.
2. I like paying for inexpensive APIs for mid-tier models like deepseek v4 flash, gcp-5-mini, gemini-2-flash for things that option 1. fails at. This option is almost free.
3. Pay for more expensive APIs like deepseek v4 pro, gemini 3.5 flash, etc. This option is not too expensive.
4. If all else fails on a class of tasks, then pay for awesomeness of Claude Opus. $$ expensive, I try not to use unless absolutely necessary.
I think developers and companies that just cram everything into Claude Opus are unprofessional.
I also pay for Proton's Lumo+ private chat and for what it is it is also good.
The free plans from all the providers are bad, which is fare enough.
I use Apple devices and I expect to be paying for Gemini tokens after the integration.
Do you think Google doesn't protect privacy for large paying customers?
For years I have enjoyed using Google products that I pay for, and they are clear about privacy guarantees.
Slow is Fast.
I had horrible luck with Gemma 4 12B with a variety of coding harnesses - but as usual Qwen 3.5 9B did OK.
EDIT: CORRECTION: I pulled a fresh copy of Gemma 4 12B and inference code and the tool use problems in my test harnesses are fixed. Gemma 4 12B is slow on my 16B MacBook Air, put produces OK results.
I do care about: how useful their products are vs. cost and how secure are their businesses. Actually I only care about the first thing since these services are hot swap-able with some effort.
Then I saw this in the article:
>> I've discovered that Claude Desktop supports MCPB. The MCPB provides a Node runtime. Therefore, our application would only contain our code. This means the size of the application would be ~1MB.
I don't know why this is so appealing to me but it is. I currently use Claude Code with a DeepSeek v4 Pro API backend. I will check out if Claude Code itself has MCPB support.
It is easy for me to change providers. Right now I use the open source Claud Code harness with two paid API venders for DeepSeek v4 (flash and Pro). I like seeing how much each session costs.
BTW, Google is my pick for the winner in the USA tech giants AI race. I worked at Google about 12 years ago and was impressed by their use of renewable energy, etc.
I spent a month comparing Gemini Ultra plan to using much lower cost DeepSeek v4 with open source coding harnesses and, spoiler alert: I was happier using the much cheaper and more environmentally friendly open models: https://marklwatson.substack.com/p/my-evaluation-of-ai-agent...
I have read a few references that humming or ‘ohming’ help sinus health and breathing so I guess it makes sense playing the didgeridoo would help also. Blowing bubbles through a straw won’t cause vibration, so probably in itself won’t help.
Breaking code up into composable chunks has worked well for me over 50+ years as a professional software developer, and I can't get away from the idea that it is still usually the way to go using agentic coding tools.
That is the question. I love using OpenCode with paid inference providers and seeing the cost of every little thing I do. On the other hand, right now I am flipping between Antigravity CLI and the two Antigravity apps burning Claude Opus tokens like crazy, knocking off a ton of work. Google must be losing money on me.
My 1 month subscription of Gemini Ultra is finished in three days and I revert to their $20/month plan. Assuming that daily and weekly quotas are OK for casual use, I will probably use AntiGravity CLI most of the time.
Off topic, but maybe interesting: during my one month test of Gemini Ultra I did several tasks in parallel to compare (old) Antigravity+Claude Opus vs. OpenCode with a fast provider for deepseek-v4-pro, kimi-k2p6, and minimax-m2p7. In almost all cases I could get stuff done in about 60% to 70% of the time using Antivravity+Claude Opus -- but!, OpenCode with the open models is so much cheaper. I get that in a work environment when someone else is paying for tokens, why not burn someone else's money. Paradoxically, I felt more relaxed after the tests with OpenCode with the open models even though I was actively doing more work myself.
EDIT: two months ago I wrote my own small coding agent in Emacs Lisp that I enjoy using. I am researching redoing my Emacs project using the new Antigravity SDK.
EDIT: I run with context wired at 64K
I have worked with old fashioned neural networks, deep learning, and now LLM-specific deep learning: wonderful technology, but over hyped, and advice to go a little slowly, with firm use cases that are financially viable is great advice!