HNHacker News
TopNewBestAskShowJobs

mark_l_watson

21,526 karma · joined August 19, 2009

I am an author of 20 books and a practitioner specializing in artificial intelligence, deep learning, natural language processing, and the semantic web. I have 55 US patents. I code in Common Lisp, Clojure, Swift, Python, Haskell, Java, and Scheme. My web site is https://markwatson.com

My recent books can be read for free online on my web site or optionally you can pay for them at https://leanpub.com/u/markwatson

Twitter: mark_l_watson and Mastodon: @mark_watson@mastodon.social

submissionscomments
mark_l_watson··on Qwen 3.8
I like that v4 flash is so fast! I run it on both FireWorks.ai in the US and bought some tokens directly from DeepSeek as an experiment. I only work on Open Source projects, so I don’t have to worry about my work being used to train models - I welcome AI’s being trained on my open content books and code (but not my conventionally published books: I am a party to the copyright suit against Anthropic).
mark_l_watson··on Qwen 3.8
There is a ton of headroom (or room for improvement) in smaller locally runnable models. Some of the Gemma 4 models were re-released this week with better tool support and the improvement in using it with pi for a local coding harness is very noticeable.

I have had my 32G mac mini for 2 1/2 years and I have enjoyed watching one technology advance after another improve the quality of work I can do locally. I bet that what I will be able to do in one year on my old hardware will be even more awesome.

mark_l_watson··on Qwen 3.8
I found qwen3.6:26b slightly better on my 32G mac mini than the same sized gemma until gemma was updated with better tool support 4 or 5 days ago.

It is like a ping-pong game: the advantage flips back and forth between providers.

mark_l_watson··on The state of open source AI
You may be correct but the combination of local models when they are fast enough and work, combined with paying close to zero for deepseek v4 flash from US providers is pretty good. When you need it, glm5.2 is cheap to use and very good for working with larger projects.
mark_l_watson··on The state of open source AI
I worry about that but here are two positive things:

1. it would put us (USA) at a competitive disadvantage, and cooler heads will prevail in this fight

2. there are good US open models. I have the latest gemma4:27b with better tool support functioning at a high level in the pi coding harness. Thinking Machines seems to be on a good path, we will see what they and other US companies can do.

mark_l_watson··on UIUC AI Teaching Assistant
Very true. The code is interesting and educational since the overall design is simple. Also true that the PDF to vector database code reminded me of something I would have done a couple of years ago. (I almost always now combine good old fashioned BM25 search with using chunk matching with embedding vectors.)
mark_l_watson··on UIUC AI Teaching Assistant
Interesting looking project, but I just looked at the repo. UIUC is such a great school. I managed an AI team in the UIUC Research Park for two years and interactions with professors and hiring a lot of student interns was a good experience.
mark_l_watson··on LM Studio Bionic: the AI agent for open models
Apple’s System Model is actually pretty good but is hard-limited to a 4K context length. This is OK for small Python utilities that need a model for applications that operate on small amounts of data, but is a disappointing limitation.

That said, and this is off topic: Siri on the newest iOS beta is surprisingly good now. I asked it what model it was using yesterday and it said Gemini for difficult problems, then secure Apple model in cloud, and local Apple model.

mark_l_watson··on Kimi K3: Open Frontier Intelligence
re:

> Right now unless you're paying by the token, there's no cost based reason to use the open weight models for daily coding work because the monthly coding plans from Anthropic and OpenAI are a better deal.

Maybe. I am on a $20/month Anthropic subscription this month but I also use Claude Code frequently with Deepseek v4 flash and pro, GML5.2. For simple work Deepseek v4 flash is so nice because it is fast.

What you say is true however, the US hyper-scalers are still (desperately?) subsidizing subscriptions for market share to boost there valuations.

I really want to see AI inference costs approach zero, and I think I just need to wait a few years to see that.

mark_l_watson··on Inkling: Our Open-Weights Model
That is a nice business niche, but I would hope that companies might also care about security. There are just a few inference providers who I trust to not save data outside the scope of maintaining inference cache, but I do use untrusted providers sine everything I work on is open source or open content.
mark_l_watson··on Inkling: Our Open-Weights Model
I was blown away by how effective their latest model drop that works with their own coding harness pool is. I tested it extensively with Haskell, Python, and TypeScript for small coding projects. The functionality is good but the inference speed is too slow for most of my work: I would set up a problem, take a walk, then return later to evaluate the results. Note that I have an old mac mini with 32B memory; a fast modern home system would be much better.

EDIT: their business model is interesting, aiming for supporting organizations with privacy and security concerns.

mark_l_watson··on Inkling: Our Open-Weights Model
+1 I enthusiastically use Chinese open weight models for a wide range of tasks (I also love Opus and Gemini) but I am so happy to see another high quality American open model (I consider gemma to be high quality, like qwen).

I enjoyed turning off web search for Inkling to experiment with what innate knowledge is encoded in the model weights. A fun thing I do is check what innate knowledge very large models contain about me, as an individual. Inkling has an interesting concise shadow of what I do. (I have written a lot of books, so I am in training data.)

mark_l_watson··on Ask HN: What Are You Working On? (July 2026)
I retired several years ago and now I spend my allotted ‘tech time’ learning. Thirty five years ago I used to make up for having the ‘wrong’ degree (Physics instead of Computer Science) by writing books (Springer-Verlag, McGraw-Hill, etc.) and these books helped me get work and do business.

Now I write to learn new things myself and my mode of operation has changed radically: I still manually research and write code but I also ‘flesh out’ my writing and coding experiments using agentic tools. My approach is different because I don’t care at all about work efficiency so I spend a lot of time reading and studying material I generate with AI, and iterating in ways that wouldn’t make sense in a business environment. I still write open content books on whatever currently interests me but I am thinking, after 35 years of being an author, of winding down this activity because I think my readers might be better off doing their own research and coding, and with agentic design, research, and coding tools this gets easier.

BTW, I miss the old way of doing things, and some of what I do now is attempting to make AI less noticeable to me as I work, and using AI to work on my own work environment.

Everyone has different interests and abilities, and if what we might call the current AI hype cycle has any lasting value (open question) it is in letting us as humans ‘go our own way’ and customize tech and our life style to match what we want.

mark_l_watson··on Building and shipping Mac and iOS apps without opening Xcode
Well you can use Apple Containers - more convenient than Docker. Alternatively: I am not enthusiastic about Anthropic as a company but I use Claude Code with the three best open coding models (for low cost and speed) and opus when I need it: if you configure Claude Code correctly you protect !/.ssh, etc. I would rather use OpenCode but I am more comfortable with CC (and of course turn off telemetry).

re: this article: excellent! I always do about 90% of my Swift dev on the command line and I learned new tricks from the article. Bookmarked.

mark_l_watson··on Vint Cerf, “father of the Internet”, is retiring
Exactly right. I have only had one long conversation with him, but he was friendly and very interesting.
mark_l_watson··on A road to Lisp: Why Lisp
I use LLM_based coding agents frequently with Lisp languages and my way of working is different with Lisp languages: 1) if generated code ever has a syntax error, I like to quickly fix the syntax error myself. 2) for some reason I usually prefer to run tests myself in another terminal (with Python, Typescript, etc. I let the coding harness run tests).

I haven’t let coding harnesses run REPLs. When I do let a coding harness run tests I specify in, for examp,e, my Common Lisp skills file to run ‘sbcl -load …” so bash test commands are one liners.

It would be interesting to work on skills and a harness to use REPLs - nothing bad about that idea, I just haven’t tried it.

mark_l_watson··on Apple Silicon Exec Explains Mac Mini AI Demand and On-Device Future
Those things also work for me but Siri is often so slow that it would be faster to do myself.
mark_l_watson··on Apple Silicon Exec Explains Mac Mini AI Demand and On-Device Future
I am on the developer beta for iOS/iPadOS/macOS and all I see is still just the OpenAI option.

That said, Siri seems a little bit better now - my subjective opinion. It is a little bit less frustrating.

mark_l_watson··on Show HN: Davit, a Apple Containers UI
Nice, I like the functionality and that it is a tiny native SwiftUI app. I recently wrote a blog [1] on using Apple containers for agentic coding; I just updated it to mention Davit. I so much prefer using Apple containers to using Docker on my two home Macs for personal projects.

[1] https://open.substack.com/pub/marklwatson/p/running-opencode...

mark_l_watson··on GLM 5.2 and the coming AI margin collapse
re: "So, first, by no measure is GLM5.2 as good as Opus."

I accept that for you and your work this is true.

I have a different experience: for a month I paid big money for Opus and got a lot done. Now I am gorging on GLM 5.2 running on Fireworks.ai and I am also getting a lot done for about 15% of the money.

Everyone should do their own evals on their own work.

mark_l_watson··on Claude Sonnet 5
Good point, I also like to do the work myself, with an assistant under my control. I am usually really happy with DeepSeek v4 Flash that I feel just mostly does what I tell it to do, but I do switch to Pro for harder tasks.

There are so many models, and I personally ignore benchmarks so it takes some time to try different models on my use cases. Fortunately, it is ‘good enough’ to do the work to find a few models that work for me, and just use them for a month or two before re-investing time for my own evals to possibly change models.

People should evaluate what works for them and ignore other people and benchmarks. (Apologies if that sounds snarky.)

mark_l_watson··on Lumo 2.0
I like Lumo, glad to pay about $12/month for a private LLM with good conversation management.
mark_l_watson··on Qwen 3.6 27B is the sweet spot for local development
I disabled many built in skills and increased the context size. I also use little-coder that is based in pi.
mark_l_watson··on Qwen 3.6 27B is the sweet spot for local development
There are several general types of tasks that a Gemma 4 12B class model works for me, including: 1) design a large project composed of small libraries that can be coded and tested in isolation. 2) clean up old coding projects: add README files, comment code, show an example of using a new API and have it update API use, etc.

All small-scale stuff. For large integrated projects I am finding DeepSeek v4 Pro commercial API to be very inexpensive and helps me produce good results.

mark_l_watson··on Qwen 3.6 27B is the sweet spot for local development
I can come close to agreeing because queen-3.6-27b is my second favorite for local coding. I am using gemma4:26b-a4b-it-qat-48k (the "-48k" is from my modifying a model run with Ollama to always use a 48K context size). On a 32G Mac I use gemma4:26b-a4b-it-qat-48k and OpenCode and on my 16G MacBook Air I use gemma4:12b-it-qat-16k ("-16k" is my resizing context size) and little-coder. I break up projects into small libraries because local coding works better for me using small code bases.

I find that for local coding, I need to spend a lot of time building concise SKILLs for specific things I work on and try to only enable one or two skills per coding session.

To the author of the linked article nice job, and if you feel like adding to it, please add details on your setup.

mark_l_watson··on Google limits Meta's use of its Gemini AI models
Misleading title on HN but an interesting article, a reminder of why the hyper scalers are investing heavily in infrastructure.

That said, I expect much of the AI bubble to pop. Google Gemini with Antigravity is a good product, as is a Claude Code subscription but I have switched to using DeepSeek v4 Pro with the Claude Code harness and DeepSeek v4 Flash with the OpenCode harness (when I am not using local models with little-coder/pi) and at least for the foreseeable future I don’t think I am going back. Fast APIs at low cost trumps having to spend a little more time to get the same quality of results.

mark_l_watson··on Previewing GPT‑5.6 Sol: a next-generation model
That sounds correct: Pro for longer agentic tasks, Flash is fine for writing short programs, finding things for me in a large code base, etc.
mark_l_watson··on Previewing GPT‑5.6 Sol: a next-generation model
Your experience with DeepSeek v4 Flash differs from mine: while I usually use DeepSeek v4 Pro (that is also inexpensive), I find using DeepSeek v4 Flash with the Fireworks.ai API and properly configured OpenCode to be very good for routine work, and it is pleasantly very fast. Admittedly I use DeepSeek v4 Pro for difficult problems.

I encourage people to at least once a month to do a quick evaluation with their own problems and workflows. Estimate cost as both what inference tokens cost for a task and also how much human effort it takes to get required results.

I disregard benchmarks.

mark_l_watson··on Apple raises prices of MacBooks, iPads
I was expecting this. Glad I just upgraded my wife and myself in December.

One fix for this problem: Allow US companies to buy memory chips from China. I saw an article about a month ago, that if my memory is correct in this, said that China is ramping up high-end memory manufacturing.

Fix number two: my country (USA) should cease and desist with the craziness that is data center buildouts for AI.

Clearly ‘BIG MONEY’ always needs a new thing (cloud -> crypto -> AI) and the powerful get what they want.

If the US Congress acted to benefit regular people rather than special interests (both party's are corrupt, disbelieve that if you want to live in a fantasy land) then anti-dumping laws would be passed.

If all companies and individuals paid the real price for tokens, then we collectively would work more efficiently. As is, the filthy rich get even filthier, and regular people will get screwed.

mark_l_watson··on GLM-5.2 is a step change for open agents
OpenCode Go looked intriguing and I spent time reading their docs and pricing but didn’t purchase services. Do you think they are running it at a loss to get market share? (Probably not.) I have been happy buying tokens directly from DeepSeek (I am retired and everything I do is open source code and writing open content books (the manuscript files are available along with the source code) so I have no privacy issues). I also use FireWorks.ai to try different models. Both API services are excellent, but I may try OpenCode Go for a month or two to support the devs of OpenCode.
← PreviousPage 5 of 34Next →