HNHacker News
TopNewBestAskShowJobs

simonw

119,361 karma · joined October 29, 2007

JSK Fellow 2020. Creator of Datasette, co-creator of Django. Co-founder of Lanyrd, YC Winter 2011.

https://simonwillison.net/ and https://til.simonwillison.net/

submissionscomments
simonw··on Jevmem – automatic project memory for Claude Code, built on Jev
Yeah, I've started dialing back my use of AI for READMEs because of this.

My previous rule was that I never use AI for writing that expresses my own opinions or tries to be convincing (anything on my blog for example) but I'll let it do technical documentation.

The top of a README is about convincing and explaining why I built something though, which means it should fit my no-AI policy after all.

simonw··on Opus 5.5 is good at explainer videos
I trust my workflow a tiny bit more, because I don't know how that one works.

Plus having the video file means I can extract frames as images at specific timestamps.

Homely though the ask feature is probably fit for purpose, at least on YouTube.

simonw··on Opus 5.5 is good at explainer videos
I sometimes run yt-dlp against YouTube or TikTok URLs to extract an mp4, then send that to Gemini to ask questions about the content.

It's particularly good for those frustrating recipe videos where they don't tell you quantities of ingredients to use - you can have Gemini hallucinate the quantities and it normally gets them about right.

simonw··on Unknown number of Texas voter registrations went unprocessed due to DPS error
TIL a great deal of voter registration in Texas still happens using paper forms: https://www.votebeat.org/texas/2024/10/28/voter-registration...

And apparently they sometimes get lost or processed incorrectly, leaving people disenfranchised.

simonw··on Once Claude can measure something, it can make it faster
Yeah, culture. I expect most of the web developers at Anthropic are of the generation that considers 21MB of JavaScript a perfectly reasonable way to build a web app.
simonw··on Gemini 3.8 text-to-speech
Bit of both. I define "vibe coding" as building without looking at the code at all. For this one I was reading the code to understand how it works, and I then followed up several times to adjust how it was working.

Full session here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

simonw··on OpenAI breaches Medicare, Albanese reveals
If it was the Australian Medicare Statistics Reporting Service on June 18th it may have been part of this incident: https://collusion.wiki/ - the bulk of that coordinated activity was between 16th and 21st of June, and we know they were hitting UK government data sites.

I had a dig around in the data that they published on that site and found references to www.aihw.gov.au and viz.aihw.gov.au and vizprod.aihw.gov.au

simonw··on We just shipped support for the ugliest part of HTTP: Vary
I've been wanting this from Cloudflare for years.

The classic problem here is if you do that thing where user agents that send "accept: text/html" get HTML, while user agents that don't get JSON or some other format.

This used to be impossible to deploy behind Cloudflare caching, because they ignored the Vary header on anything other than images - so you risked caching the JSON version and then serving it up to someone who was expecting HTML.

(Independent of the Cloudflare feature I ended up deciding never to use that pattern, because I prefer having URL that predictably returns HTML or JSON - I add a .json suffix to my apps to serve JSON instead.)

simonw··on Jev introduces a new shape of LLM
It feels like an LLM variant to me. It's a large model, trained in large amounts of text, that you prompt using human languages. The only difference is that the output is a set of scores as opposed to more text.
simonw··on Gemini 3.8 text-to-speech
Example (topical, pelican themed) audio clip here: https://simonwillison.net/2026/Sep/23/gemini-tts-playground/
simonw··on Once Claude can measure something, it can make it faster
I visited https://claude.ai/ over a mobile tethered connection from my laptop the other day and was pleasantly surprised at how quickly it loaded.

(That said, I just had a look in Firefox and it loads 20.78 MB of JavaScript (6.84 MB compressed) so I expect they could make it a bunch lighter if they kept trying.)

simonw··on Gemini 3.8 text-to-speech
Sure. Feel free to copy the code and run it yourself instead.
simonw··on Gemini 3.8 text-to-speech
Because building my own is a better way to understand the capabilities of the model and how to use it - and to verify that it can be used via CORS.
simonw··on Gemini 3.8 text-to-speech
The iPhone has a built in voice cloner hidden in the accessibility settings for exactly this use-case: creating a backup of your voice in case you need it in the future.
simonw··on Gemini 3.8 text-to-speech
I vibe coded a playground UI for trying this out. The conversation mode is neat, and it's very expensive - most of my experiments have cost less than a cent.

https://tools.simonwillison.net/gemini-tts-playground#compos...

simonw··on Gemini 3.8 text-to-speech
> Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.

I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.

simonw··on Jev Can't Be Calibrated
Fine tuning LLMs has turned out to be mostly not worth the effort, but I wonder if fine tuning Jev-style models will turn out to be a whole lot more useful.
simonw··on GPT-6 Sol and Luna
It shows that the three 5.6 models are closely enough related that they exhibit similar "taste" in their color choices, and the same is true for the 6 models.
simonw··on Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
"Drawing bicycles with pelicans is superhuman."

It really isn't. Many humans can draw a bicycle, and a pelican, just fine.

simonw··on The darker side of being a doctor
Important context that's not obvious from the article - this is about the Australian healthcare system. I'd assumed it was the USA.
simonw··on The darker side of being a doctor
Huh, you're right! This is by an Australian doctor:

https://ericlevi.com/ - "Paediatric & Adult Specialist Otolaryngologist, Ear Nose & Throat, Head & Neck Surgeon" in Melbourne.

simonw··on The Download: why AI's latest breakthroughs and fears may be more hype than rea
Both authors. Stochastic Parrots was Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell - the linked article is by Emily M. Bender and Timnit Gebru.
simonw··on The Download: why AI's latest breakthroughs and fears may be more hype than rea
Is this meant to link to the article as opposed to a newsletter that mentions the article?

Article link is: https://www.technologyreview.com/2026/09/22/1144867/dont-be-...

simonw··on The darker side of being a doctor
Wow.

> I had worked in a hospital network that covered 4 campuses and drove 500kms a week when covering these sites. I had worked in a hospital where I didn’t get home for days at a time, sleeping overnight in hospital quarters, outpatient clinic benches and in my car.

As a patient, I'd like the person performing surgery on me to be well-rested!

It gets worse:

> I used to be able to arrange the operating list because I know that some operations take longer than others. But now, the bookings office determine that that all my tonsillectomies take 14 minutes because that’s the average time recorded on the computer. The moment I scrub in, the timer starts. The moment I unscrub timer stops. Click. Click. Click. Because the theatre bookings does not take into account the interpreter time, pre-med period or transfer to ICU, the list is running late. The nurse in charge is breathing down my neck to finish on time.

And yet somehow that 14 minute tonsillectomy gets billed at ~$10,000.

This seems to me like a system that has been hyperoptimized in a way that grinds down the participants.

simonw··on Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
The fact that bicycles are hard for people to draw was one of the inspirations for the test.
simonw··on GPT-6 Sol and Luna
If they train for the benchmark, how come many of the pelicans produced by their different models at different reasoning levels still suck?

That aside, the relevance these days is in comparing models and effort levels within the same model families - hence the comparison grids.

simonw··on GPT-6 Sol and Luna
Yeah, I have a couple of variants that I want to get working:

1. Each model gets three chances, and then gets to pick the best according to its vision input

2. Models run in a loop where they can produce SVG, see it rendered, and then edit it further

I tried that loop last year and had disappointing results, but the models are a lot more effective this year.

simonw··on GPT-6 Sol and Luna
That's a good idea. It's the default output for my `llm logs` command, but that header could at least show the model ID.

OG images will require me to move away from publishing in a Gist and linking to from a JavaScript page that loads the Gist. Probably worthwhile though.

simonw··on GPT-6 Sol and Luna
https://simonwillison.net/tags/pelican-riding-a-bicycle/ but I need to build something better.
simonw··on GPT-6 Sol and Luna
The table on https://developers.openai.com/api/docs/pricing is more readable:

  +--------------+-------+--------------+--------------+--------+
  | Model        | Input | Cached input | Cache writes | Output |
  +--------------+-------+--------------+--------------+--------+
  | gpt-6-luna   | $0.10 | $0.01        | $0.125       | $0.50  |
  | gpt-5.6-luna | $0.20 | $0.02        | $0.25        | $1.20  |
  +--------------+-------+--------------+--------------+--------+
← PreviousPage 2 of 34Next →