HNHacker News
TopNewBestAskShowJobs

simonw

120,393 karma · joined October 29, 2007

JSK Fellow 2020. Creator of Datasette, co-creator of Django. Co-founder of Lanyrd, YC Winter 2011.

https://simonwillison.net/ and https://til.simonwillison.net/

submissionscomments
simonw··on Claude Haiku 5.5
Pelicans riding bicycles for Haiku at the different thinking levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Low messes up the bicycle frame, but medium/high/xhigh/max all get the bicycle frame right.

The max one took 5 minutes 9 seconds and cost 3.3826 cents. The cheapest one (low) cost 0.0936 cents and took 7 seconds.

The most recent release of my llm-anthropic plugin queries the Anthropic model listing API directly, so I didn't have to upgrade the plugin to add support for this model:

  llm install llm-anthropic -U                                
  llm anthropic refresh
  llm -m claude-haiku-5.5 'prompt goes here'
EDIT: Here's the Haiku 4.5 pelican from a year ago for comparison, it was terrible: https://simonwillison.net/2025/Oct/15/claude-haiku-45/
simonw··on Claude Haiku 5.5
My complaint about Haiku 4.5 was that it was 10x the price of GPT-6 Luna.

> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens

Haiku and Luna now have the exact same price up to 100,000 tokens. Luna is now cheaper for anything after 100,000 tokens, even after Luna's own price increases at 270,000 it's still less than Haiku.

So it sounds like they've directly addressed that problem. Their self-reported benchmarks are all higher than Luna too.

simonw··on Why were Victorian elites so effective?
> Academic life was dominated by Latin and, to a lesser extent, Greek. The exact share varied, but the classical languages took up generally at least half of teaching hours, and sometimes as much as 80 percent, especially earlier in the century.

That's so weird to me. How did that tradition hold on for so long? Where did it come from? Who thought that was a good idea?

simonw··on Decisions API is in public beta
Depends on the quality of the results. These things are driven by text prompts. If it turns out the OpenAI one returns better quality results than open weight variants they'll be rewarded by the market.

Anyone using a decision model like this is going to have to spin up their own evals - these are far harder to vibe-check than regular text output LLMs.

simonw··on Decisions API is in public beta

  curl https://api.openai.com/v1/decisions \
    -H "Authorization: Bearer $(llm keys get openai)" \
    -H "Content-Type: application/json" \
    --data '
  {
    "model": "gpt-6-luna",
    "input": [{
      "role": "user",
      "content": [
        {"type": "input_text", "text": "I am angry about the new product feature"}
      ]
    }],
    "questions": [{
      "type": "predicate",
      "name": "complaint",
      "instructions": "Is this a complaint?"
    }, {
      "type": "predicate",
      "name": "compliment",
      "instructions": "Is this a compliment?"
    }]
  }'
Returned:

  {
    "model": "gpt-6-luna",
    "answers": [
      {
        "type": "predicate",
        "name": "complaint",
        "probability": 0.91
      },
      {
        "type": "predicate",
        "name": "compliment",
        "probability": 0.06
      }
    ],
    "usage": {
      "input_tokens": 310,
      "input_tokens_details": {
        "cached_tokens": 0,
        "cache_write_tokens": 0
      },
      "output_tokens": 0,
      "output_tokens_details": {
        "reasoning_tokens": 0
      },
      "total_tokens": 310
    }
  }
That https://api.openai.com/v1/decisions endpoint is notable because usually when OpenAI define an endpoint like that it ends up as a defecto standard for other providers.

(I turned this all into a new llm plugin: https://github.com/simonw/llm-openai-decisions)

simonw··on Paramount Skydance has completed its $111B merger with Warner Bros. Discovery
Nilay Patel: https://www.theverge.com/podcast/1004286/senator-adam-schiff...

> One of the longest-running tropes at The Verge is that a workable antitrust policy for the United States would be simply to forbid companies from buying Time Warner because it has never worked and it may never work again, and that we should just pass one law.

Checks out:

Jan 11, 2001: AOL and Time Warner merged to create AOL Time Warner

June 14, 2018: AT&T acquired Time Warner, renamed WarnerMedia

April 8, 2022: WarnerMedia spun out of AT&T, merged with Discovery

simonw··on A sustainable web career, for when all this blows over
Assuming you lived through those cycles, can you tell us why?
simonw··on EmbeddingGemma 2: An open, lightweight multimodal embedding model
I really appreciate that EmbeddingGemma 2 is under the Apache 2.0 license.

For embedding models in particular, I don't think it makes sense to use a closed, proprietary, hosted-only model.

Most applications of embedding models involve calculating thousands or even millions of embedding vectors and storing them for later comparison.

If your model is proprietary, the vendor is likely someday going to decide to stop offering that model. They'll have a better model to replace it, but you still need to pay to re-calculate those millions of stored existing vectors.

(In April 2024 OpenAI offered to "cover the financial cost of users re-embedding content with these new models" - https://openai.com/index/gpt-4-api-general-availability/ - but I don't think that's something we can rely on from every provider.)

Notably, I don't want to host the model myself. I'd much rather pay a provider for a hosted model while knowing that if they ever stop hosting it I can run the open weights version myself - or find another vendor who can do that for me.

simonw··on Mistral Large 4
That's a GitHub rate limit. Try again now, I just pushed a hopeful fix: https://github.com/simonw/tools/commit/7793fb74c2d37bd613cdc...
simonw··on Show HN: Parseable, an open observability datalake, handles 100M time-series/min
The open source version doesn't accept protobufs, but does accept JSON.

I found this out because I set Codex the task of running this locally agains another of my apps and it worked around the limitation by running this proxy: https://gist.github.com/simonw/b0e61a0aa8e3f7d30f27ce2f747c9...

... but it turns out my stack can emit JSON just fine, so I switched to that instead. Here's me TIL write-up of getting Parseable running locally https://til.simonwillison.net/datasette/datasette-parseable-...

simonw··on Mistral Large 4
> The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.

OK well I couldn't resist this one:

  llm -m claude-opus-5.5 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
  llm -m gpt-6.1-sol 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
  llm -m gemini-3.8-flash 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
  llm -m mistral/mistral-large-4 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
Default reasoning levels for each: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
simonw··on Vibecoding isn't as fun as writing code by hand
It's heavily biased by the circles I hang out in, but I'm seeing a loose correlation between how much people enjoy working with coding agents and their depth of experience. I've even seen people come out of retirement because this stuff is so much fun for them.

AI tools amplify existing coding experience, and reward management experience too. Engineers with a decade+ of experience are more likely to have been engineering leads or spent time in engineering management.

Related observation: Anthropic are developing a reputation for hiring former CTO/CEO/founders and putting them back to work as individual contributors.

simonw··on Mistral Large 4
I think they're still visually pretty different. The most common shared details are:

- Pelican cycling to the right - that's been discussed at length, images of bicycles online always show that side of the bike because that's where the chain is.

- Bicycle is usually red. No idea! Red ones go faster?

simonw··on The era of software quality, or the era of ostriches?
Insecure code is incorrect code.
simonw··on Mistral Large 4
Surprisingly it only supports reasoning "none" or reasoning "high".

That setting didn't seem to make any real difference - it added a tiny bit of thinking trace and high actually produced less output tokens than none.

The high bicycle frame is better then the none one though.

Pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

(Definitely the best I've seen from any Mistral model: https://simonwillison.net/tags/pelican-riding-a-bicycle+mist... )

simonw··on Ephemeral Testing
I've been using this pattern quite a bit recently for API design, and I really like it.

The big challenge with designing an API is that the only way to be confident in the design is to build a bunch of different things on top of it. But why invest all that effort in an API that you don't think is ready yet?

With coding agents the cost of building those prototypes drops to almost nothing. I can exercise a proposed API design five different ways before I commit to the shape.

simonw··on The era of software quality, or the era of ostriches?
Have you tried having a recent model audit your code for potential security issues recently?

I found that quite humbling. I thought I was a lot better at writing secure code than that.

simonw··on People are asking ChatGPT to help them decide how to vote in the midterms
That's why I said "xAI may be the exception here."
simonw··on People are asking ChatGPT to help them decide how to vote in the midterms
We heard their concerns. My argument here is that any attempts within an AI lab to bias political answer would come to light.
simonw··on People are asking ChatGPT to help them decide how to vote in the midterms
The employees of the AI labs. They're a pretty vocal bunch, and would be very likely to quit, whistleblow, or both.

(xAI may be the exception here.)

simonw··on "I'm Embarrassed on Behalf of the Tech Industry"
Watching regular people try to use technology is so painful.

The other day I helped a man in a cafe get dictation working on his new MacBook Neo. It took three of us (two very technical) 10 minutes to figure out.

We had to turn OFF Siri so that the F5 button would enable dictation in the way that he wanted. We then had to coach him on the difference between tapping F5 to toggle dictation on and off as opposed to holding down F5 constantly.

And Apple are meant to be good at this stuff!

simonw··on Web Search API
The phrase "reasonably necessary" is infuriatingly vague.
simonw··on Anthropic wants your thoughts on AI
Surprising number of comments in here complaining about a company asking their users for ideas, when that's pretty much the first piece of business advice any startup will get.

Talk to your users!

simonw··on Web Search API
My number one question about search APIs is always if they allow you to store and resyndicate results you get from them.

If I'm running an agent system but I'm not allowed to store the responses - or provide a "share transcript" button - that's a pretty significant limitation.

The answer to that question is inevitably buried deep in the terms. Here's the relevant section I found for Ceramic, in their list of things you can't do:

> (n) collect, aggregate, store, or compile Output, including search results, relevance scores, or rankings, for the purpose of creating or contributing to any database, dataset, index, or corpus, whether or not such database, dataset, index, or corpus is used for a purpose that competes with Ceramic; (o) resell, syndicate, or otherwise make Output available to any third party on a standalone basis or as a separately accessible component of another product or service; provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query, and is not independently accessible, extractable, or downloadable by end users or third parties; or (p) retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use, unless expressly permitted in an applicable Order Form.

https://www.ceramic.ai/terms-of-service

Am I alone in caring about this?

simonw··on We're going to need default hard budget caps on pretty much everything
Today's hobbyist is tomorrow's decision maker at work over which service to use.
simonw··on We're going to need default hard budget caps on pretty much everything
See https://docs.aws.amazon.com/accounts/latest/reference/create... - they pause your access but don't delete your data for a 90 day grace period.
simonw··on We're going to need default hard budget caps on pretty much everything
Amazon's new feature for this specifically says that it won't delete any of your data for 90 days:

> If you take no action within 90 days of your project being paused, AWS permanently deletes your project data.

From https://docs.aws.amazon.com/accounts/latest/reference/create...

simonw··on We're going to need default hard budget caps on pretty much everything
A hospital should select the checkbox that says "no spending limit".
simonw··on We're going to need default hard budget caps on pretty much everything
It's definitely technically difficult. You can't easily estimate how much an operation is going to cost before you kick off that operation, which means as soon as you get close to the limit you are at risk of tripping it.

Consider something like a "select * from bigtable" SQL query that might process a trillion rows. Hard to know that's going to cost $100 until after you have run it.

simonw··on We're going to need default hard budget caps on pretty much everything
I don't really understand that argument. This seems pretty obvious to me, as a customer. Is this really something that companies don't understand?

Sending an email when your budget gets low shouldn't be a big lift.

← PreviousPage 3 of 34Next →