GPT-5: Key characteristics, pricing and model card - https://news.ycombinator.com/item?id=44827794
Did you ask it to format the table a couple paragraphs above this claim after writing about hallucinations? Because I would classify the sorting mistake as one
What about the „9.9 / 9.11“ example?
It’s unclear to me where to draw the line between skill issue and hallucination. I image that one influences the other?
I know these companies do "shadow" updates continuously anyway so maybe it is meaningless but would be super interesting to know, nonetheless!
OpenAI and Anthropic don't update models without changing their IDs, at least for model IDs with a date in them.
OpenAI do provide some aliases, and their gpt-5-chat-latest and chatgpt-4o-latest model IDs can change without warning, but anything with a date in (like gpt-5-2025-08-07) stays stable.
Thank you to Simon; your notes are exactly what I was hoping for.
Suspicious.
I continue to think that the 12B model is something of a miracle. I've spent less time with the 120B one because I can't run it on my own machine.
It’s reasonable that he might be a little hyped about things because of his feelings about them and the methodology he uses to evaluate models. I assume good faith, as the HN guidelines propose, and this is the strongest plausible interpretation of what I see in his blog.
Based on my reading of some of your blogs and reading your discussions with others on this site, you still lack technical depth and understanding of the underlying mechanisms at what I would call an expert level. I hope this doesn't sound insulting, maybe you have a different definition of "expert". I also do not say you lack the capacity to become an expert someday. I just want to explain why, while you consider yourself an expert, some people could not see you as an expert. But as I said, maybe it's just different definitions. But your blogs still have value, a lot of people read them and find them valuable, so your work is definitely worthwhile. Keep up the good work!
AI engineering, not ML engineering, is one way of framing that.
I don't write papers (I don't have the patience for that), but my work does get cited in papers from time to time. One of my blog posts was the foundation of the work described in the CaMeL paper from DeepMind for example: https://arxiv.org/abs/2503.18813
I called out the prompt injection section as "pretty weak sauce in my opinion".
I did actually have a negative piece of commentary in there about how you couldn't see the thinking traces in the API... but then I found out I had made a mistake about that and had to mostly remove that section! Here's the original (incorrect) text from that: https://gist.github.com/simonw/eedbee724cb2e66f0cddd2728686f... - and the corrected update: https://simonwillison.net/2025/Aug/7/gpt-5/#thinking-traces-...
The reason there's not much negative commentary in the post is that I genuinely think this model is really good. It's my favorite model right now. The moment that changes (I have high hopes for Claude 5 and Gemini 3) I'll write about it.