My understanding is that this workload is very uncommon: the Tailscale blog says they used a custom unusual configuration to have many checkpoints like this. So without knowing the bug, it seems unlikely one would build this workload and then find the bug. If that makes sense.
Edit: just want to say that you being able to repro it is awesome, but that the overall claim seems a little overstated to me.
https://news.ycombinator.com/item?id=49278424 https://news.ycombinator.com/item?id=49278521
Also appreciate the nice words at the end :) I'm feeling a bit ganged up on.
I guess what would have been an even more cool thing is "we ran some more general testing with Antithesis, and it found five other bugs". Have you thought along those lines or explored something like that? There have to be other, similar bugs lurking in SQLite :)
Like some others have mentioned, one of my earliest thoughts was “how much of a hint was the LLM given about the bug?” I think if the prompt used was stated clearly/verbatim near the beginning of the article, that would probably dispel a good amount of the criticism.
It states you replicated the bug once the SQLite team fixed it, and published it.
Not sure what’s difficult about replicating behaviour when it’s spelled out for you.
As-is, this is just 20/20 hindsight with concerns about leading the AI on through the prompt hand-waved away. Come on.
Oh btw, we also found some other ones… Stay tuned!