Finding Miscompiles for Fun, Not Profit
newsletter.semianalysis.com
newsletter.semianalysis.com
I wonder if we get to a world where a full repo sweep like this is a default Github action after commit.
Why are you using phrasing that equates AI and humans? You used Codex to write a fuzzer. It didn't decide to join you.
I had heard LLMs were finding a lot of bugs very quickly and now I can see what that looks like from a user perspective.
Does MRI low level code produce wrong images? Do some kind of unexpected http connection quirks happen? Does (LL)M inference produce randomly wrong and non reproduceable output? Graphical artifacts in video games? Application crashes that happen once every billionth request? Security vulnerabilities? Race conditions?