They are not good at that. The space of possibilities can be massive and LLMs are terrible at exploring such space because they predict from the prior tokens they made. They are inherently bad at exploring new space because it is antithetical to how they work.
To me it's deeply concerning that so many people are getting fooled into thinking that LLMs are actually good at covering their bases like you're describing here. It's one of their weakest qualities.
> Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible.
It usually does quite a bad job at this too, often the tests it wrote feel like that of a lazy student that didn't really want to do the task and just sort of cheats at it or does a really shallow job. It certainly cannot run the software in every scenario possible.
> Anyway all of this reads like someone who is not actually using LLMs to build software or hasn't tried them in a while
This is something that is said by everyone who contests anyone pointing out the risks in being overly trusting of AI or otherwise points out their flaws. I use the latest and greatest all the time and all the time I'll point out something that it totally overlooked and get hit with the classic "you're absolutely right". This occurs because I actually read the code and can see the myriad of blatant issues that still occur when using LLMs and know better than to trust them. You will find so many issue by delving into the details.