> Or you think some human or team of humans could have manually parsed some logs to provide an unsloppy analysis?
When we have an error or issue in the $WORK codebase on LIVE/PROD, that's precisely what we do. We sit down, analyze the logs for our services over the relevant date ranges and try to piece together exactly what happened and why. We have a huge number of logs too, but thanks to the magic of proper SWE (which you'd think OAI would have with their magic AIs) we've managed to partition our observability tooling so that you can digest only what you need.
That's basically how any serious organization does things, instead of just throwing a non-deterministic black box at the problem. Especially because logs are by their very nature noisy, and they will saturate any model's context window very quickly leading to massive hallucinations and what ultimately amounts to making shit up that isn't anywhere in the logs (ask me how I know)