well sure, but i mean realistically speaking, we cannot step debug an llm's output to find out what happened given the way we currently execute inference
We absolutely can with mechanistic interpretability & companies like Anthropic, OpenAI, Meta, and Google do precisely this do debug their models.