Thats the exact issue with gpt. You don't know how its making the summary. It could very well be wrong in parts. It could be oversummarized to a bla bla state like you say. There's no telling whether you have outputted garbage or not, at least not without secondary forms of evidence that you might as well use anyway and drop the unreliable language model. You can summarize everything with traditional statistical methods too. On top of that people understand what tradeoffs are being made exactly with every statistical methods, and you can calculate error rates and statistical power to see if your model is even worth a damn or not. Even just doing some ML modelling yourself you can decide what tradeoffs to make or how to set up the model to best fit your use cases. You can bootstrap all these and optimize.