If LLMs were good at summarization, this wouldn't be necessary. Turns out a stochastic model of language is not a summary in the way humans think of summaries. Thus all this extra faff.
I think you are still stuck with try if it works for you and hope it generalizes beyond your evaluation.
The task itself is not very well-defined. You want a lossy representation that preserves the key points -- this may require context that the model does not have. For technical/legal text, seemingly innocuous words can be very load-bearing, and their removal can completely change the semantics of the text, but achieving this reliably requires complete context and reasoning.
[information content of summary] / [information content of original] for summaries of a given length cap?
Examples: https://eugeneyan.com/writing/evals/#summarization-consisten...