Evaluating faithfulness and content selection of LLMs in book-length summaries
arxiv.org
arxiv.org
I would imagine that summaries of non-fiction books are evaluated quite differently from summaries of fiction.
I've been trying to figure out what prompts they used. The https://github.com/mungg/FABLES GitHub repo says this:
Summary -- (str) Entire book summarized
by one of five models: Mixtral,
GPT-3.5-Turbo, GPT-4, GPT-4-Turbo, and
Claude-3-Opus, using the hierarchical
merging method described in Chang et al..
With a link to https://arxiv.org/pdf/2310.00785.pdf - which then links to another GitHub repository, https://github.com/lilakk/BooookScore which has a bunch of prompts in https://github.com/lilakk/BooookScore/tree/main/promptsWhich makes me think that this original paper isn't evaluating LLMs so much as it's evaluating that one particular prompting technique for long summaries.
Gemini Pro 1.5 has 1 million token context length, which should remove the need for weird hierarchical summary tricks. I wonder how well it would score?
I suspect that they were preparing this for press by the time Gemini 1.5 Pro was released.