Biorecap: An R package for summarizing bioRxiv preprints with a local LLM
blog.stephenturner.us
blog.stephenturner.us
Evaluating the quality of summaries is also very difficult. How would you evaluate whether a bioRxiv paper is well summarized or not? In the scientific literature you will find that one way is to let human domain experts score various dimensions of the produced summaries e.g. using a 4-point Likert scale. Alternatively, professional summarizers produce "reference summaries" that serve as a model, and automatic summaries can be evaluated automatically against these using string overlap metrics.
People could remain in a "development environment" not just to develop something or to manipulate numbers but for more generic data usage. Aside Quarto reports who can mix text, formatting and live code, like org-mode is another classical thing in a modern environment.
It's VERY positive to me, because it's a simple sign of another step toward classic computing, a thing we have lost and that's desperately needed.