Smarter summaries with finetuning GPT-3.5 and chain of density
jxnl.github.io
jxnl.github.io
There's still tons of very low hanging fruit in the summarization work. I'm not aware of significant followup work to pointer networks besides pointer-generator networks, which these days are considered old news. Pointer based architectures are the ideal system for word level extractive summarizers, yet the very best extractive summarization systems today are usually nothing more than sentence selectors using some kinds of embeddings and cosine similarity.
Happy to see such success with abstractive summaries, but the kind that myself and most other humans are interested in is still far from solved.
Most fine tunes will have much larger datasets (I am under the impression you want 10’s of thousands of examples for most runs).
So I’m similarly impressed 20 examples would make such a big difference.
But also note entity density decreases as example count increases. This is counterintuitive — maybe something else is going on here?
imo summarization is also a fairly simple task -- I wouldn't be surprised if a fine-tuned open source model (eg llama 13 / mistral 7) would get to similar performance.
However, thanks to this article I might revisit the summarization techniques used a try a fine tuned 3.5.
It would be great to see these techniques compared to Chat GPT 4 Turbo.
The issue was always quality, length of summarized text, and control of exactly what kind of summary.
one way i have heard is to send a chunk of data to gpt4 and ask for questions to be generated. unsure of other ways. what has worked well ?
Without seeing some kind of demonstration otherwise, my feeling is that it would be like regressing stock price on inflation, then trying to generate more data using the regression model and random inflation numbers. All you'd learn is the model that you put in to generate the data.
Even in the main article here, the model did better with fewer fine tuned examples. To us, the auto-generated examples might look different enough and might look good enough, but they were all generated algorithmically. Feeding more examples in might easily be leading it to focus on some artifact of the embeddings or generating model that we just don't perceive.
If you have a small student model and a large teacher it makes sense, the student is better off after this distillation.
If you have a way to filter out low quality synthetic examples then it would be useful to generate a bunch more and take the best.
If your LLM is an agent, then it can generate feedback signals from the environment. Even a human-AI chat is a form of environment for the model. Every human response can be evaluated as positive or negative reward.
More fundamentally, organic datasets are very unbalanced, LLMs need more complex reasoning chains than what is usually available. There are some exceptions - in scientific papers, manuals and code you get very complex reasoning chains. But not in general. This issue can be fixed with synthetic data.
And even in principle, if you have a model at level N and want to make a dataset at level N+1, then you need to boost your model. You can give it more tokens, more attempts or more tools.
Article: {{ARTICLE}}
You will generate increasingly concise, entity-dense summaries of the above Article.
Repeat the following 2 steps 5 times.
Step 1. Identify 1-3 informative Entities (";" delimited) from the Article which are missing from the previously generated summary.
Step 2. Write a new, denser summary of identical length which covers every entity and detail from the previous summary plus the Missing Entities.
Although this appears to refer to a "previous" summary, from a previous run, would would appear to suggest it is run anew for each iteration, in fact it generates an "Initial Summary:" on the first run, and generates a Step 1 and Step 2.For me though, it stops there. I can, however, simply say:
repeat again
And it does another Step 1 and Step 2, and stops.I will note that the third iteration, on an article such as ...
https://www.reuters.com/technology/cybersecurity/payments-ap...
... is indeed increasingly superior each iteration. By the third iteration (the fourth summary), it did seem ideal. The fourth iteration and fifth summary added entities that feel extraneous.
> Answer in JSON. The JSON should be a list (length 5) of dictionaries whose keys are “Missing_Entities” and “Denser_Summary”
Also I think it doesn’t make sense to write in the prompt for gpt to iterate if it is not doing the iteration. There is not templating of the step number or recursive summary injection in the sample prompt either.
Article: {{ARTICLE}}
You will generate increasingly concise, entity-dense summaries of the above Article.
Repeat the following 2 steps 5 times.
Step 1. Identify 1-3 informative Entities (";" delimited) from the Article which are missing from the previously generated summary.
Step 2. Write a new, denser summary of identical length which covers every entity and detail from the previous summary plus the Missing Entities.
A Missing Entity is:
- Relevant: to the main story.
- Specific: descriptive yet concise (5 words or fewer).
- Novel: not in the previous summary.
- Faithful: present in the Article.
- Anywhere: located anywhere in the Article.
Guidelines:
- The first summary should be long (4-5 sentences, ~80 words) yet highly non-specific, containing little information beyond the entities marked as missing. Use overly verbose language and fillers (e.g., "this article discusses") to reach ~80 words.
- Make every word count: re-write the previous summary to improve flow and make space for additional entities.
- Make space with fusion, compression, and removal of uninformative phrases like "the article discusses".
- The summaries should become highly dense and concise yet self-contained, e.g., easily understood without the Article.
- Missing entities can appear anywhere in the new summary.
- Never drop entities from the previous summary. If space cannot be made, add fewer new entities.
Remember, use the exact same number of words for each summary.
Answer in JSON. The JSON should be a list (length 5) of dictionaries whose keys are "Missing_Entities" and "Denser_Summary".
And yes, the answer as a JSON list length 5 causes 5 summaries to get spit out!However, it's not fully clear to me that it's considering the prior summaries on the later summaries in a good/useful way. The expressly iterated results I get are superior to the inline list of results.
"More research is needed." -- https://www.explainxkcd.com/wiki/index.php/2268:_Further_Res...
My original expectation was the same as the way instructor software implemented it. But I found the prompt in the article confusing toward that perspective. I’m sure it can work either way, but it should be a lot more performant (and less expensive) as a single pass.
- Never drop entities from the previous summary. If space cannot be made, add fewer new entities.
Specifically, for me it randomly drops entities.> a lot more performant
Faster? Absolutely. But I'm not having luck getting it smarter.