GPT-3! Thank god I don't have to hand write them lol.
How are you getting past the token limit?
Likely only summarizing the abstracts.
Yup only abstracts for now, but I think it should be possible to parse the paper PDF to extract just the approach + experiment results w/ GPT-3 (despite 4k token limit). Will add this soon.
I would suggest sentence tokenization followed by clustering to pull the most representative sentence from each cluster. Unless you want to go abstractive, that is.
If you run into issues with the token limit, there are some nice architectures for handling longer inputs (e.g., Longformer, Reformer, etc.).