- They benchmarked against general models like GPT-3 but not well-established specific models that have been trained for specific tasks like SPECTER[0] or SciBert[1]. Specter outperformed GPT-3 on tasks like citation prediction two years ago. Nobody seriously uses general LLMs on science tasks, so nobody who actually wants to use this cares about your benchmarks. I want to see task-specific models compared to your general model, otherwise whats going to happen is I either need to run my own benchmarks or, much more likely, I shelve your paper and never read it again. If you underperform some that's fine! If you don't compare to science-specific models all you're claiming is that training on science data gives better science results... thats not exactly an impressive finding. Fine-tuning is a separate thing, I get it, but pleeeeeease just give the people what they want.
- Not released on huggingface. No clue why not. On the back-end this appears to be based on OPT and huggingface compatible, so I'm really confused.
- Flashy website. Combine 1&2 with a well designed website talking about how great you are and most of my warning lights got set off. Not a fan.
@authors, if you're lurking, please release more relevant benchmarks for citation prediction etc. Thanks.
[0] - https://arxiv.org/abs/2004.07180 [1] - https://arxiv.org/abs/1903.10676