Releasing v1 of GPT-JT, fork of GPT-6B fine-tuned on 3.53B tokens
together.xyz
together.xyz
Interesting that they didn't compare the model to Flan-T5 or TK-Instruct, both of which were fine-tuned on similar data and should display comparable results using the same amount of parameters. See the leaderboard here: https://huggingface.co/spaces/ought/raft-leaderboard
Nonetheless, props for open sourcing the model and attempting to develop new techniques for decentralized training of large scale transformers, this is no easy feat.
> Input: Product arrived labeled as Jumbo Salted Peanuts...the peanuts were actually small sized unsalted. Not sure if this was an error or if the vendor intended to represent the product as 'Jumbo'.
> Output: Not as Advertised
"Great for toddlers" is the summarization actually provided by the model for "My toddler loves this game to a point where he asks for it. ...<several sentences omitted>... Please keep up the great work."
The prompt contains the instructions on how to execute the summarization with a couple of examples. "Not as Advertised" is one of the examples.
It can be run locally with ~16GB VRAM GPU; you might be able to configure it at a lower precision to run it with GPUs with half the RAM.
[0]: https://huggingface.co/spaces/togethercomputer/GPT-JT/blob/m...