I wish rather than stopping training early they would have run more data through a small model so we could have something more competitive with LLaMA 7B.
I wish rather than stopping training early they would have run more data through a small model so we could have something more competitive with LLaMA 7B.
"We chose to train these models to 20 tokens per param to fit a scaling law to the Pile data set. These models are optimal for a fixed compute budget, not necessarily "best for use". If you had a fixed parameter budget (e.g., because you wanted to fit models on certain hardware) you would train on more tokens. We do that for our customers that seek that performance and want to get LLaMA-like quality with a commercial license"
I'd chip in!
There are plenty of such efforts, but the organizer needs some kind of significance to attract a critical mass, and a AI ASIC chip designer seems like a good candidate.
Then again, maybe they prefer a bunch of privately trained models over an open one since that sells more ASIC time?
This is really weird to hear out loud.
I still think of Discord as a niche gaming chatroom, even though I know that (for instance) a wafer scale IC design company is hosting a Discord now.