At $DAY_JOB nowadays we run 128x H100 runs without thinking twice nowadays. Only takes a few days to train a small-ish LLM with that to test out some ideas.
To preempt other queries, maybe just paste the `set` output.
I don’t get why people do this at all. It adds no clarity only awkwardness over “At my job”
"At my job" is annoying and awkward too if you're not going to specify where when making grand claims. It's not quite as annoying as "In my country we..." without specifying where, but it's close.
Ah yes, like people saying "back in my day" and not even giving a precise date, smh.
> echo $DAY_JOB
FAANG AI Lab
Where are they hosted?
AWS and GCP both.
Out of curiosity, what leads you to train models from the ground up rather than fine tuning existing models?
We do both. You can’t just fine tune if you’re trying a different model architecture, or even change some of the hyperparameters on an existing one. Every now and again you might be able to reuse some of the weights, but that’s about it. That’s part of the reason research is so incredibly expensive and time consuming in this field. I bet that $80k is only a fraction of the overall cost for the model described in the article, too.