What’s better, train a model with 10X parameters once on some default hyperparameter setting or to search for a good hyperparameter configuration by training on X parameters 10 times? While I’m at it, how many LLMs of the size of GPT3 were trained until they landed on the capability of GPT3? How much of this is dependent on the data, or do good settings transcend the type of text that a model is trying to train on?