I would have imagined such a thing would be smaller and thus run on smaller configurations.
But since I am only a layman maybe someone can tell me why this isn't the case?
I would have imagined such a thing would be smaller and thus run on smaller configurations.
But since I am only a layman maybe someone can tell me why this isn't the case?
Also, the software you’re working on will generally in some way have a real-world domain - without knowing it the AI all likely be a less effective assistant. Design conversations with it would likely be pretty non-fun, too.
Finally, the “bitter lesson” article[0] from a couple years ago is I think somewhat applicable too.
[0]: http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Meta does have some specialized models though, llamaguard was released for llama 2 and 3.
The expensive part is building the dataset, training itself isn't too expensive (you can even fine-tune small models on free colab instances!), and when you have your dataset, you can just fine tune the next generalist model as soon as it's released and you're good to go ago.
It's just a fancy name for sparse evaluation of the total network to save compute and memory bandwidth.