Calling these models open source is like calling a binary open source because you can download it.
Which in this day and age isn't far from where were at.
Calling these models open source is like calling a binary open source because you can download it.
Which in this day and age isn't far from where were at.
Synthetic data can be more diverse if you sample carefully with seeded concepts, and it can be more complex than average web text. You can even diff against a garden variety Mistral or LLaMA and only collect knowledge and skills they don't already have. I call this approach "Machine Study", where AI makes its own training data by studying its corpus and learning from other models.
Mistral models are one example, they never released pre training data and there are many fine tunes.
Is anyone else just assuming at this point that virtually everyone is using the pirated materials in The Pile like Books3?