I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.
We have a lot of details in the tech report if you want to go deeper.
I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!
Why expose yourself to this liability?