The problem is that there are a lot, at least 30, of these small projects scattered around, funded for a few years as some ad-hoc temporary coalition of universities and businesses. Those simply cannot compete with businesses spending tens of billions on developing these. Especially when you have to bring a spoon to a gunfight restricting to "clean" data.
Multilinguality is essentially a solved problem, and restricting too much on one language with more limited resources is gonna make the model worse in that language too.
I'd rather have smaller european labs try to give it a go at distributed training. If multiple countries got together and said, "look, we tried training a distributed model that speaks in all of our local languages and that is comparable to 1-year-old Chinese open-source models", that, at least, I would find interesting.
I don't think your approach would work because you can't create a strong model from distilling several weak models.
Maybe I'm more attuned to this type of thing having grown up as a national of a smaller state living in the shadow of a bigger state but you constantly see actors from the bigger state belittling and condescending anything contributed socially or economically from the smaller state.
And I see this sort of dynamic here in this forum where Americans very frequently talk condescendingly like this about Europe generally and European tech especially (they did it to China too but China smartly ignored this self-interested nonsense and carried on anyway which is what Europeans should do).
It really grates on me and presumably many others. But it serves an agenda too of a lot of the founders and financiers that hang out here that have big fat customers in Europe they'd like to keep sweet and competitors they'd like to keep down.
Sure… they can, except at the end of the day it’s a bit late, regulatory burden will make it comparatively useless, and because of that nobody will ever use it. It will be spending a bunch of taxpayer dollars for press releases.
The running joke is that when these “sovereign” EU models launch, they’re going to refuse to answer anything that might involve personal information such as Elon Musk’s birthday.
I challenge the assumption you can do meaningful work in this field without blatant disregard for intellectual property.
The idea that it’s all down to training size is clearly incorrect, as every expert human learned their craft without nearly the sum total information of the internet. Clearly there are architectural wins to be found.
Besides that, why would everyone just be fine with Opus level AI at best, as that’s all the US is willing to export, and I doubt China will share beyond that.
Sovereign AI is more important than ever after Friday.
I agree there is likely some hubris in this sort of announcement, but investing in European expertise and industrial base in this area is important.