No, the model has nothing do to with Llama. We are using our own architecture, and training from scratch. Llama also does not have open training data, and is non-compliant, in contrast to this model.
Source: I'm part of the training team
Source: I'm part of the training team
Good luck though, very needed project!
Can you comment on how the filtering impacted language coverage? E.g. finweb2 has 1800+ languages, but some with very little actual representation, while finweb2-hq has just 20 but each with a subdsantial data set.
(I'm personaly most interested in covering the 24 official EU languages)