FYI Mistral at launch just dropped the weights without any model architecture mentioned.
Most of the OSS models follow the same architecture which is Llama +- a few things, so it wasn't too hard for people to make it work.
Most of the OSS models follow the same architecture which is Llama +- a few things, so it wasn't too hard for people to make it work.
It used to be the case until last year, but now almost every Chinese model come with their own linear attention mechanism.