There's also a model comparison spreadsheet that you can compare sizes and such https://weavers.neocities.org/architecture-encyclopedia/mode...
If you'd like any additional models to be added I can add them in.
1,127 karma · joined May 15, 2018
There's also a model comparison spreadsheet that you can compare sizes and such https://weavers.neocities.org/architecture-encyclopedia/mode...
If you'd like any additional models to be added I can add them in.
You're right the T5 stuff is very important historically but they're below 11B and I don't have much to say about them. Definitely a very interesting and important set of models though.
When you train it to be an assistant model, it's better at compressing assistant transcripts than it is general text.
There is an eval which I have a lot of interested in and respect for https://huggingface.co/spaces/Jellyfish042/UncheatableEval called UncheatableEval, which tests how good of a language model an LLM is by applying it on a range of compression tasks.
This task is essentially impossible to 'cheat'. Compression is a benchmark you cannot game!
> it somehow merged Llama 4 Maverick's custom Arena chatbot version with Behemoth
I can clarify this part. I wrote 'There was a scandal as facebook decided to mislead people by gaming the lmarena benchmark site - they served one version of llama-4 there and released a different model' which is true.
But it is inside the section about the llama 4 model behemoth. So I see how that could be confusing/misleading.
I could restructure that section a little to improve it.
> Llama 405B was also trained on more than 15 trillion tokens[1],
You're talking about Llama 405B instruct, I'm talking about Llama 405B base. Of course the instruct model has been traiend on more tokens.
> why is there such a focus on token training count?
I tried to include the rough training token count for each model I wrote about - plus additional details about training data mixture if available. Training data is an important part of an LLM.
Thank you for spotting the error.
How?
It's not the function that is wrong, it's those people.
An LLM doesn't know anything about itself - it can be pre-prompted with facts about itself, but this is going to be an example of it just making plausible text up.
I think you need to quantize the model yourself from the float/huggingface versions. My understanding is that the quantization formats have changed recently. and old quantized models no longer work.
A good place to dig for prompt structures may be the 'text-generation-webui' commit log. For example https://github.com/oobabooga/text-generation-webui/commit/33...