Super cool project! Though, aren't most modern LLM's ar transformers internally?
Super cool project! Though, aren't most modern LLM's ar transformers internally?
A large language model is just what it says--a very large statistical model trained for language tasks. This covers the spectrum of GPT-style models, but also those hard to classify ones, like Liquid's "Liquid Foundation Models", which can get up to 24 billion parameters and use grouped query attention, but are closely related to state-space models as well: https://huggingface.co/LiquidAI/LFM2-24B-A2B
Also, as others have pointed out, a Transformer isn't inherently a language model. So really they're sort of two different axes, one classifying the model size and task, the other referring to a specific architecture.
Till then, transformers were being used primarily for stuff like translation and such and no one was even pretraining at scale, even tho transformers and attention existed.
Openai and google if you count T5 persisting with pretrained generative models was what led to the LLM boom. Yes they used transformers, but that's just one IMO minor aspect.
The “L” in LLM’s generally refers to human-language specifically. You’d expect to feed it…human text. Nitpicking that the human text also constitutes a mathematical language is like, correct, but so general as to be unhelpful.