From a random web search, it seems the sizes above Large are: Extra Large, Jumbo, Extra Jumbo, Giant, Colossal, Super Colossal, Mammoth, Super Mammoth, Atlas.
Followed by another company introducing their "Plus Ultra" model.
And my experience isn't unique in any way here and it's really hard to not see it pervasive through our culture.
They're not merely real values, they're also rational.
I should say though, that's the only place I've seen this particular localization.
You mean the EU, right? The UK isn't covered by the AI act.
/s
We could take a page from Trump’s book and call them “Beautiful” LLMs. Then we’d have “Big Beautiful LLMs” or just “BBLs” for short.
Surely that wouldn’t cause any confusion when Googling.
- Extremely Low Frequency (ELF)
- Super Low Frequency (SLF)
- Ultra Low Frequency (ULF)
- Very Low Frequency (VLF)
- Low Frequency (LF)
- Medium Frequency (MF)
- High Frequency (HF)
- Very High Frequency (VHF)
- Ultra High Frequency (UHF)
- Super High Frequency (SHF)
- Extremely High Frequency (EHF)
- Tremendously High Frequency (THF)
Maybe one day some very smart people will make Tremendously Large Language Models. They will be very large and need a lot of computer. And then you'll have the Extremely Small Language Model. They are like nothing.
https://en.wikipedia.org/wiki/Radio_frequency?#Frequency_ban...
https://en.m.wikipedia.org/wiki/Overwhelmingly_Large_Telesco...
Horrendous being based on the Latin root for "trembling with fear", tremendous on another Latin root meaning "shaking from excitement" and terrible deriving from a Greek root for, again, "trembling with fear".
I’ve seen corporate slogans fired off from the shoulders of viral creatives. Synergy-beams glittering in the darkness of org charts. Thought leadership gone rogue… All these moments will be lost to NDAs and non-disparagement clauses, like engagement metrics in a sea of pivot decks.
Time to leverage.
Alternatively, just make sure you keep things consensual, and keep yourself safe, no judgement or labels from me :)
XXLLM: ~1T (GPT4/4.5, Claude Opus, Gemini Pro)
XLLM: 300~500B (4o, o1, Sonnet)
LLM: 20~200B (4o, GPT3, Claude, Llama 3 70B, Gemma 27B)
~~zone of emergence~~
MLM: 7~14B (4o-mini, Claude Haiku, T5, LLaMA, MPT)
SLM: 1~3B (GPT2, Replit, Phi, Dall-E)
~~zone of generality~~
XSLM: <1B (Stable Diffusion, BERT)
4XSLM: <100M (TinyStories)
teensy 4B to 29B
smol 30B to 59B
mid 60B to 99B
biggg 100B to 299B
yuuge 300B+
Gotta leave room for future expansion.
LLM 3.0, LLM 3.1 Gen 1, LLM 3.2 Gen 1, LLM 3.1, LLM 3.1 Gen 2, LLM 3.2 Gen 2, LLM 3.2, LLM 3.2 Gen 2x2, LLM 4, etc...
I really appreciated the way they managed to come up with a new naming scheme each time, usually used exactly once.
(Obviously ∞ is for the actual singularity, and ℵ₁ is the thing after that).
https://en.m.wikipedia.org/wiki/Continuum_hypothesis
;-)
But the quality of Apple Intelligence shows us what happens when you use a tiny ultra-low-wattage LLM. There’s a whole subreddit dedicated to its notable fails: https://www.reddit.com/r/AppleIntelligenceFail/top/?t=all
One example of this is “Sorry I was very drunk and went home and crashed straight into bed” being summarized by Apple Intelligence as ”Drunk and crashed”.
Personally, I think the summaries of alerts is incredibly useful. But my expectation of accuracy for a 20 word summary of multiple 20-30 word summaries is tempered by the reality that’s there’s gonna be issues given the lack of context. The point of the summary is to help me determine if I should read the alerts.
LLMs break down when we try to make them independent agents instead of advanced power tools. Alot of people enjoy navel gazing and hand waving about ethics, “safety” and bias… then proceed to do things with obvious issues in those areas.
- Passed out drunk
- Crashed in bed
- Slacking because drunk
...
The issue isn't a lack of context; it's that even the available context was handled poorly.
I actually applied to YC in like ~2014 or such for thus;
-JotPlot - I wanted a timeline for basically giving a histo timeline of comms btwn me and others - such that I had a sankey-ish diagram for when and whom and via method I spoke with folks and then each node eas the message, call, text, meta links...
I think its still viable - but my thought process is too currently chaotic to pull it off.
Basically looking at a timeline of your comms and thoughts and expand into links of thought - now with LLMs you could have a Throw Tag od some sort whereby you have the bot do work on research expanding on certain things and plugging up a site for that Idea on LOCAL HOST (i.e. your phone so that you can pull up data relevant to the convo - and its all in a timeline of thought/stream of conscious
hopefully you can visualize it...
So in that sense, maybe people would prefer a private alternative.
(Fyi I was a designer at fb and while it was luxious I still hated what I saw in zucks eyes every morn when I passed him.
Super diff from Andy Grove at intel where for whateveer reason we were in the sam oee schekdule
(That was me typing with eues ckised as a test (to myself, typos abound
It's like saying Automated ATM. Whoever wrote it barely knows what the acronym means.
This whole article feels like written by someone who doesn't understand the subject matter at all
I.e. when pretty much every tool or script I used before doesn't work anymore, and need a special tool (gsutil, bq, dusk, slurm), it's a mind shift.