Gemma 3n is a model utilizing Per-Layer Embeddings to achieve an on-device memory footprint of a 2-4B parameter model.
At the same time, it performs nearly as well as Claude 3.7 Sonnet in Chatbot Arena.
Gemma 3n is a model utilizing Per-Layer Embeddings to achieve an on-device memory footprint of a 2-4B parameter model.
At the same time, it performs nearly as well as Claude 3.7 Sonnet in Chatbot Arena.
What's the catch?
I can give more details if you (or anyone else!) is interested.
TL;DR: it is scoring only for "How authoritative did the answer look? How much flattering & emojis?"
Not sure if they've shared more since.
IMVHO it won't help, at all, even if they trained a perfect model that could accurately penalize it*
The main problem is its one off responses, A/B tested. There's no way to connect it into all the stuff we're using to do work these days (i.e. tools / MCP servers), so at this point its sort of skipping the hard problems we'd want to see graded.
(this situation is a example: whats more likely, style control is a small idea for an intractable problem, or Google has now released multiple free models better than Sonnet, including the latest, only 4B params?
To my frustration, I have to go and bench these things myself because I have an AI-agnostic app I build, but I can confirm it is not the case that Gemma 3-not-n is better than Sonnet. 12B can half-consistently make file edits, which is a major step forward for local tbh)
* I'm not sure how, "correctness" is a confounding metric here: we're probably much more likely to describe a formatted answer in negative terms if the answer is incorrect.
In this case I am also setting aside how that could be done, just saying it as an illustration of no matter what, it's the wrong platform for a "how intelligent is this model?" signal, at this point, post-Eliza post-Turing, couple years out from ChatGPT 1.0
In general, all those scores are mostly useful to filter out the models that are blatantly and obviously bad. But to determine whether the model is actually good at any specific thing that you need, you'll have to evaluate them yourself to find out.
edit: I seem to be the only one excited by the possibilities of such small yet powerful models. This is an iPhone moment: a computer that fits in your pocket, except this time it's smart.
Leave that part out, I'm excited. I'd love to see this plays some roles in "inference caching", to reduce dependencies on external services.
If only agents can plan and match patterns of tasks locally, and only needs real intelligence for doing self-contained/computationally heavy tasks.
If you don't believe me, here is a fun mental exercise: define "understand" and "reason" in a measurable way, that includes humans but excludes LLMs.
> The `foobar` is also incorrect. It should be a valid frobozz, but it currently points to `ABC`, which is not a valid frobozz format. It should be something like `ABC`.
Where the two `ABC`s are the exact same string of tokens.
Obviously nonsense to any human, but a valid LLM output for any LLM.
This is just one example. Once you start using LLMs as tools instead of virtual pets you'll find lots more similar.
But a LLM is sometimes the genius and sometimes the idiot.
That doesn’t happen often if you always talk to the same person
But for LLMs these kinds of illogical statements are common.
LLMs don't think; what they do cannot be in good faith called "thinking" by any definition.
They trained models on only task specific data, not on a general dataset and certainly not on the enormous datasets frontier models are trained on.
"Our training sets consist of 2.9M sequences (120M tokens) for shortest paths; 31M sequences (1.7B tokens) for noisy shortest paths; and 91M sequences (4.7B tokens) for random walks. We train two types of transformers [38] from scratch using next-token prediction for each dataset: an 89.3M parameter model consisting of 12 layers, 768 hidden dimensions, and 12 heads; and a 1.5B parameter model consisting of 48 layers, 1600 hidden dimensions, and 25 heads."
Ask these questions again in two years when the next winter happens.
I'd go as far as saying LLMs are meaning made incarnate - that huge tensor of floats represents a stupidly high-dimensional latent space, which encodes semantic similarity of every token, and combinations of tokens (up to a limit). That's as close as reifying the meaning of "meaning" itself as we ever come.
(It's funny that we got there through brute force instead of developing philosophy, and it's also nice that we get a computational artifact out of it that we can poke and study, instead of incomprehensible and mostly bogus theories.)
Bruh. Do you need to be paid to interact in good faith or were you raised to be social
Now would I take AI as a trivia partner? Absolutely. But that's not really the same as what I look for in "smart" humans.
Note that "smarter than smart humans" and "smarter than most humans" are not the same. The latter is a pretty low bar.
If not, I strongly encourage you to discuss your area of expertise with it and rate based on that
It is incredibly competent
But potentially maybe I'm just not looking for a trivia partner in my software.
compared to what you are used to right?
I know it's elitist but most people <=100 iq (and no, this is not exact obviously, but we have not many other things to go by) are just ... well, a lot of state of the art LLMs are better at everything compared, outside body 'things' (for now) of course, as they don't have any. They hallucinate/bluff/lie as much as the humans and the humans might know they don't know, but outside that, the LLMs win at everything. So I guess that, for now, people with 120-160 iqs find LLMs funny but wouldn't call them intelligent, but below that...
My circle of people I talk with during the day has changed since I took on more charity which consists of fixing up old laptops and installing Ubuntu on them; I get them for free from everyone and I give them to people who cannot afford, including some lessons and remote support (which is easy as I can just ssh in via tailscale). Many of them believe in chemtrails, vaccinations are a gov ploy etc and multiple have told me they read that these AI chatbots are nigerian or indian (or so) farms trying to fraud them out of 'things' (they usually don't have anything to fraud otherwise I would not be there). This is about half of humanity; Gemma is gonna be smarter than all of them, even though I don't register any LLM as intelligence and with the current models, it won't happen either. Maybe a breakthrough in models will be made that changes it, but it has not much chance yet.
This is incorrect, IQ tests are normally scaled such that average intelligence is 100, and such that they are approximately normally distributed so that most people will be somewhere between 85-115 (66% on average).
> average intelligence is 100
You both are saying the same thing.
IQ is defined such that both average and mean would be equal 100. The combination of sub-100 and exactly-100 would be more people than above-100, hence "most people <=100 iq".
85 is special housing where I live... LLMs are far beyond that now.
Not only is this unbelievable, it's reprehensible
I refuse to touch the IQ bait