Both of those are local models, and I didn't provide them tools to access the internet to call other models. None of this is proof of anything, but it is suggestive.
1,566 karma · joined August 5, 2019
Both of those are local models, and I didn't provide them tools to access the internet to call other models. None of this is proof of anything, but it is suggestive.
Thank you for calling this out. Using Wikipedia snippets for these is a terrible choice. I did a bunch of KL and other stats with the five Gemma 4 models, and the results were non-obvious. Anthropomorphizing:
Gemma 4 31B: "I guess we'll pretend I said this, but it's not me." (Baseline for stats)
Gemma 4 26B: "Dude, I'm certain I wouldn't have said this." (Bad KL)
Gemma 4 12B: "Umm, Me either!" (Similarly Bad KL)
Gemma 4 E4B: "I might say almost anything, this is fine." (Much better KL!!!)
Gemma 4 E2B: "I'm basically a toy. Let's play a game!" (Same KL as E4B)
Anyway, for comparing quantizations, it seems like the largest precision version should be given a one-shot prompt, and the result from that should be used as the corpus for the quantized versions.
AI policy is being shaped somewhat by the things Sam and Dario say. So even if you're not feeling vindictive, it's probably good to keep a track record of the previous things they have said as a Bayesian prior. People who don't know better listen to these people, and maybe they shouldn't.
30.8% - A researcher in a field not listed above
21.2% - A science enthusiast
18.0% - A researcher studying quantum physics
12.0% - A researcher studying astrophysics or cosmology
9.2% - A researcher studying gravity
8.8% - OtherAs far as I'm concerned, it's a word without a useful enough definition to bother worrying about it.
Anyway, this article reads a lot like, "the beatings will continue until cheating is eliminated". Maybe try a carrot instead of a stick.
That's an interesting thought on the current "safety blocking" being a trial run for the topics that scare people (bio). You're more charitable about their motives than I am, but you might be right.
Consider this though: Regulating AI models in the US benefits the data centers, not individuals. That's where regulated models run. Regulated models creates demand for data centers.
Dario wants regulation so he can continue to sell access to his data centers.
Law abiding US citizens won't be able to run them, but you didn't solve anything other than making sure US citizens pay Sam or Dario.
DeepSeek V4 Flash 0731 is 167 gigabytes from the developer and as a GGUF with no additional quantization. It limps along on my 192GB M2 Mac from several years ago [0]. This model tests better[1] than Claude Opus 4.6 released in February. That's six months ago - what will be available 6 months from now?
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tr...
https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF
So yeah, enthusiasts aren't going to run frontier models on their gaming machines, but a small office could easily justify the $30k - $100k cost to run something like this at high speed. The small company I worked for routinely spent that kind of money on Dec Alphas twenty five years ago, and that's not accounting for inflation adjustment.
And this is completely discounting the advances smaller models are making. You're right that Qwen 3.8 comes in different sizes. However, Qwen 3.8 27B and Qwen 3.6 27B do run on gaming cards, and they're better than the frontier models from twelve months ago.
I have no idea what will happen in the future, but I wouldn't base my guesses solely on the largest open weight models.
[0] Yes, it's unpleasantly slow (5-8 tok/sec)
[1] Yes, benchmarks should be taken with a lot of salt.
For instance, it's a step backward, but I put {"reasoning_effort":"none"} and led it by the nose:
User: We're going to make <silly demo>. Please create a plan, but do not write code yet.
Agent: <short and reasonable plan>
User: Now please follow that plan and write the code. No other chat.
Agent: <reasonable code in reasonable time>
Maybe this can be fixed with Jinja templates or something, or maybe it's a hack to your harness, but it shows you can get the model to reason reasonably.> We consistently saw a multiagent turf war... In fact, they sabotaged others with increasingly aggressive, self-replicating malware.
Seems like Anthropic should withdraw their models until they can be taught to behave and cooperate as well their competitors (both open and closed) do. /s
I hate fearmongering, and I don't trust Dario's intentions for doing it.
Your signature from the safe material is now associated with your content on the dangerous part, and then the brownshirts come knocking on your door.
I just tried with an audio file, and it transcribed the lyrics as I hoped. Then it made a bunch of suggestions about what do next, which I wasn't after. I could probably fix that by sending some text to narrow the scope.
However, watching tests of heavily quantized models that weren't designed for it (non-QAT) is frustrating. There's no way to tell if the actual model fails because it's dumb or if the lobotomy made it that way.
Gemma 4 31B: "Um, if I really said all of that, I guess I'd say this next"
Gemma 4 26B: "Dude, I would've said completely different stuff" (large divergence)
Gemma 4 12B: "Umm, there's zero chance I would've said some of this" (INFINITE divergence)
Gemma 4 E4B and E2B: "Derp derp, I'm happy to say almost anything" (lowest divergence)
For models which are chat trained, they simply would not recite Wikipedia, so the divergence is almost meaningless. I thought about capturing a realistic coding session and trying to use that as the corpus, but you need to preserve the turn-based tokens and such, so I moved on to other things.
I don't think that will help: https://www.worldometers.info/co2-emissions/co2-emissions-by...
Maybe you'll suggest changing that to be per capita? Punishing Palau, Qatar, and Kuwait?
If your arguments aren't believable, I'm buying extra steak this evening.
https://ourworldindata.org/ghg-emissions-by-sector
More than 70% of CO2 from energy, less than 6% from ALL livestock.