HNHacker News
TopNewBestAskShowJobs

xscott

1,566 karma · joined August 5, 2019

submissionscomments
xscott··on Detecting and countering misuse of AI: September 2026
I had Muse Glimmer (from Meta / Facebook) quoting OpenAI's safety guidelines to me, and I had Poolside's Laguna (a smaller US company) with thinking traces about obeying Chinese law.

Both of those are local models, and I didn't provide them tools to access the internet to call other models. None of this is proof of anything, but it is suggestive.

xscott··on Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
> A common issue is that it's rarely mentioned on which dataset KL-divergence is computed. It seems the most common dataset is wikitext

Thank you for calling this out. Using Wikipedia snippets for these is a terrible choice. I did a bunch of KL and other stats with the five Gemma 4 models, and the results were non-obvious. Anthropomorphizing:

Gemma 4 31B: "I guess we'll pretend I said this, but it's not me." (Baseline for stats)

Gemma 4 26B: "Dude, I'm certain I wouldn't have said this." (Bad KL)

Gemma 4 12B: "Umm, Me either!" (Similarly Bad KL)

Gemma 4 E4B: "I might say almost anything, this is fine." (Much better KL!!!)

Gemma 4 E2B: "I'm basically a toy. Let's play a game!" (Same KL as E4B)

Anyway, for comparing quantizations, it seems like the largest precision version should be given a one-shot prompt, and the result from that should be used as the corpus for the quantized versions.

xscott··on GLM-5.3 is now open-weight
Maybe I'm reading too much between the lines, but I suspect the reason is to rub his nose in the duplicity or naivety depending on how generous you're feeling. Publishing the model would be a confession that he was wrong.

AI policy is being shaped somewhat by the things Sam and Dario say. So even if you're not feeling vindictive, it's probably good to keep a track record of the previous things they have said as a Bayesian prior. People who don't know better listen to these people, and maybe they shouldn't.

xscott··on Most Physicists Don't Agree on Most Things
There's a link you can click to see the personal background of the respondents. It's tough to know what "researcher in a field not listed" means, but it's possible that over 60% are not even physicists:

    30.8% - A researcher in a field not listed above
    21.2% - A science enthusiast
    18.0% - A researcher studying quantum physics
    12.0% - A researcher studying astrophysics or cosmology
     9.2% - A researcher studying gravity
     8.8% - Other
xscott··on NanoGPT Speedrun Frontier
This seems very cool, but I'm not sure I understand exactly what it's doing. Are they making a new speculative drafter for Qwen 3.8 27B? Maybe they're optimizing the MLX code for the decoder itself? Thank you in advance.
xscott··on Could AIs Become Conscious?
Your definition is not useful enough. There are people missing any or all of those, and any or all of those can be approximated by a machine.
xscott··on Could AIs Become Conscious?
I've thought about turning this upside down. To any person who is sure they know what has and doesn't have consciousness: If I say I don't have it, can you prove me wrong?

As far as I'm concerned, it's a word without a useful enough definition to bother worrying about it.

xscott··on Every Model Cheats
I'm not claiming to have any expertise in this area, but I've got a list of things I try to apply when working with LLMs. Possibly relevant here is, "don't tell the model what NOT to do, show it what TO do". I think guard rails should be implemented outside the model with an isolated system. The models seem to like patterns to follow.

Anyway, this article reads a lot like, "the beatings will continue until cheating is eliminated". Maybe try a carrot instead of a stick.

xscott··on Ornith-1.5: From Self-Scaffolding to Self-Improvement
People over-quantize things, muck with the temperature and other settings based on superstitions or results from models they think are similar. There's lots of ways to make 3.8 27B dumber.
xscott··on Bun 1.4 Rust rewrite is not looking good
Not that my opinion matters much, but I like Deno. I never tried Bun.
xscott··on GPT-5.6 Sol Pricing Cut by 50% on OpenRouter
Same here - my questions were sophomore level. I think it's notable that when I edited my question to say it was about Gemma 4, it answered without blocking. A cynic like myself would interpret that as evidence they don't care about sharing information if it involves their competitors.
xscott··on GPT-5.6 Sol Pricing Cut by 50% on OpenRouter
What's the distinction between "protecting their turf" and "competitive reasons"? I see them as the same, but I could be missing something.

That's an interesting thought on the current "safety blocking" being a trial run for the topics that scare people (bio). You're more charitable about their motives than I am, but you might be right.

xscott··on GPT-5.6 Sol Pricing Cut by 50% on OpenRouter
I've gotten flagged for asking questions about tokens and tensors. That makes me believe it's not about safety, it's about protecting their turf. I cancelled my subscription - same fear about getting flagged too much leading to a ban.
xscott··on On AI regulation and messaging
All good, but that seems unrelated to what I said, and I'm not sure why you replied to me.

Consider this though: Regulating AI models in the US benefits the data centers, not individuals. That's where regulated models run. Regulated models creates demand for data centers.

Dario wants regulation so he can continue to sell access to his data centers.

xscott··on On AI regulation and messaging
I think you might misunderstand. The regulation isn't about what OpenAI and Anthropic can do. It's about what you, a citizen, can do.
xscott··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
Yeah, I think I'm seeing the same thing. I don't have all the answers, I just think it'd be a mistake to throw the baby out with the bath water on this model. It seems significantly better than the other dense models near the same size (Gemma 4, Muse Glimmer, etc...). Maybe harness changes, maybe fine tunes or LoRAs.
xscott··on On AI regulation and messaging
That's definitely not an argument I was making.
xscott··on On AI regulation and messaging
Any argument about regulation in the US which doesn't mention that China and other countries aren't bound by that regulation should be heavily questioned. Exactly who are you stopping from doing what you don't like?!?

Law abiding US citizens won't be able to run them, but you didn't solve anything other than making sure US citizens pay Sam or Dario.

xscott··on On A.I. regulation and messaging
I don't really understand the argument you're making, but just to add a data point:

DeepSeek V4 Flash 0731 is 167 gigabytes from the developer and as a GGUF with no additional quantization. It limps along on my 192GB M2 Mac from several years ago [0]. This model tests better[1] than Claude Opus 4.6 released in February. That's six months ago - what will be available 6 months from now?

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tr...

https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF

So yeah, enthusiasts aren't going to run frontier models on their gaming machines, but a small office could easily justify the $30k - $100k cost to run something like this at high speed. The small company I worked for routinely spent that kind of money on Dec Alphas twenty five years ago, and that's not accounting for inflation adjustment.

And this is completely discounting the advances smaller models are making. You're right that Qwen 3.8 comes in different sizes. However, Qwen 3.8 27B and Qwen 3.6 27B do run on gaming cards, and they're better than the frontier models from twelve months ago.

I have no idea what will happen in the future, but I wouldn't base my guesses solely on the largest open weight models.

[0] Yes, it's unpleasantly slow (5-8 tok/sec)

[1] Yes, benchmarks should be taken with a lot of salt.

xscott··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
It won't satisfy the people who just want to drop a model into their existing toolset and run, but I think there are a lot of ways to deal with this overthinking problem.

For instance, it's a step backward, but I put {"reasoning_effort":"none"} and led it by the nose:

   User: We're going to make <silly demo>.  Please create a plan, but do not write code yet.

   Agent: <short and reasonable plan>

   User: Now please follow that plan and write the code.  No other chat.

   Agent: <reasonable code in reasonable time>
Maybe this can be fixed with Jinja templates or something, or maybe it's a hack to your harness, but it shows you can get the model to reason reasonably.
xscott··on Patterns and problems in emerging multi-agent systems
> [...] we evaluated the behavior of various Claude models in a setting with contradictory objectives.

> We consistently saw a multiagent turf war... In fact, they sabotaged others with increasingly aggressive, self-replicating malware.

Seems like Anthropic should withdraw their models until they can be taught to behave and cooperate as well their competitors (both open and closed) do. /s

I hate fearmongering, and I don't trust Dario's intentions for doing it.

xscott··on Anthropic shares details about how Claude's new watermarks will work
One possibility is that you have Claude write something that would not get you in trouble and concatenate that with something you wrote by hand that would.

Your signature from the safe material is now associated with your content on the dangerous part, and then the brownshirts come knocking on your door.

xscott··on RISC-V: They Should Have Known Better
2^N I think, but who's counting.
xscott··on GLM-5.3: Frontier coding with emergent cyber capabilities
So many possibilities for how you could glue it all together. However, when I send Gemma 4 12B in llama.cpp an image with no accompanying text, it assumes I want a description and gives me one.

I just tried with an audio file, and it transcribed the lyrics as I hoped. Then it made a bunch of suggestions about what do next, which I wasn't after. I could probably fix that by sending some text to narrow the scope.

xscott··on Qwen 3.8 27B
So much potential for that channel. He's got a nice range of tests and a no nonsense presentation style.

However, watching tests of heavily quantized models that weren't designed for it (non-QAT) is frustrating. There's no way to tell if the actual model fails because it's dumb or if the lobotomy made it that way.

xscott··on Qwen 3.8 27B
You're very right about KL divergence. I spent a couple days playing with the Gemma 4 models. That's 10 separate models (varying weights, MoE, QAT or not, etc...) with identical tokenizers. I treated 31B at BF16 as the gold standard, feeding Wikipedia snippets, and anthropomorphizing a bit:

Gemma 4 31B: "Um, if I really said all of that, I guess I'd say this next"

Gemma 4 26B: "Dude, I would've said completely different stuff" (large divergence)

Gemma 4 12B: "Umm, there's zero chance I would've said some of this" (INFINITE divergence)

Gemma 4 E4B and E2B: "Derp derp, I'm happy to say almost anything" (lowest divergence)

For models which are chat trained, they simply would not recite Wikipedia, so the divergence is almost meaningless. I thought about capturing a realistic coding session and trying to use that as the corpus, but you need to preserve the turn-based tokens and such, so I moved on to other things.

xscott··on GLM-5.3: Frontier coding with emergent cyber capabilities
Probably not what you're after, but I've considered having a separate small mm-model act as a seeing-eye dog for the bigger more capable one.
xscott··on Rapid warming may tip AMOC at 2°C, slower warming may avert collapse
You're going to punish Bolivia, Uzbekistan, Vietnam, Sri Lanka, and Congo?!?

I don't think that will help: https://www.worldometers.info/co2-emissions/co2-emissions-by...

Maybe you'll suggest changing that to be per capita? Punishing Palau, Qatar, and Kuwait?

xscott··on Beef and dairy drive 41% of biodiversity damage linked to global farmland
How much more than 6 percent? Support your claim with reputable sources.

If your arguments aren't believable, I'm buying extra steak this evening.

xscott··on Beef and dairy drive 41% of biodiversity damage linked to global farmland
Actually focusing on CO2, same site:

https://ourworldindata.org/ghg-emissions-by-sector

More than 70% of CO2 from energy, less than 6% from ALL livestock.

Page 1 of 25Next →