HNHacker News
TopNewBestAskShowJobs

anon373839

4,781 karma · joined May 21, 2023

submissionscomments
anon373839··on Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Yep, can confirm that is NOT normal. Are you using Nvidia’s NVFP4 quant? There are other NVFP4s floating around but they are not as good. The quality of the calibration data really matters.

Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)

anon373839··on As A.I. makes law firms more efficient, clients ask: 'Where's my discount?'
> Email and phone calls are considered confidential, even though it is possible for vendors to inspect the communication.

It's a weak analogy. Ordinary comms infrastructure providers actually have pretty robust policies, technical, and contractual measures in place that restrict employee access to customer communications. In contrast, in the wild west of generative AI, companies actively monitor session data for the content itself, in order to exploit it for their own business purposes. There is zero expectation of privacy.

So I don't share your expectation that precedent will uphold the use of consumer-tier services (in their current form) for handling privileged material.

anon373839··on As A.I. makes law firms more efficient, clients ask: 'Where's my discount?'
This is a completely straightforward application of existing law on privilege. To maintain privilege, among other things, communications must be confidential.

Chatting with Claude breaks confidentiality: chats with Claude are subject to arbitrary inspection by Anthropic employees, not to mention the issue of model training.

You can use self-hosted LLMs without breaking privilege. And funny enough, law firms like Latham & Watkins are now buying Nvidia GPU clusters for this purpose.

anon373839··on "As a Language Model": Chat Template Switches LLM Self-Referential Voice
God almighty, yes. Make it stop.
anon373839··on California is chasing wealth that has feet
> the argument that billionaires will leave if taxed at a higher rate isn't compelling

I agree with this. California’s climate and culture will keep many a billionaire within tax nexus reach of the state. Of course, this isn’t a strategy every locale can pursue, but I don’t see a reason for California not to exploit its advantages.

anon373839··on California is chasing wealth that has feet
> Rent is a function of supply and demand

Is there really a market dynamic in rent pricing anymore? I thought that algorithmic collusion had eliminated the need for landlords to compete on price.

anon373839··on Claude Opus 5.5
The fix for this is to tap the overflow menu icon and choose “Reduce privacy protections”. (Wtf, Alibaba?) This appears to be related to use of iCloud Private Relay.
anon373839··on MiMo v2.6
The gap really is closing, though. I don't remember the source, but there was a publication recently stating that it's 4 months at this point. So, 4 months where you can charge premium prices for the small subset of tasks where only a frontier model will do. I have to think the big AI duopoly is at an inflection point where most of their customers haven't just yet realized that they're being fleeced.
anon373839··on Fable 5 – Median thinking declined in August
Qwen 3.8 Flash-Next is not dumb. If you've used it and that was your experience, your workload is either ultra-ultra-sophisticated or you're dealing with a broken quant/buggy chat template/other issue. That model is a smart, reliable workhorse.
anon373839··on Alibaba open-sources AI model that can detect cancer and nearly 150 conditions
The model weights are only ~5GB, so this is small.
anon373839··on Claude Code now reads AGENTS.md if there is no Claude.md
The world has moved on.
anon373839··on Qwen 3.8 Omni Flash
This is a serving bug or quantization issue. I had all kinds of issues that were like this on DGX Spark until I found a single-GB10 vLLM recipe [1] that uses Nvidia's NVFP4 quant. The community quants did not work well.

Another failure mode you may see is inordinately long CoT. Properly served, the model is good at calibrating its CoT length to the difficulty of the immediate task.

[1] https://github.com/blazux/qwen3.8-Flash-DGX

anon373839··on HarnessTax: How Much Does the Harness Matter for Coding Agents?
> Personally, I’ve replaced OpenCode with a thin wrapper around Pydantic-AI as the pythonic analogue to Pi-Agent for headless use via Hermes

That's really interesting. I like Pydantic AI a lot and wondered why all of the harnesses seem to be written in Javascript instead of it. What do you use it for headless, though? I haven't tried Hermes or similar yet, so don't have a handle on what you do with them.

anon373839··on A warning about 'model welfare'
Model welfare, much like AI xrisk, is a concept born from evidence-free “what if?” questions. Some people ran with these what-ifs and developed ornate belief systems around them. And now they demand the rest of us take them seriously.
anon373839··on Due to concerns about malicious applications, GPT2 will not be released (2019)
That isn’t any kind of antitrust violation. Agreeing not to undercut each other’s prices would be, however.
anon373839··on Everyone should slow down AI development except for me
Ah, no, that’s not cheaper. Renting GPUs adds up quickly and leaves you with nothing in the end.

Renting tokens from open model providers is cheaper but it incurs the same issues: unexpected changes in model quality, inconsistent speeds, service outages.

anon373839··on AI models don't kill people – people kill people
I will say that LLMs are somewhat unlike guns in that they emit text.
anon373839··on Everyone should slow down AI development except for me
Privacy is a great reason, but independence is another. It’s very nice knowing that you’re going to get the same reliable product every time you call the model. Nothing is going to change unless you decide to change it.
anon373839··on We must pace the frontier
It’s absolutely laughable that he refers to METR as if they were neutral observers. They are ex-Anthropic employees and others with direct financial interests in Anthropic.
anon373839··on Everyone should slow down AI development except for me
It is costly, especially right now. I don’t think you can make a case for it on cost savings!

The throughput in a single stream is about 50 tokens/sec (a bit less for prose, a bit more for code due to speculative draft acceptance rates) and about 2,000 tokens/sec for prefill. Both numbers are flat and stable as context accumulates. That’s what finally tilted me away from the Mac Studio despite its much superior memory bandwidth.

I think these numbers may improve because the model is pretty new and optimizations aren’t done.

anon373839··on Everyone should slow down AI development except for me
> If someone put in frontier AI models from like .... last june I guess? in a box and let me run it with "decent" token throughput I would be happy.

You can have that! Qwen 3.8 Flash-Next is ~Opus 4.6 and runs nicely on a DGX Spark. And that’s just an architecture preview. The Qwen 4 family is expected to arrive this fall.

anon373839··on Nvidia is the central bank of AI
No, it’s not the active parameters. Qwen 3.8 Flash has 6B active and it smokes both models.
anon373839··on ChatGPT: "Super shady" data controls settings
https://xcancel.com/edoardocontente/status/20976073405059281...
anon373839··on Do you think it happened? Research stolen from their Codex private chats
I don’t use ChatGPT anymore, but when I did, the training opt-out setting was frequently silently resetting itself to off.
anon373839··on We have a year to fix security everywhere
> GLM 5.3-flash released last week, and that means Project Glasswing and Daybreak are running out of time. Cheap models capable of dangerous hacking are now available to anyone, without the normal safeguards for refusing malicious actions. We need to fix vulnerabilities across the industry so that we aren't caught unawares. … We can use frontier LLMs that move faster than a human to find and fix these issues in the time we have left.

I think the author has this backwards. In the timeline I’ve been living in, it’s the frontier models that have been carrying out attacks on third parties, and Chinese open source models doing the defending! During the Huggingface incident, HF was denied use of frontier models to fend off the intrusion, but was fortunately able to turn to its self-hosted instance of GLM-5.2. And it did the job.

anon373839··on Corporate America is getting hooked on open-source AI
Claude is different from S3. AWS doesn’t need to rifle through your files to stay ahead of the competition or to mine them for business ideas because the core business is overvalued and rapidly commoditizing. AI labs, on the other hand, have an incentive to exploit every last drop of data they can lay their ethically challenged hands on. And they’ve already demonstrated that they will do this, even when it involves blatant fiduciary violations (see eg, Anthropic / Figma board member scandal).
anon373839··on Artificial Analysis Intelligence Index v4.2
Hallucinations are very damaging to a model’s utility. But doesn’t the Omniscience Index focus on knowledge-based queries? To me, using LLMs for their memorized knowledge is very 2023 and suboptimal.

IMO, what really makes a model useful is its ability to process information within its context reliably and faithfully. I don’t care if it hallucinates George Washington’s favorite color, but I do care about it hallucinating the results of tool calls.

anon373839··on Corporate America is getting hooked on open-source AI
Jensen Huang?
anon373839··on Claude for Commerce Agents
Ah, I see Anthropic is back at that “saving humanity” game…
anon373839··on How concerned should we be about Astra's recurrent architecture?
Sebastian Raschka posted about this architecture:

> A lot of hype around OpenAI's Astra model here on my timeline today. Apparently, this goes back to a new article from The Information, which said Astra is a "recurrent depth or looped transformer".

> It's always interesting to read about new or different approaches (including rumors about what the closed labs may be up to), but let's debunk this a bit.

> About 2 months ago, I shared the architecture details of Nanbeige, for example, where "Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters."

> Yes, that's it. The looped transformer idea is just reusing layers in the transformer block.

> In the case of Nanbeige, the main idea is to reuse the same 22-layer stack (=transformer block) twice instead of once. So, effectively it extends the 22-layer architecture to 44 layers, but without duplicating the weights.

> In simple terms, this roughly doubles the size of the model (if we ignore the embedding and output layers for a second). But instead of requiring 2x the storage and RAM to host this model, it stays at the same size since we reuse the components. However, it's almost 2x as expensive in terms of compute, because we run the embedded text through almost 2x as many layers.

> Why? In the Nanbeige 4.2 technical report, the researchers found that two passes gave the best trade-off and retained about 75% of the token efficiency of a standard architecture. (More passes gave barely any gains but made the training much slower and much more expensive.)

> While, as far as I know, Nanbeige 4.2 is the first notable open-weight model that adopted this approach, the idea goes back to the NeurIPS paper "Mixture-of-recursions: Learning dynamic recursive depths for adaptive token-level computation". Actually, this paper proposes a mechanism that is a bit more sophisticated by adding a learned router that determines whether each token receives one, two, or more passes. So, easy tokens can exit early while harder tokens receive additional computation.

> In sum, Astra may be a really good model, but this shouldn't be about this "looped transformer aspect," which is just a tiny architectural tweak.

https://x.com/rasbt/status/2095141254958858496

Page 1 of 32Next →