Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)
4,781 karma · joined May 21, 2023
Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)
It's a weak analogy. Ordinary comms infrastructure providers actually have pretty robust policies, technical, and contractual measures in place that restrict employee access to customer communications. In contrast, in the wild west of generative AI, companies actively monitor session data for the content itself, in order to exploit it for their own business purposes. There is zero expectation of privacy.
So I don't share your expectation that precedent will uphold the use of consumer-tier services (in their current form) for handling privileged material.
Chatting with Claude breaks confidentiality: chats with Claude are subject to arbitrary inspection by Anthropic employees, not to mention the issue of model training.
You can use self-hosted LLMs without breaking privilege. And funny enough, law firms like Latham & Watkins are now buying Nvidia GPU clusters for this purpose.
I agree with this. California’s climate and culture will keep many a billionaire within tax nexus reach of the state. Of course, this isn’t a strategy every locale can pursue, but I don’t see a reason for California not to exploit its advantages.
Is there really a market dynamic in rent pricing anymore? I thought that algorithmic collusion had eliminated the need for landlords to compete on price.
Another failure mode you may see is inordinately long CoT. Properly served, the model is good at calibrating its CoT length to the difficulty of the immediate task.
That's really interesting. I like Pydantic AI a lot and wondered why all of the harnesses seem to be written in Javascript instead of it. What do you use it for headless, though? I haven't tried Hermes or similar yet, so don't have a handle on what you do with them.
Renting tokens from open model providers is cheaper but it incurs the same issues: unexpected changes in model quality, inconsistent speeds, service outages.
The throughput in a single stream is about 50 tokens/sec (a bit less for prose, a bit more for code due to speculative draft acceptance rates) and about 2,000 tokens/sec for prefill. Both numbers are flat and stable as context accumulates. That’s what finally tilted me away from the Mac Studio despite its much superior memory bandwidth.
I think these numbers may improve because the model is pretty new and optimizations aren’t done.
You can have that! Qwen 3.8 Flash-Next is ~Opus 4.6 and runs nicely on a DGX Spark. And that’s just an architecture preview. The Qwen 4 family is expected to arrive this fall.
I think the author has this backwards. In the timeline I’ve been living in, it’s the frontier models that have been carrying out attacks on third parties, and Chinese open source models doing the defending! During the Huggingface incident, HF was denied use of frontier models to fend off the intrusion, but was fortunately able to turn to its self-hosted instance of GLM-5.2. And it did the job.
IMO, what really makes a model useful is its ability to process information within its context reliably and faithfully. I don’t care if it hallucinates George Washington’s favorite color, but I do care about it hallucinating the results of tool calls.
> A lot of hype around OpenAI's Astra model here on my timeline today. Apparently, this goes back to a new article from The Information, which said Astra is a "recurrent depth or looped transformer".
> It's always interesting to read about new or different approaches (including rumors about what the closed labs may be up to), but let's debunk this a bit.
> About 2 months ago, I shared the architecture details of Nanbeige, for example, where "Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters."
> Yes, that's it. The looped transformer idea is just reusing layers in the transformer block.
> In the case of Nanbeige, the main idea is to reuse the same 22-layer stack (=transformer block) twice instead of once. So, effectively it extends the 22-layer architecture to 44 layers, but without duplicating the weights.
> In simple terms, this roughly doubles the size of the model (if we ignore the embedding and output layers for a second). But instead of requiring 2x the storage and RAM to host this model, it stays at the same size since we reuse the components. However, it's almost 2x as expensive in terms of compute, because we run the embedded text through almost 2x as many layers.
> Why? In the Nanbeige 4.2 technical report, the researchers found that two passes gave the best trade-off and retained about 75% of the token efficiency of a standard architecture. (More passes gave barely any gains but made the training much slower and much more expensive.)
> While, as far as I know, Nanbeige 4.2 is the first notable open-weight model that adopted this approach, the idea goes back to the NeurIPS paper "Mixture-of-recursions: Learning dynamic recursive depths for adaptive token-level computation". Actually, this paper proposes a mechanism that is a bit more sophisticated by adding a learned router that determines whether each token receives one, two, or more passes. So, easy tokens can exit early while harder tokens receive additional computation.
> In sum, Astra may be a really good model, but this shouldn't be about this "looped transformer aspect," which is just a tiny architectural tweak.