610 karma · joined May 4, 2025
Into world models, reinforcement learning, evolutionary methods and fast ML code
I could be misreading this, but I hope the author doesn't think GANs are used in LLMs. They are cool, though.
I used to work at Tortoise (who now own/operate The Observer) and our motto was 'Slow News'. We operated on a pretty transparent membership model, but I got the impression that it was difficult to make pure membership work financially, leading to a reliance on other revenue streams like podcasting and business partnerships over time.
On your questions - I've spoken to a number of execs and seniors behind closed doors but nothing public I can point to. Anecdotally, I've spoken to senior leaders at banks spending billions of tokens on one-off tasks like prepping execs for earnings calls or piloting end-to-end agent workflows for specific use cases (but mostly piecemeal/one-off).
On Fable, financial analysts I know are using it to produce research docs, models and decks - I hear that it's a big improvement for these tasks. This lot have been blindsided by the spend growth [2], the same as for coders in enterprise (e.g. Uber blowing annual budget in 4 months [3]), so I do think there's appetite and budget for a capable, cheaper open model - but, due to the price, Kimi does not obviously fill that role the way it might for coding. That said, I still largely agree with you on share - where coding has seen a broad deployment across software development, most of the office work stuff is still fairly piecemeal and certainly lower compute-spend.
I think it's fair to say I could've focussed on coding more rather than taking AA's benchmark distribution as representative - perhaps a more balanced title would be "Kimi K3 is not cheap across the board"? I guess there's also some ambiguity about what 'cheap' means - as I said elsewhere in this thread, I think when some people talk about the price of Chinese models, they imagine Deepseek competing with o1 for 1/20th of the price. Even though it is better priced for coding, Kimi isn't Deepseek-level cheap.
I do, however, think you could debate whether coding will remain at >50% total token usage going forwards - big enterprises are hunting for ways to get value out of LLMs, and the labs are investing a correspondingly large amount in generating demonstrations and RL environments to get the models up to par. At the end of the day, programmers make up ~5% of all white collar work. Of course, it's also possible that Chinese labs will shift focus to white collar applications now they've demonstrated a lead on coding cost efficiency, so, I mean who knows - it'll be interesting to get some detail when Anthropic IPOs.
Sorry for the long reply! Appreciate it's quite meandering...
[1] https://cdn.openai.com/pdf/5d1e1489-21c0-43e4-9d42-f87efdbf0...
[2] https://www.reuters.com/business/finance/australias-cba-flag...
[3] https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-c...
General office work is one of the big frontiers the labs are pushing on, and it's part of how they're justifying the value proposition to enterprise customers. It's also accounts for a big portion of the spend on RL; tasks/environments designed to train agents to navigate Slack or Salesforce. If you're Anthropic pitching Claude to a bank (taking an example I'm familiar with), coding probably accounts for ~20% tops of the workforce, and it doesn't drive direct revenues. The 'agentic coding bump', but for all your analysts, traders, and wealth managers, would be a much more attractive prospect.
I don't disagree that coding is the most successful use case so far (and probably more relevant to a HN audience). But I think the future of the labs is also contingent on them making progress on more general white collar work. I suspect that's why the Opus 5 release blog lists 3 coding benchmarks (FrontierBench, DeepSWE and FrontierCode) to 3 or 4 more general ones applicable to office work - depending on how you slice it (GDPVal, AutomationBench, Legal Agent Benchmark, BrowseComp).
I don't mean to imply that Kimi is not at all cheaper than U.S frontier models. I more wrote this because I believe - since Chinese LLMs entered the public consciousness via DeepSeek R1, which was genuinely ~20x cheaper than o1 - there's a bit of a halo effect around Chinese models which causes people to overestimate the scale of the discount. And relative to that price anchor, Kimi is less extraordinarily cheap.
At the moment Kimi is ~10% cheaper than GPT-5.6 on the AA benchmark, and as you say that could go down to 20-30% cheaper (although I don't know how inference provider discounts play out on real world usage once you account for quantisation etc...). I'm not trying to suggest that that's nothing, but I do think some of the people driving the Chinese AI discourse would have a harder time pitching their conclusions if they were saying "this new Chinese model is 10% cheaper on some tasks, and it might get another 20% cheaper in the future".
"Despite being a highly competitive model overall, K3 nonetheless exhibits a noticeable gap in user experience compared with Claude Fable 5 and GPT 5.6 Sol."
Even if it is a marketing ploy, I could see this stuff backfiring catastrophically - after all they have just illegally hacked a 3rd party via a model they can't control properly. Any serious person in government (US or otherwise) will look at this and say "these guys have no idea what they're doing"
I don't think JEPA is necessarily the solution (although I'm bullish on world models), but I don't understand why people feel so infuriated about LeCun. The reason he is famous today is because he spent years taking a contrarian stance, working on neural networks when they were seen as dead and buried. He was eventually proven right. Nowadays he holds strong views which contradict the LLM zeitgeist --- what's shocking about that?
I think it should be obvious that he understands that LLMs are trained via next-token prediction.
[0]: https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-...
In flax nnx, what's the idiomatic way to store state on a Module. For example, if I'm handling the carry manually for an nnx.RNN.
Or one asking about a checkpointing package: How do I restore one of the orbax checkpoints into NNX from this script?
I also got flagged for asking about syntax highlighting in the Helix editor.It's a shame - I like Fable for writing tasks over ChatGPT and I do believe Anthropic is a more ethical outfit than OpenAI. But with the safeguards (and Fable access expiring in a few days) there's no reason to pay for draconian guardrails and harsh rate limits.
And in the case the previous poster describes, the other model doesn't generate datasets, it generates environments which the next generation interact with to learn from.
Agreed that the neocortex uses fewer layers because of looping - I also suspect it's partly because neurons are more complex than the neurons used in ANNs, so in principle they should be capable of more sophisticated computation in a single forward pass (especially considering that handling multiple neurotransmitters could mean superposed functions).
The point about there being one correct way to model something does seem backed up by the Platonic Representation hypothesis (https://arxiv.org/pdf/2405.07987). I've even seem some work that shows you can find a bijective map between the latent spaces of different transformers (https://arxiv.org/pdf/2505.12540).
Your point about Finsler spaces is fascinating - I hadn't come across the term before. It'd be interesting to see if there's work that specifically allows latents to exhibit that kind of directional metric behaviour, and whether that improves generalisation or something.
On the entire brain being dedicated to prediction, that's more of a high-level comment. I was thinking of work like Andy Clark on predictive processing, which suggests that even regions of the brain which we think of as receptive (eg. the visual cortex) may be implementing a generative/predictive model which is corrected by sensory input.
As far as prediction - I mean sure the cortex and LLMs do prediction, but then so can RNNs or diffusion models or any other generative model. Really any ML architecture is learning to compress its environment in pursuit of modelling. More broadly, the predictive brain model would suggest that all of the brain, not just the neocortex, is dedicated to prediction. What would you say makes LLMs similar to the neocortex, rather than the basal ganglia or Broca's area?
Similarly, if you agree with the Manifold Hypothesis, then all machine learning models operate on manifolds. I agree it's an exciting thought, but then I don't know what would distinguish an LLM from a VAE or SVM in terms of operating over a low-dimensional manifold embedded in high dimensional spaces - maybe just scale?