HNHacker News
TopNewBestAskShowJobs

ainch

611 karma · joined May 4, 2025

Studying a DPhil in Robotics World Models for Nuclear Fusion applications

Into world models, reinforcement learning, evolutionary methods and fast ML code

submissionscomments
ainch··on What Emily Bender meant by "stochastic parrots"
I'm afraid the precise connection you're making isn't totally obvious to me.

As far as prediction - I mean sure the cortex and LLMs do prediction, but then so can RNNs or diffusion models or any other generative model. Really any ML architecture is learning to compress its environment in pursuit of modelling. More broadly, the predictive brain model would suggest that all of the brain, not just the neocortex, is dedicated to prediction. What would you say makes LLMs similar to the neocortex, rather than the basal ganglia or Broca's area?

Similarly, if you agree with the Manifold Hypothesis, then all machine learning models operate on manifolds. I agree it's an exciting thought, but then I don't know what would distinguish an LLM from a VAE or SVM in terms of operating over a low-dimensional manifold embedded in high dimensional spaces - maybe just scale?

ainch··on What Emily Bender meant by "stochastic parrots"
Could you provide more detail? My understanding is that the neocortex is predominantly focused on forwards simulation, which seems distinct to how transformers operate.
ainch··on Box3D, an open source 3D physics engine
As an ML researcher, I know box2d because it underpins many of the standard reinforcement learning environments (in OpenAI Gym) that we use to benchmark methods, like Lunar Lander or Car Racing: https://gymnasium.farama.org/environments/box2d/car_racing/

Thanks to Erin for such a useful piece of software!

ainch··on Un-0: Generating Images with Coupled Oscillators
This method is cool and the post explains it well. It would, however, be good to get more detail on the energy efficiency they flag as their motivation: is this model actually more energy efficient than the comparators they highlight?
ainch··on Qualcomm to Acquire Modular
They've said that Mojo is still on track to be open-sourced this year, post-acquisition.
ainch··on Sakana Fugu
Indeed. The world models research many labs are now chasing was to some degree ignited by David Ha and Schmidhuber's 2018 paper.

More broadly, Sakana is pursing a refreshingly distinct research path, with their focus on evolutionary methods, biological intelligence (e.g. continuous thought machines) and open publication.

ainch··on The Commodore Callback 8020 smart flip phone
Basically all the images and videos on their website are AI-generated. It doesn't inspire much confidence.
ainch··on Gram Newton-Schulz: A Fast, Hardware-Aware Newton-Schulz Algorithm for Muon
Tri Dao's lab must have saved countless watts with FlashAttention. Great to see them continuing to open-source massive efficiency gains.
ainch··on Anthropic apologizes for invisible Claude Fable guardrails
Here's one that was flagged for me: a question about a niche Reinforcement Learning paper from 2012

I've been reading the option-option model paper by David Silver. It appears that they achieved quite an effective result. Why hasn't there been more work on it since?

ainch··on πFS
I agree it's an oversimplification. The example I think of is something like Newton's law of gravitation vs Ptolemaic epicycles: one simple explanation replaced many layers of tweaks.

It's also a relevant example for AI - one paper tested the ability of Transformers to model planetary orbits: unlike Newton's Law, the implicit forces they learn are nonsense.

https://arxiv.org/pdf/2507.06952

ainch··on Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
Anthropic's claim was that Deepseek collected ~150k conversations.

https://www.anthropic.com/news/detecting-and-preventing-dist...

I think the extent of distillation by Deepseek specifically is overstated. For comparison, Minimax collected over 13m 'exchanges', which starts to sound a lot more like large-scale distillation.

ainch··on πFS
In some sense, science is the most extreme form of compression - Newtonian mechanics explains an incredible number of phenomena in a few lines of text.
ainch··on Claude Fable 5
AISI did also say that GPT-5.5, which has been public for months, scores basically the same as Mythos on their cybersec evaluation. But there wasn't as much media about about that for some reason.

https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5...

ainch··on Claude Fable 5
Token prices have increased, but it's not really the whole story at this point, given some models will use far more tokens to complete a task than others. One of the charts in Anthropic's blog posts shows Fable at 'low' reasoning achieving better results for less money than Opus on 'high'.
ainch··on Apple reveals new AI architecture built around Google Gemini models
They say it's because the EU's DMA would require them to open up device data to third-party assistants, and they'd no longer be able to guarantee user privacy.

https://www.apple.com/newsroom/2026/06/due-to-dma-siri-ai-de...

ainch··on AI is slowing down
As WIRED reported[0], despite constantly writing about how an AI collapse is just about to come, Zitron privately does PR for AI firms on the side. The man is an obvious hack, and it's disappointing that he has become one of the mainstream faces of AI skepticism.

[0]: https://www.wired.com/story/ai-pr-ed-zitron-profile/

ainch··on Ask HN: What was your "oh shit" moment with GenAI?
I'd spent 6 hours solving a gnarly RL problem (mathematically solving divergence of off-policy TD-Lambda for any value of lambda or behaviour policy).

As a punt I gave it to o3 (remember LLMs were 'bad at maths') - after 15 minutes it returned with the answer that had taken me hours.

ainch··on Artificial intelligence is not conscious – Ted Chiang
There's a growing body of evidence that most of what the brain does is constantly predicting the world around us - look into the predictive brain hypothesis if you're interested.

As Ilya Sutskever has pointed out, if you read a mystery novel up until the reveal of the culprit, and then fill in "The killer was _____", don't you need to understand the novel to accurately predict the next word?

ainch··on Nvidia Cosmos 3
You can fine-tune it so, given an image and a task description, it generates a corresponding set of actions.
ainch··on Dune's Butlerian Jihad and the Future of AI
Deep Blue's strength was leveraging massive compute to execute a task-specific, human-written algorithm. The problems which LLMs are tackling don't elude mathematicians because they require too much number-crunching, but because they demand creative problem solving. The latter seems more profound, even if it doesn't imply sentience.
ainch··on Agora-1: The Multi-Agent World Model
Very cool work, the learned world state is a smart way of getting consistent generation across all the views (and not having the map vanish when you 180 like some other models). Multi-agent is such an interesting field, because it's clear that humanity benefits from distributed intelligence, but I don't think MARL has really had a big breakthrough like AlphaGo or RLVR for single-agent RL.

Two thoughts about where this could go: first, the internal world state would need to be learned to transfer to real-life robotics, since you can't query the internals of a game engine in training. Second, an enormous challenge for many of these world models is going to be truly unbounded environmental interactivity - Agora is still mostly about a few agents interacting in a static environment. Learning interaction will be hard, because the interactions in games are intentionally added in, by hand. But we (human learners) acquire a strong model for environental interaction very efficiently, which is part of what helps us generalise so effectively.

ainch··on Futhark by example (2020)
Thanks for the reply - you're clearly much more experienced with the internals here, but I believe we're still talking at cross-purposes. I believe you're talking about compilers like jax.jit or torch.compile performing symbolic shape inference. I'm talking about the ergonomics of tracking shape information while writing Python code that calls these libraries. I don't use Torch much, so I'll just comment on the Jax side.

> Jax is explicitly mentioned in your pyrefly link as having a parallel (but slightly weaker) system

Jaxtyping is limited to runtime-only checks (which might as well be assert statements), and doesn't infer shapes based on operations. I'm interested in Pyrefly because I've run into the limitations of Jaxtyping in my own usage.

> Jax is built on stablehlo which uses shape dialect, which is part of the compiler (and therefore statically known).

It's true that JAX does shape inference when it compiles down to HLO - but that isn't available to the Python typing system. The Pyrefly development is addressing that, so you get static analysis before even running anything, or without having to add eval_shape calls all over your codebase. I think that's helpful, and will catch bugs. When I say Jax does inference at runtime, I mean that you have to run for the jit compiler to kick in - you don't get feedback as you edit.

> the fact that people do not know how to use the tools does not mean the tools are lacking... almost everyone that is employed to work with these tools is aware of these features and therefore eschews those kinds of comment strings.

The examples I took are from Andrej Karpathy and Noam Shazeer - maybe the disconnect is that they're more on the research side. Perhaps only unsophisticated users rely on these hacks - but as one such user I'm very excited that Pyrefly is addressing a problem I have. I suspect part of the misunderstanding that's evolved here is that these tools serve audiences with different needs.

ainch··on Every AI Subscription Is a Ticking Time Bomb for Enterprise
At the same time, the training paradigm being scaled, Reinforcement Learning, is significantly less data-efficient than next-token prediction. You basically need to run an agent for minutes (or longer if you want good long-horizon performance), only to give it a binary pass/fail - one bit of information.

Inference compute is definitely scaling fast, but to scale RL, training and R&D compute also needs to scale hard. I don't think it's obvious that inference will overtake R&D/training, unless there's a reputable source that states that.

ainch··on Every AI Subscription Is a Ticking Time Bomb for Enterprise
The price for a given level of capability will fall, but the frontier has recently been getting more expensive. If you compare GPT-5 to GPT-5.5 on the Artificial Analysis benchmark, it's ~4x more expensive, but achieves a higher score. Claude 4.7 is also more expensive than predecessors because of a tokenizer change.

As the AI labs become more reliant on enterprise adoption, it makes sense to push capabilities at a cost that makes sense for businesses. Even if it prices out consumers or hobbyists.

ainch··on AI subscriptions are a ticking time bomb for enterprise
Tokens can be sold at profit, but 70% of compute expenditure goes to R&D and model training[0]. Inference needs to cover all of that as well as being profitable in a vacuum.

[0] https://epoch.ai/data-insights/openai-compute-spend

ainch··on Futhark by example (2020)
I use jaxtyping as documentation, but the fact it can only be used for runtime checking (in a slightly clunky manner) and can't infer shapes based on ops really limits its utility imo.
ainch··on Futhark by example (2020)
I think we're talking at cross purposes. The reason I'm excited about the Pyrefly work is that it leverages the type system to infer array shapes statically, which makes it simpler to reason about them when you're writing the code and catch bugs. The fact that people have developed these janky approaches to shape tracking suggests that there's a gap to be filled.

Jax and Torch don't do that statically. They obviously have to do it at runtime, but that doesn't address this particular issue. I mention Eigen because array shape hinting is generally useful for any linalg library, not purely for ML applications.

ainch··on Futhark by example (2020)
I didn't know that, thanks for sharing. It makes sense, but then it also makes me wonder why none of the deep learning libraries (Torch, Jax/NNX, Eigen etc...) make this information available. Instead, ML people all have their own schemes for tracking shape information, like commenting '# (b, n, t)' on every line, or suffixing shapes to variable names - and in my experience it's a common source of bugs.
ainch··on Futhark by example (2020)
The Pyrefly type checker is starting to work on this kind of shape hinting - so far it only works on Torch but I believe the plan is for it to work with other array packages (eg. JAX, NumPy)

https://pyrefly.org/en/docs/tensor-shapes/#how-it-works

ainch··on The bird eye was pushed to an evolutionary extreme
I've only discovered Quanta this year but it's quickly become my favourite publication. The focus on quality articles across science, and especially pure maths, feels very unique. I don't know whether it's profitable or reliant on the Simons Foundation funding - but hopefully it's a sustainable business model that will stick around.
← PreviousPage 3 of 7Next →