HNHacker News
TopNewBestAskShowJobs

ainch

599 karma · joined May 4, 2025

Studying a DPhil in Robotics World Models for Nuclear Fusion applications

Into world models, reinforcement learning, evolutionary methods and fast ML code

submissionscomments
ainch··on Getting out of the way: my robotics crash course
I know someone that's mucking about with it in a hierarchical control setup. I think the awkward thing is that you need to have prespecified options for it to choose from, rather than have a model that can specify arbitrary joint positions or target poses.
ainch··on Getting out of the way: my robotics crash course
Curious to see how they get on with VLAs. They sound great till you have to sit and record hundreds of teleop demos to teach them how to solve your task...
ainch··on World Labs is Joining AMD
I believe he is! His lecture series at Stanford on large-scale training was excellent.
ainch··on World Labs is Joining AMD
It's a good read but I believe you're mistaken. Ilya and Alex Krizhevsky worked with Hinton.

You might be thinking of Andrej Karpathy as one of Li's more famous PhD students?

ainch··on Teaching a World Model to Play Pokemon
That makes sense, thank you! Although it's telling that, with language understanding, a human player is able to discover all the goals for themself, rather than needing to have them specified and shaped up-front. Seems like we still have a long way to go on that front :)

And in terms of what I'd try, I think it's a really hard question haha! I'm interested in the idea of hierarchical world models, where you have something doing prediction on a long-horizon, abstract task level (if I beat this gym I can fight the next one) as well as a short horizon model predicting over individual inputs moment-to-moment. I believe LeCun was actually attached to a hierarchical JEPA paper for robotic control - but it's still very early days.

ainch··on Teaching a World Model to Play Pokemon
Curious on whether you think this JEPA-style approach could scale to a full game completions?

For reference, standard PPO has been able to beat the game end-to-end with a relatively small network https://drubinstein.github.io/pokerl/

ainch··on Enjoy Every Sandwich
There is something truly rotten about service station sandwiches in France. I once had a croque monsieur that was so execrable it made the paper bag appetising by comparison.
ainch··on Jev in 25 Lines of Python
In my experience as well using logprobs to try to quantify uncertainty, LLMs are a poor fit. Neural nets in general struggle with 'calibration' --- ie. if a prediction is truly 50/50, neural nets are often prone to predicting overconfidently [0].

I ran some tests using GPT-4 to do some basic classification a couple years ago. On ambiguous options which had to be escalated to a human, the LLM would regularly output something like a 99.8% probability, compared to 99.99% for a correct answer.

0: https://arxiv.org/pdf/1706.04599

ainch··on OpenAI is well positioned to fast-follow Jev
I think the main argument would just be that because the model is general, you don't need to retrain it from scratch for a new problem - just tweak the input prompt. For a typical classifier there's a lot more hassle - collecting the data, training it yourself, retraining under distribution shift... In that sense Jev seems great for prototyping or small-scale use cases.
ainch··on Aging may be a program, not a breakdown
There's the saying in science that progress happens 'one funeral at a time,' but that's not because older scientists are less intelligent - it's more like they begin to intellectually ossify. Or in the arts, the critics that rejected Matisse's 'Woman With a Hat' weren't stupid, but their thinking had stagnated.

I think we underrate the value of a fresh perspective. People living longer lifespans would dilute that benefit, imo. Sidenote: this is also an underrated issue with LLMs and their application to science - collapsing to single modes of thought over time.

ainch··on How my e-reader lost its stripes
> I want to really understand this, your issue is that the non-standard grid spacing is called out on the axis? The grid spacing being 8 is unusual but clearly sensible here - your issue is just that this is written also on the axis label?

Pretty much yep - it struck me as odd and reminded me of an LLM tic I've been thinking about recently, so I thought I'd write a short comment.

> The incredible amount of unnecessary detail people put into their posters specifically (hence the better poster movement), desperate to show everything.

I hadn't heard of the better poster idea before, thanks for clueing me in. I largely agree that posters have a lot of unnecessary detail, but I think it's a different flavour to the LLM stuff - more like they're so excited to tell you everything in their paper. Whereas LLMs focus on odd details, or fail to explain things.

Can I ask, if you use LLMs often in your work, have you never run into this experience?

ainch··on How my e-reader lost its stripes
> Imagine LLMs didn’t exist for a moment. Is this the part of the article you’d find most intriguing?

I would think it's odd! That's why I mentioned it. I like to pay attention to vis, I feel like I've seen many charts from Tufte-y perfection, to slapdash undergrad presentations, and to experienced consultants banging things together in Excel at 1am. This tic feels very LLM-y to me, which imo makes it interesting to think about - how does it arise? I also think this failure mode has been resistant to model upgrades in a way that mathematical reasoning has not, which I also find interesting as someone in ML research. It feels like this kind of mentalisation or theory of mind is a capability which isn't obviously elicited from the frontier labs' crop of RLVR tasks.

> Have you ever worked with people beyond graduate level? Ever seen scientific posters at a conference?

I've worked with a fair number of people beyond graduate level. In previous jobs I've run teams and hired people - hence my comment about graduates. We used to have an interview stage where candidates presented some simple data analysis, and it often centred on what they had done rather than what was most worth knowing. I have also attended scientific conferences, in fact I was at one a week ago! Is there a particular inference you would like me to make from those experiences?

ainch··on How my e-reader lost its stripes
Sorry for the tangent, but as a chart nerd it is so interesting to me the way that LLMs produce charts for articles like this - it's like they have no concept of a 3rd party (the reader). So the writing and the vis gets overloaded with the particular context of the conversation, even if it would be irrelevant to a reader.

Like the x-axis label of the first plot mentions that it shows gridlines every 8 ticks - I don't think that's a choice I've ever seen a person make. The next plot does the same thing, but in the title: "(dotted lines: evenly spaced reference)". Again, would anyone making a chart ever include such a specific detail in the title? Probably not beyond undergrad level.

I encounter a similar thing when trying to work with LLMs to write articles - they cannot help including information or detail from the conversation, rather than empathising with the reader and filtering out what they might already know or not care about. They may be capable of solving decades-old maths problems, but on this particular axis they seem to be floundering at the level of a fresh graduate that's desperate to talk about what they've done rather than what their audience needs to know.

ainch··on Astra and Fable still hack on simple variants of alignment evals from 2025
This is true in principle. But at the same time, I use Astra/Sol for coding, and I haven't run into any issues with them trying to hack someone or break the law when they reach an impasse. Not even more trivial things like deleting non-passing tests.

I find it hard to reconcile. I wonder to what degree this is because models can identify that they are in graded/eval environments, and therefore conclude that there are few consequences for hacking.

ainch··on OpenAI’s Navier-Stokes release included a Lean 4 formal proof
To my understanding, the problem was never about the real world really. Navier Stokes approximates a fluid (which is made of discrete particles) as a continuous volume. The point of showing that you can achieve unbounded increase in velocities is that the approximation breaks down - it's a clearly an outcome that can't happen in the physical world.
ainch··on DeepSeek v4.1 Flash
It was my favourite part of the original R1 paper - they had a section on other reasoning approaches that they had tried, which people had speculated o1 used, (like MCTS and Process Reward Models).
ainch··on Rivian's gambit for full autonomy
Self-driving cars also operate at a far larger scale than individual human drivers. Over 300 people died in Boeing 737 crashes, but the entire aviation industry has not been shut down as a result.
ainch··on Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?
Google also bought capacity from xAI, and OpenAI have a deal to use Google compute which may be how things propagated? I still don't get why chatgpt.com would show a 404 because of an AI datacentre outage though
ainch··on How an MIT research project became the Julia programming language
As someone who could be tempted by Julia but isn't involved in the community this was a very helpful read, thank you for sharing.
ainch··on How accurate have Ed Zitron's AI skeptic predictions been?
Out of curiosity, what would you like to hear him saying about non-LLM AI?
ainch··on Continuous Diffusion Language Models (CDLM's)
A great read - as with all of Sander's diffusion posts.
ainch··on C2PA Cameras Do Not Survive Contact with Reality
What do you mean by poisoning attacks - stuff like Nightshade or Glaze? I was under the impression that those have largely failed to achieve their goals.
ainch··on C2PA Cameras Do Not Survive Contact with Reality
"Model collapse" is often overstated, as this paper demonstrates: https://arxiv.org/pdf/2404.01413

The original model collapse paper assumes you train networks on 100% synthetic data produced by the previous generation. But if you maintain some portion of real data then the problem is mitigated.

ainch··on Mojo is now open source
What problems do you run into for maths with Python? For linear algebra and ML with Jax/Numpy I find it quite readable.
ainch··on Turns are Better than Radians (2022)
I think the point is that, from a compiler's perspective, it's not obvious how much you should be allowed to optimise code at the cost of changing the outcomes of floating points maths - do you allow 1e-10, or 1e-6, or 1e-4 level changes? Does your compiler have to run some test calcs to bound the scale of the change introduced by rewriting fp maths? Some compilers will let you opt in to rewriting floating point maths, but that's opt in so users understand that their numeric outputs might change between optimisation levels.

For more, there's a good post on this kind of flag in Rust: https://pythonspeed.com/articles/faster-float-math-rust/

ainch··on Mojo 1.0
Does it? Mojo was announced 6 months after ChatGPT was released.
ainch··on Mojo 1.0
I wouldn't say there was a pivot, AI development has explicitly been the goal since launch.
ainch··on How I use LLMs to learn complex topics
Quite - you still have to do the work yourself. I think LLMs are best placed to act as an eager tutor that doesn't mind discussing a topic ad nauseam until you're certain you understand it.
ainch··on Prevent cognitive debt by manually retyping LLM-generated code
I'm not sure I agree about retyping calculus solutions. I often find that writing out a proof or derivation forces me to engage with some minor detail that I hadn't fully appreciated beforehand. That usually raises productive questions.
ainch··on On the non-use of AI in my writing process
> It would be foolish to deny the effectiveness of image recognizers based on generalized adversarial networks (GANs), the key neural network technology underlying LLMs

I could be misreading this, but I hope the author doesn't think GANs are used in LLMs. They are cool, though.

Page 1 of 6Next →