HNHacker News
TopNewBestAskShowJobs

ainch

611 karma · joined May 4, 2025

Studying a DPhil in Robotics World Models for Nuclear Fusion applications

Into world models, reinforcement learning, evolutionary methods and fast ML code

submissionscomments
ainch··on A 0-click exploit chain for the Pixel 10
Did you publish this anywhere? Would love to read more.
ainch··on Show HN: Needle: We Distilled Gemini Tool Calling into a 26M Model
Apple already have this working on iOS - as discussed in this recent post.

https://unix.foo/posts/local-ai-needs-to-be-norm/

ainch··on Pyrefly v1.0.0 is here (Type Checker / Language Server for Python) [video]
I'm glad it's 1.0, but for me the most exciting thing is the tensor shape typing they mention in the blogpost. I've been looking for a good solution to this for ages, and the Pyrefly implementation seems to have everything I'd like - incl. static typing (unlike jaxtyping's runtime-only setup), and shape inference (so passing a 2,3 tensor through a 3,10 linear layer gives you a 2,10 output tensor).
ainch··on Kraftwerk's radical 1976 track
I love Kraftwerk, but contributing to anti-nuclear sentiment in Germany hasn't been a major success. If only more European countries had followed the French example and developed substantial nuclear fleets.
ainch··on A recent experience with ChatGPT 5.5 Pro
Agreed, Gemini is clearly a capable model, but the tool use is lagging behind the other two. Ironically it regularly gets things wrong (ie. the current version of some software) because of an unwillingness to use web search.
ainch··on Mojo 1.0 Beta
They've said they'll open source the compiler alongside the 1.0 release.
ainch··on Mojo 1.0 Beta
In their original pitch that was definitely part of it: take Python code, add type hints, get a big speedup. As they've built it out it seems to have diverged.
ainch··on Mojo 1.0 Beta
As someone in ML who's interested in performance, I'm keen for Mojo to succeed - especially the prospect of mixing GPU and CPU code in the same language. But I do wonder if the changes they're making will dissuade Python devs. The last time I booted it up, I tried to do some basic string manipulation just to test stuff out, but spent an hour puzzling out why `var x = 'hello'; print(x[3])` didn't work, and neither did `len(x)` (turns out they'd opted for more specific byte-vs-codepoint representations, but the docs contradicted the actual implementation).

Hopefully they get Mojo to a good place for more general ML, but at the moment it still feels quite limited - they've actually deprecated some of the nice builtins they had for Tensors etc... For now I'll stick with JAX and check in periodically, fingers crossed.

ainch··on AlphaEvolve: Gemini-powered coding agent scaling impact across fields
The AI CEOs love to pontificate about AI curing cancer, but it seems like DeepMind is the only one actively working on these research problems, while OpenAI/Anthropic largely chase enterprise/coding revenue.
ainch··on Uber wants to turn its drivers into a sensor grid for self-driving companies
They're two different models - you can use the world model to train (or test like Wayve) a different car-driving model.

The world model is basically intended as a more true-to-life simulator.

ainch··on Softmax, can you derive the Jacobian? And should you care?
And importantly it's got nice properties like being differentiable and monotonic, unlike eg. taking |x|.
ainch··on How to Build the Future: Demis Hassabis [video]
I'm generally suspicious of IQ but still - 2 standard deviations above average would include about 2.5% of the population. Hassabis is in a significantly more exclusive slice.

His parents weren't particularly wealthy. More likely: he is exceptionally intelligent, hardworking, visionary, and grew up in an environment which fostered those attributes in a precocious child.

ainch··on There Will Be a Scientific Theory of Deep Learning
On tabular datasets less than ~250k samples, tabular foundation models now outperform boosting. Of course it remains to be seen how they'll scale to significantly larger datasets as the models improve.

https://huggingface.co/spaces/TabArena/leaderboard

ainch··on There Will Be a Scientific Theory of Deep Learning
Check out Andrew Gordon Wilson's excellent paper "Deep Learning is Not so Mysterious or Different" for a discussion of the ways in which existing learning theory does and doesn't work neural nets.

https://arxiv.org/pdf/2503.02113

ainch··on Claude Opus 4.7
I would expect to see a significant wall clock improvement if that was the case - Meta's Coconut paper was ~3x faster than tokenspace chain-of-thought because latents contain a lot more information than individual tokens.

Separately, I think Anthropic are probably the least likely of the big 3 to release a model that uses latent-space reasoning, because it's a clear step down in the ability to audit CoT. There has even been some discussion that they accidentally "exposed" the Mythos CoT to RL [0] - I don't see how you would apply a reward function to latent space reasoning tokens.

[0]: https://www.lesswrong.com/posts/K8FxfK9GmJfiAhgcT/anthropic-...

ainch··on Exploiting the most prominent AI agent benchmarks
I think that should be true, but doesn't hold up in practice.

I work with a good editor from a respected political outlet. I've tried hard to get current models to match his style: filling the context with previous stories, classic style guides and endless references to Strunk & White. The LLM always ends up writing something filtered through tropes, so I inevitably have to edit quite heavily, before my editor takes another pass.

It feels like LLMs have a layperson's view of writing and editing. They believe it's about tweaking sentence structure or switching in a synonym, rather than thinking hard about what you want to say, and what is worth saying.

I also don't think LLMs' writing capabilities have improved much over the last year or so, whereas coding has come on leaps and bounds. Given that good writing is a matter of taste which is beyond the direct expertise of most AI researchers (unlike coding), I doubt they'll improve much in the near future.

ainch··on How Do You Find an Illegal Image Without Looking at It?
Germany has an anonymous support programme for people who feel paedophilic urges but don't wish to offend. I believe they've used that network for research, but I think it's probably quite a limited, and potentially biased, sample.
ainch··on Maine is about to become the first state to ban major new data centers
Oh sure, I see what you mean - thanks for clarifying. On top of your point, it's true that CO2 has a prolonged impact on global temperature even after it's been 'removed' from the atmosphere, so even once solar pays back the original carbon investment its impact lingers for a while.

I guess at a certain point you're getting at a more fundamental question about the value of AI (plus technology and everything else) - what level of environmental tradeoff is acceptable? One thing I slightly lament about the discourse is that tradeoff is widely discussed in the case of AI, but not in the context of stuff we do. I suspect most people aren't aware that the water use associated with eating a burger dwarves a year of ChatGPT, that a long-haul flight wipes out the emissions savings of a couple years' veganism, or that renewables have their own impacts, like the demolition of Chile for copper.

ainch··on Maine is about to become the first state to ban major new data centers
Sorry, I'm not picking up on the connection - could you expand? Do you think they should also pay for offsets alongside developing energy infrastructure?
ainch··on Maine is about to become the first state to ban major new data centers
Carbon offsets are a sham, but you could just require them to directly pay for the actual energy infrastructure required. If you need 1GW of electricity, develop 1GW of solar.
ainch··on ML promises to be profoundly weird
Transformers do have a fixed input/output size though - that's what a context window is. It's just that, via scaling and algorithmic improvements, the length of usable context windows has increased to the point that they're much less of a bottleneck.

I think your points around parallelisation and the flexibility of quadratic attention are spot-on though.

ainch··on Project Glasswing: Securing critical software for the AI era
This opens up an interesting new avenue for corporate FOMO. What if you don't partner with Anthropic, miss out on access to their shiny new cybersec model, and then fall prey to a vuln that the model would have caught?
ainch··on Sam Altman may control our future – can he be trusted?
Great piece. And a good excuse to read up on the use of diaeresis in English (eg. coördination, reëlection) to distinguish repeated vowels - I hadn't seen the New Yorker's usage before.
ainch··on Artemis II Launch Day Updates
Randall Munroe of xkcd? I like his work but I'm not sure I'd call him a philosopher...
ainch··on AI for American-produced cement and concrete
Sure, it's clearly marketing. I think a private company pursuing marketing via open research with open source code (including datasets) is a good trade. A hypey blogpost + research is better than no blogpost and no research.
ainch··on AI for American-produced cement and concrete
The Gaussian Processes underpinning this work are hardly a product of the 'AI Hype Machine' - they've been around for decades, have strong statistical underpinnings, and are being widely explored for experimental design across many disciplines. Reflexive and poorly-informed backlash to any variety of machine learning is no more productive than blindly hyping up LLMs.
ainch··on Take better notes, by hand
A sidenote along these lines - I've recently done an MSc, and found that the default approach to lectures is now to present slide decks. One of the profs, however, delivers a more traditional lecture, writing everything on a blackboard. I've found the second style far more effective, largely because writing caps the rate at which information can be conveyed. Because slides have no such bottleneck, I've found they're often misused and overladen with information which is skipped over too quickly.
ainch··on How the AI Bubble Bursts
It's definitely true that they've increased their revenue rapidly. But at the same time the 'scaling laws' that the labs were first built around require exponentially-scaling cost (10x flops for a fixed reduction in training loss).

If anything, a better look at the economics is a reason to look forward to one of them IPO-ing. I suspect the labs probably could cut R&D and turn a profit, but that might only work for one generation, until they get superseded by the competition.

ainch··on How the AI Bubble Bursts
Do you have any evidence that inference revenue is growing faster than training costs? RLVR is significantly less compute-efficient than token-prediction pretraining - especially as labs are trying to train models to achieve agentic tasks which take tens of minutes per rollout.
ainch··on Show HN: Veil – Dark mode PDFs without destroying images, runs in the browser
As a PhD student doing my fair share of midnight paper-reading I think I'm the exact target market - thank you for sharing!
← PreviousPage 4 of 7Next →