HNHacker News
TopNewBestAskShowJobs

rsfern

1,761 karma · joined February 5, 2015

Working as a researcher in materials science (physical metallurgy specifically).

My research interests include applied machine learning for quantifying materials substructure, which we call microstructure.

submissionscomments
rsfern··on An Emacs-style visual undo tree for Pi session branching
Thanks! This seems really cool. If I’ve got it right, your UI builds and displays these diffs (and implements undo/redo) by parsing the edit tool calls?

If so that seems really nice, one of the things I don’t like with coding agents is having to defensively commit changes to roll back if the agent goes off the rails. I’m not sure if that’s a problem with my workflow, but your tool seems great for exploratory stuff

How does it change the way you personally use pi? How git aware is it, do you have to commit at the end of a branching session, and does it handle manual edits?

rsfern··on An Emacs-style visual undo tree for Pi session branching
Interesting project idea! The link seems to be 404, is the repo still private?
rsfern··on Navier-Stokes – Tristan Buckmaster [pdf]
Why would mining chat transcripts for ideas be untenable? They already run a summarization model to auto-title the chat, and to run a bunch of safety filters, and presumably to score transcript quality for A/B testing and to collect more finetuning data. Seems like evaluating for open research questions and approaches would be pretty trivial extension of this, after all it’s kind of their core business model
rsfern··on Trump Ripped by Reporters Used as Decoys in Secret Escape
The bit you quoted doesn’t capture why the reporters are upset:

> White House journalists are outraged that a threat credible enough to force Donald Trump to escape from Air Force One using an airport catering truck wasn’t relayed to them.

rsfern··on U.S. Department of Energy Launches the Genesis Open Models Initiative
Right, I did specifically say that most of the DOE scientists are contractors, but I concede the phrase “government scientist” is a bit ambiguous. I appreciate the extra detail you added. I think the distinction between political appointee and scientist/researcher stands.

As an added complication, some of the DOE labs do have civil servant scientists, for example National Energy Technology Lab and National Renewable Energy Lab are like 50/50 civil servants and contractors. And most of the funding arm of DOE are career civil servants. LANL, Sandia, Livermore, Argonne are all staffed by contractors

rsfern··on U.S. Department of Energy Launches the Genesis Open Models Initiative
Let’s distinguish a bit. There are political appointees (Trump’s government employees as you say) who are mostly upper management, and there are career civil servants (all the government scientists are under this category) who have a strong culture of apolitical dedication to the mission of their agency and to the American people and Constitution, regardless of who the current president is. And in the DOE labs in particular most (not all) of the scientists are actually employed as government contractors, but they have a similar non-partisan ethos.

That doesn’t necessarily mean there’s no need to be concerned with potential impact of policy and priority changes from the administration, but it does temper the threat model because the government employees you’re considering trusting have given oaths of office to protect and defend the Constitution.

rsfern··on Position: LLMs Can't Jump
I found this paper really thought provoking, but I think the conclusion of “world models are the solution” leaves something to be desired. People are already equipping agentic systems with physical simulation tools and exploring action-conditioned world models. This is cool because you can change the rules of the simulation and observe what happens, but it doesn’t address the core question of what to change the rules to, or even what the goal should be in the first place.
rsfern··on The session you cannot take with you
Or they could store the reading traces and validate the user hasn’t edited them server-side? They could sign reasoning traces so they can’t be counterfeited?
rsfern··on Marimo now runs in PyCharm
The cell DAG enforces that there’s no implicit state, which reduces cognitive load for me a lot and provides some pressure to abstract experimental code into functions. In Jupyter this is left to user discipline and restart-and-run-all workflow

The reactive components are also really nice for interactive plotting and exploratory data analysis. You can do this in Jupyter but it feels less seamless somehow. Interactive marimo workflow feels like streamlit but in a notebook interface

One thing I miss from Pluto.jl workflow is `let` for lowering friction for exploratory or plot cells. In marimo you have to name a `_` prefixed function and then call it which is better than nothing but not as clean as `let`. This is a minor complaint that’s more down to language features though

rsfern··on LoRA Speedrun – a public wall-clock leaderboard for fine-tuning techniques
My point with the force field example wasn’t to argue against neural scaling as a valid strategy, it totally is effective and a lot of groups are doing it. But I feel like we might be talking past each other a bit.

What I’m pushing back on is what I think is a sort of one-dimensional view of Sutton’s bitter lesson. People seem to equate it with model scaling, but there are lots of general ways to leverage computation that don’t involve just scaling models and supervised training datasets up. For example Sutton’s first example is straight up search, no parameters at all.

The point of the force field example is that it seems you don’t need billions of parameters to represent the functions we’re interested in, but with small models it’s harder to find those functions by pushing harder on the standard training algorithms, and that maybe some different algorithm that leverages computation more effectively could do so.

rsfern··on LoRA Speedrun – a public wall-clock leaderboard for fine-tuning techniques
I don’t think there’s a fundamental reason that performance has to be monotonic in model size or even training FLOPs. At least I don’t think it’s been proved to be so, so I think “misinformed” is a bit premature and sort of makes GP’s point.

There’s evidence that model size and representational capacity are not exactly the same, and that scale is maybe more important for learning than it is for representation (past a point). Consider the early work from the current neural scaling paradigm. The Chinchilla scaling study shows that smaller models can match the performance of larger models by training longer.

To GP’s point, if everyone is exploiting the scaling lever, few resources are being allocated to finding more efficient training algorithms that could let us work with right-sized models instead of pulling the scaling lever as hard as we can afford to.

I’ll end with a dramatic example from my field of materials science (which admittedly might not strictly generalize to LLMs). A lot of the field is pursuing the model scaling strategy, and it’s still paying off. But [0] recently reported competitive accuracy with much smaller models that run faster and can address much larger problems. The model architecture is pretty much the same, but they use a different training strategy and really focus on data quality

0: https://arxiv.org/abs/2504.21286

rsfern··on I burned all my tokens researching how to save tokens
I don’t mean to pick on you in particular here, but this approach has been bothering me a lot lately, and it seems like it’s super common.

I get that it’s an early prototype and not all the design choices are made yet, but I struggle with “I can’t afford human-readable documentation yet”. Isn’t human readable documentation important for efficiently planning and deciding what you want to build? I feel like I can’t afford not to have human readable docs and plans while the project is taking shape

Related, if a user can prompt an agent to translate to a human-readable summary, wouldn’t it be better to just do this in place? Sure, models can deal with noisy LLM outputs but shouldn’t a document that’s easier for humans also be easier for bots?

rsfern··on Ask HN: Do you say please and thank you to your LLMs?
I often start prompts with “please”, but I usually don’t thank the model. Framing a question or a request for help with “please” is in distribution for me, it’s a distraction from composing a thoughtful prompt about my actual question to go back and edit out politeness.

I don’t reply “thanks” like I would to a person though, I just close the chat if I have no more follow-ups

rsfern··on US postgraduate student visas will be limited to 4 years
The DHS secretary seems to me to have the point of hosting international students backwards

> This final rule ensures that foreign students remain focused on their primary purpose: completing their studies and returning home.”

Especially at the PhD level, why would educating foreign students and sending them away be the primary motivation for granting visas? Historically it’s been about attracting the best and brightest in the world, training them to be excellent researchers and scholars, and giving them a path to becoming citizens and contributing to growing our economy and technology development and all that. Sending them away after investing in training them for years makes no sense! Especially given the administrations adversarial stance on technology development as a competition between nations more than as an opportunity for collaboration.

I mentor a postdoc who might be affected by this, he’s on a J visa and it’s already been a nightmare of paperwork. He’s really good. Not disputing that there’s some amount of people gaming the student visa system, but this just doesn’t seem well thought out to me :(

rsfern··on Designing APIs for Agents
True. But it has no idea that it has no idea, so it might be able to look back at the session trace and pattern match its way to actionable feedback?
rsfern··on LeMario: Training a JEPA World Model on Super Mario Bros
I think JEPA is super interesting, but I feel like this example highlights some of the challenges of long horizon planning. For one, chunking the planning stage into a bunch of intermediate goals seems really limiting, because a lot of what makes model based control interesting is that we don’t want to impose a solution strategy (because we want to solve problems we don’t know how to solve)

Another thing that has been bothering me is that you have to write the goal in input space. That doesn’t align with all problems, for some problems there could be many different states that satisfy a goal. For Mario maybe it’s ok, but there’s some weirdness still, like should the goal state be Mario at the finish line of the level with a specific timer state in the frame header? What about optimizing the number of points?

Also it’s interesting to think about how you would get Mario to reliably jump on koopas and goombas. IIUC JEPA models are usually trained with random rollouts, and then you’d handle this sort of intermediate goal in the planning optimizer? But that seems inefficient, and including some planning in the pretraining rollouts might be necessary to get enough relevant intermediate states. And then it starts feeling like reinforcement learning…

I’d be happy to have a check on my intuition here, or pointers to interesting writing on these topics

p.s. on topic, I liked the debugging strategies used in the blog post, that was my favorite part of the writeup

rsfern··on World-First 'Super Alloy' Could Transform the Way Metals Are Made
What aspect do you consider basic? I haven’t had a chance to read more than the abstract because of the paywall, but the really interesting thing here is the mechanically induced transformation that leads to a 3-phase nanocrystalline alloy. I haven’t understood the “fully coherent” part from just the abstract, but I think it’s a very novel report. It means they have three different crystal structures with seamless interfaces because all the atoms line up at where the crystals meet. To do that with three structures is remarkable. I’m not sure if it’s the first example, but I’m only aware of alloys with two coherent phases (some superalloys for example)

There’s some precedent for mechanical deformation to get nanocrystalline grain structure, and some precedent for mechanically induced phase transformations (see TrIP steels) but I consider both of those concepts pretty advanced

TrIP is probably the closest thing, but I’m not sure how widely known there are among the “metal forging” community? TrIP is usually targeting well known phase transformations to two-phase microstructures, here we have three nanostructured phases.

Finally high entropy alloys are absolutely not well understood, even if the idea of mixing a lot of elements and getting a disordered solution seems simple on its face.

TrIP: https://en.wikipedia.org/wiki/TRIP_steel

rsfern··on World-First 'Super Alloy' Could Transform the Way Metals Are Made
The science daily article is just incorrect to call this a superalloy, which it is not. This is a high entropy refractory alloy (HfNbTaTiZr), superalloys are usually based on lighter metals and they usually have only one dominant element while HEAs have 4+ dominant elements
rsfern··on World-First 'Super Alloy' Could Transform the Way Metals Are Made
This is really cool metallurgy. They start with an alloy and deform it and because of elemental size mismatch they can cause the alloy to self assemble into nanoscale crystals with three different structures

The paper: https://www.science.org/doi/10.1126/science.aec4995

As an aside, “super alloy” is not the best wording choice on the part of the author of this sciencealert article, superalloys are an established alloy family that follow a different design strategy and have a very different composition profile https://en.wikipedia.org/wiki/Superalloy

rsfern··on Old and new apps, via modern coding agents
Terry Tao has actually been one of the more prominent voices in the math community exploring AI for cutting edge mathematical discovery. This particular post is a bit softer but he has also written a lot about using AI assistance for serious core research

Nov 2025: https://terrytao.wordpress.com/tag/artificial-intelligence/

https://academy.openai.com/public/blogs/terence-tao-ai-is-re...

rsfern··on Unexpected Solidlike Fracture in Simple Liquids
But this is true for lots of solids, not just glass. Consider creep deformation [0] which is deformation mediated by diffusion of defects in solid materials. It’s a big problem in metal turbine blades, it limits the maximum usable temperature to well below the melting point.

The physical mechanism would be similar in glass flowing in this way, so I don’t think evidence of glass flowing like this should make us think of it as a liquid instead of an amorphous solid

0: https://en.wikipedia.org/wiki/Creep_(deformation)

rsfern··on Materials innovation has a scale-up problem, not discovery
Lab to product scale-up is a well known hard problem in materials (and chemistry), lots of public and private investment has been aimed at accelerating this for decades

I think the distinction between discovering a material and processing / scaling it up is a bit artificial. A lot of people think of a new material as just the crystal structure or something, but really all the defects and complex multiscale structure is just as much part of what defines a material, and controlling all that is why materials development is hard, and why you need so many different complementary measurements to understand what’s going on

I was a bit underwhelmed by this writeup because it’s a bit generic. I didn’t really see any specific new ideas on how to accelerate this process, or to differentiate from the main stream of materials discovery research which has been pretty dominantly AI forward for at least five years now

EDIT: I checked out some of their case studies and they’re pretty interesting and exploring some new materials characterizations territory. They’d be more impactful if they were more than just text IMO but much more concrete and less generic than the linked post

rsfern··on Professor denounces mass AI fraud on an exam at Brown
Maybe I have too optimistic a mindset, but “just be honest” in academia isn’t about being a rule-follower, it’s about not short-changing yourself by coasting though on autopilot instead of learning to think and solve problems for yourself.

Whether that really matters if your goal is to climb the social ladder and have power and influence, I don’t know.

rsfern··on Professor denounces mass AI fraud on an exam at Brown
> You allowed a take-home exam which means students are able to use any and all resources.

It was a closed-book exam. The professor shouldn’t have to hold students’ hands for them to act with integrity, they are all adults.

In this particular class, the professor made the final exam in-person, and didn’t count the take-home midterm because the score distribution wasn’t consistent between the two exams. I think that’s a reasonable approach, but it’s kind of sad that it was necessary

rsfern··on Codex logging bug may write TBs to local SSDs
In my experience the input field lags on short chats too, sometimes in the middle of writing the second or third prompt. Are they running some kind of prospective evaluation or something?
rsfern··on Do agents.md files help coding agents?
If the explicit role-playing prompt is just to identify multi-valent terms, then revising the question to include more specific context without a role-play prompt should work just as well right? I’d be really interested if anyone has evaluated that hypothesis

A fun (frustrating) feature of language is that we get these name collisions even with a single domain. One that I have to remember to revise myself fairly often these days when chatting with other experts in my field is “diffusion model” which can either mean generative deep learning or a differential equation describing mass transport.

rsfern··on High-Entropy Alloy
That’s the active research area GP mentioned. In startup land there are a few large outfits, Lila Sciences, Period Labs, Radical AI are all doing a mix of simulations, AI, and autonomous laboratory infrastructure specifically for materials science. (Lila does a lot of biotech but the have materials researchers too)

Also lots of interest and activity in this space in the national labs and academic research scene

rsfern··on High-Entropy Alloy
It depends what you mean by commercially interesting. There’s loads of interest in aerospace (for high temp corrosion resistant structural components) and catalysis but these alloys are pretty much across the board at a relatively low level of technical readiness. It’s developed enough that there’s significant industry R&D and not just academic and government research, but I don’t think there’s really wide-scale deployment yet of alloys with 4+ principal elements
rsfern··on An update on recent Claude Code quality reports
It seems like an opportunity for a hierarchical cache. Instead of just nuking all context on eviction, couldn’t there be an L2 cache with a longer eviction time so task switching for an hour doesn’t require a full session replay?
rsfern··on A Pascal's Wager for AI doomers
I think the exceptionalism is the other way around. What makes anyone think they understand what makes for intelligence when we barely understand our own neurology?
← PreviousPage 2 of 25Next →