HNHacker News
TopNewBestAskShowJobs

rsfern

1,745 karma · joined February 5, 2015

Working as a researcher in materials science (physical metallurgy specifically).

My research interests include applied machine learning for quantifying materials substructure, which we call microstructure.

submissionscomments
rsfern··on MentalHealthBench
I agree that accessibility is a big problem, and take your point about ignoring technology that could help, but if the technology if not there yet then I think more research and discourse is warranted before deploying it widely. As a society we need to decide what kind of requirements we expect of these systems, then agree on standards of performance

Is it ok for models to offer specific mental health interventions, or just provide neutral information? To what degree is that technically achievable? Does the answer depend on the scenario?

Health care workers have mandated reporting responsibilities (in the US, I assume other countries are similar) for some situations. Should models be bound by this responsibility as well? (I think yes but it is a thorny question). How to achieve that while preserving patient privacy for issues that don’t fall under mandated reporting rules?

We also have standards of care for health workers. Benchmarks are good, but that’s similar to a licensing requirement, and we have accountability mechanisms when people don’t adhere to their ethical and professional duties. What should be the liability situation when a model responds out of standard resulting in harm (to the standard of evidence we would hold a human health worker)?

Maybe I am not plugged in to this space enough, but my impression is these questions aren’t really in the foreground

rsfern··on MentalHealthBench
I find this troubling, there are so many potential ethical issues with this application of language models. Maybe it would be more appropriate for models to issue safety refusals and help the user learn how to get help from a licensed mental health professional
rsfern··on Did OpenAI solve the wrong Navier-Stokes problem?
I think this is dramatized to the point it’s talking past the article, the math community isn’t really making any of these arguments from what I can tell.

The discourse is (1) models are capable of making really impressive mathematical advances, usefulness is not in dispute, (2) the frontier AI companies aren’t being super transparent about information sources so it’s hard to know exactly how to evaluate the level of capability that was demonstrated, and (3) there are lots of kinds of math that is interesting and there are open questions about how to get there.

In particular this article highlights a particular open question I’ve seen discussed on HN before, which is that the particular proof strategy of finding a counterexample might be more amenable to RL than other strategies of proof that might be needed to resolve the other branches of the Navier Stokes problem (and probably other similar areas of math)

rsfern··on Chat-based Large Language Models replicate the mechanisms of a psychic's con
According to OpenAI, but they haven’t exactly been transparent about what information the prompt entailed.

The bigger question is to what extent did expert mathematicians metaprompt the model with fruitful solution strategies through their sessions finding their way into training data. Answering that question definitively is kind of important for understanding the models contribution/capability. But I feel like people want to turn this into a debate about priority and credit which is sort of secondary

rsfern··on Asking Authors About Their Own Papers
I think that’s a lot of risk of anchoring reviewer bias. I’d be more comfortable with a triaged review where the editor’s office uses models to score whether a human editor should evaluate a paper to potentially send out for review, then the editor makes their own assessment, and the reviewers continue to do their job unassisted
rsfern··on I Don't Like LLMs
For me the convenience of “hey remember this fact” is outweighed by a desire not to get stuck in a search or context bubble.

It might be nice to have better UI to control which bits of history get added to the context of a chat, but then just use a coding harness instead of web UI

rsfern··on Don't build tools for AI agents
It should resonate here. The thesis is that good design principles are transferable between humans and agents. So focus on solving important problems and build well designed tools and documentation to get there, you don’t need special design considerations just because agents

I don’t know about the startup space, but in AI for science people are spending time building MCP wrappers around poorly designed APIs instead of redesigning the API or building an abstraction layer that humans can also use. That seems like a mis-allocation of effort

rsfern··on NSF memo implementing Trump's 'golden age' of science unsettles researchers
There’s no reason it has to be like that. We could change course one more time back to a stable and non-partisan science funding landscape and then stick with it if we choose to. Even some of the alarm in this article over the proposed funding cuts isn’t set in stone, the presidential budget request from last year aimed to cut NSF by a similar amount and Congress didn’t go for it
rsfern··on Why I'm still bearish on LLMs after Navier-Stokes
On the contrary, I think the chess comparison is on point. We’re discussing observations that even the strongest models devolve into making invalid moves without scaffolding. For me that raises the question of whether these models are learning the rules and generalizing from them, or of they’re just pattern matching and flailing on this task. Maybe the reality is somewhere in between, but the benchmarks don’t seem to directly measure conceptual generalization, they measure task completion. They can disrupt a lot of people and industries by pattern matching and flailing without being AGI.

I’m sure these models know the rules and can explain them when prompted, but that doesn’t seem to be the way they actually complete this task. Will they get there? Maybe

rsfern··on Why I'm still bearish on LLMs after Navier-Stokes
Maybe it depends on the person? My six year old isn’t great at strategy but they can pretty consistently make valid moves. Sometimes they ask for confirmation on a move which is also not a trait I see in language models (at least unprompted)
rsfern··on Why I'm still bearish on LLMs after Navier-Stokes
the discussion isn’t really about whether language models can become strong chess players though, the point is they seem to struggle to consistently make valid moves. Most humans don’t need to read two books to pick that up, just a couple lines of basic instructions
rsfern··on Anecdotally, Programmers Dislike "Reduce"
For numerical code I like einops.reduce more than numpy/pytorch sum reductions because you can reduce over named dimensions. It’s much more readable than having to reason through axis indexing again every time you come back to the code
rsfern··on OpenAI have no mathematicians capable of understanding what they put out
I agree (and so does Buckmaster based on his written statement) that we are better having solved this.

But I disagree that which humans were credited is the heart of the issue in this particular controversy. The question is what do you need to bring to the table for a result like this. A pre-release frontier model trained on the open literature and $15 million of inference? Or all that plus a year of the experts finding the path to the solution for the model to run with?

I think it makes a huge difference in terms of what we think the future of mathematical research will be like, and whether we should still encourage students to go into this field, which was the original topic of this thread

rsfern··on OpenAI have no mathematicians capable of understanding what they put out
That’s the prevailing narrative, but I think this controversy calls it into question to some extent. If the OpenAI result wouldn’t have been possible without experts seeding the training data with feedback on promising solution routes, there’s less reason to believe this, IMO. More information and transparency is needed
rsfern··on More questions about whether researchers can trust OpenAI with unpublished math
I think you’re missing an important distinction. “Major damage” to the talent pipeline because models become capable of original end-to-end mathematics is what the community has been discussing. But if the models rely on sniping nearly complete work then this damage is antisocial without a lot of upside, it would be destroying a talent pipeline that would still necessary for continued progress.

Which is it? I don’t think OpenAI is being transparent enough for us to really understand whether these results would have been possible without relying on unpublished information from the solution strategies of the experts

rsfern··on More questions about whether researchers can trust OpenAI with unpublished math
The session data could be cryptographically signed. Probably easier in an open harness?
rsfern··on Is OpenAI Taking Everyone for Fools?
The screen cap of their full correspondence in the Twitter thread, which was the subject of the preceding sentence. The Twitter post has an obviously incomplete fragment of the conversation that doesn’t resolve what the author presents it as resolving, which is the dispute over how the discussion of authorship of Alpöge actually went down
rsfern··on Is OpenAI Taking Everyone for Fools?
Without seeing the full correspondence it’s hard to evaluate for sure, but parent linked to a tweet from the OpenAI employee at the center of the controversy, that’s a primary source you can read and evaluate yourself

Personally I don’t find the tweet a satisfactory explanation of their behavior, it seems like a lot of deflection without directly responding to the specific claims of front-running, and the screen cap of the correspondence doesn’t include all the relevant context. If there was really no bad behavior, why not post the whole thing?

rsfern··on How An AI math breakthrough ignited a controversy
It does, yes. So designing objections functions and making sure you can afford the training rollouts becomes really important in defining which problems are tractable. It will be really interesting to see how that shapes the kinds of problems people choose to work on
rsfern··on How An AI math breakthrough ignited a controversy
Agreed, but i think this underscores my point. We have numerical simulations in materials science too, but that doesn’t mean formally verified theorems about the underlying equations automatically translate to formal (or even informal) verification of simulation results. That’s not to say you can’t make progress with agents, but I think it’s less well defined how you write the goal and progress assessment for an agent
rsfern··on How An AI math breakthrough ignited a controversy
Yes, definitely! There’s a long history of this and I think there’s tons of opportunities for more. Both for improving the exactness/physical fidelity of models and for developing new approximate theories and simulation methods
rsfern··on How An AI math breakthrough ignited a controversy
Regardless of what you think of the priority dispute issue discussed on sibling threads, I’m highly skeptical of the closing quote that this Navier Stokes result means that the same approach of casually spending a few million on agentic computation is going to solve end to end materials design or drug development.

Those problems can’t be formally verified with an automated theorem prover. We have a lot of physics based simulation tools, but they tend to focus on small subsets of the full design problem and they make limiting approximations because otherwise they’d be too computationally expensive, or we just don’t have the right data to parameterize them beyond describing qualitative behavior. Agents are helping accelerate research in these fields but I think it’s mostly a different class of problem that’s a lot harder to specify and verify

rsfern··on An Emacs-style visual undo tree for Pi session branching
Thanks! This seems really cool. If I’ve got it right, your UI builds and displays these diffs (and implements undo/redo) by parsing the edit tool calls?

If so that seems really nice, one of the things I don’t like with coding agents is having to defensively commit changes to roll back if the agent goes off the rails. I’m not sure if that’s a problem with my workflow, but your tool seems great for exploratory stuff

How does it change the way you personally use pi? How git aware is it, do you have to commit at the end of a branching session, and does it handle manual edits?

rsfern··on An Emacs-style visual undo tree for Pi session branching
Interesting project idea! The link seems to be 404, is the repo still private?
rsfern··on Navier-Stokes – Tristan Buckmaster [pdf]
Why would mining chat transcripts for ideas be untenable? They already run a summarization model to auto-title the chat, and to run a bunch of safety filters, and presumably to score transcript quality for A/B testing and to collect more finetuning data. Seems like evaluating for open research questions and approaches would be pretty trivial extension of this, after all it’s kind of their core business model
rsfern··on Trump Ripped by Reporters Used as Decoys in Secret Escape
The bit you quoted doesn’t capture why the reporters are upset:

> White House journalists are outraged that a threat credible enough to force Donald Trump to escape from Air Force One using an airport catering truck wasn’t relayed to them.

rsfern··on U.S. Department of Energy Launches the Genesis Open Models Initiative
Right, I did specifically say that most of the DOE scientists are contractors, but I concede the phrase “government scientist” is a bit ambiguous. I appreciate the extra detail you added. I think the distinction between political appointee and scientist/researcher stands.

As an added complication, some of the DOE labs do have civil servant scientists, for example National Energy Technology Lab and National Renewable Energy Lab are like 50/50 civil servants and contractors. And most of the funding arm of DOE are career civil servants. LANL, Sandia, Livermore, Argonne are all staffed by contractors

rsfern··on U.S. Department of Energy Launches the Genesis Open Models Initiative
Let’s distinguish a bit. There are political appointees (Trump’s government employees as you say) who are mostly upper management, and there are career civil servants (all the government scientists are under this category) who have a strong culture of apolitical dedication to the mission of their agency and to the American people and Constitution, regardless of who the current president is. And in the DOE labs in particular most (not all) of the scientists are actually employed as government contractors, but they have a similar non-partisan ethos.

That doesn’t necessarily mean there’s no need to be concerned with potential impact of policy and priority changes from the administration, but it does temper the threat model because the government employees you’re considering trusting have given oaths of office to protect and defend the Constitution.

rsfern··on Position: LLMs Can't Jump
I found this paper really thought provoking, but I think the conclusion of “world models are the solution” leaves something to be desired. People are already equipping agentic systems with physical simulation tools and exploring action-conditioned world models. This is cool because you can change the rules of the simulation and observe what happens, but it doesn’t address the core question of what to change the rules to, or even what the goal should be in the first place.
rsfern··on The session you cannot take with you
Or they could store the reading traces and validate the user hasn’t edited them server-side? They could sign reasoning traces so they can’t be counterfeited?
Page 1 of 25Next →