HNHacker News
TopNewBestAskShowJobs

ag8

2,006 karma · joined July 3, 2019

runrl.com
submissionscomments
ag8··on ASML became the chokepoint for cutting-edge chips
I find this paragraph to be odd: "Wavelengths as low as 13.5 nanometers can achieve more precise patterns in a single exposure. In fact, extreme ultraviolet lithography can combine three or four photolithography patterning cycles into a single one on a seven-nanometer node. Without EUV, producing five-nanometer nodes might require as many as one hundred different steps."

But a five-nanometer node has a gate pitch of 45nm and a metal pitch of 20nm! Using different forms of the word "nanometer" in the same paragraph is very confusing...

ag8··on Thank HN: You helped save 33k lives
You're right; I should've been more precise. However, we have tools for dealing with this—that's what quality-adjusted life-years are for! I don't contest that surgeries often significantly increase QALYs, and may do so pretty cost-effectively.
ag8··on Thank HN: You helped save 33k lives
Lol, I just care a lot about saving as many lives as I can; the most effective charities I've been able to find good evidence on save one life for $6–8k. If Watsi had a credible claim at being able to save lives 10x cheaper I would redirect my entire donation budget to them!

That said, once again, Watsi is great. I really appreciate all the hard work they've put into making this happen—this is orders of magnitude more impressive and impactful than most projects I've ever seen!

ag8··on Thank HN: You helped save 33k lives
Watsi seems to be doing great work, but the title—"you helped save 33k lives"—reads as misleading to me. I guess "helped" could be doing a lot of heavy lifting here, but I would be incredibly surprised if the counterfactual number of lives saved was more than 3000. (But don't let this dissuade you from donating; concretely improving someone's life is totally a worthwhile goal, and Watsi seems very good at effecting this)
ag8··on Ask HN: Share your personal website
https://andrew.gr
ag8··on Size of Life
Not 13?
ag8··on We collected 10k hours of neuro-language data in our basement
This is a cool setup, but naively it feels like it would require hundreds of thousands of hours of data to train a decent generalizable model that would be useful for consumers. Are there plans to scale this up, or is there reason to believe that tens of thousands of hours are enough?
ag8··on Tinker
Yeah, not sure why the HN backend changed it...
ag8··on Launch HN: RunRL (YC X25) – Reinforcement learning as a service
A) You could have an additional field in the jsonl file which says which rubric to use; then, your reward function could access this via `kwargs["rubric"]` and return a reward based on that example's preferred rubric;

B) currently, pricing on the deployed API is free, but the startup time is a few minutes and it's run on a small GPU node and is therefore not awfully fast. If you would like more production-level inference, email us at founders@runrl.com and we could set you up with something much faster (where we'd charge per token depending on model size)

ag8··on Launch HN: RunRL (YC X25) – Reinforcement learning as a service
Having an RL agent that's really good at search across some space sounds very powerful in general; "proofs-as-search" make this an appealing target. Back in the day, when I did more fundamental RL research, we worked on an extension of SoRB [0] where an additional meta-level target was learning improved heuristics to explore the search space faster; would be exciting to figure out what a good setup for doing things like this in LLM-policy-gradient world is these days!

[0]: https://arxiv.org/abs/1906.05253

ag8··on Launch HN: RunRL (YC X25) – Reinforcement learning as a service
we should publish some; the high-order effect seems to be that LoRAs significantly hurt small model performance vs FFT, with less of an effect for large models. This is maybe because large models have more built-in skills and thus a LoRA suffices to elicit the existing skill, whereas for small models you need to do more actual learning (holding # parameter updates constant). In general I think it's better to get a performant small model with FFT than a performant large model with a large LoRA, which is why we default to FFT, but I agree that we should publish more details here.
ag8··on Launch HN: RunRL (YC X25) – Reinforcement learning as a service
Thanks! Our goal is to make rl "just work" with completely automated GPU provisioning/algorithm selection/SFT-warm up, but giving people the ability to switch away from the defaults if they want to.

The way tools currently work in the beta is you add tools via MCP to the configuration, and they get passed in as additional context for the model; the model might then choose to use a tool during inference; the tool is then automatically called and the output is returned as a tool message. If you really want to you could parse the tool output as part of reward calculation, but I expect you'd usually base the reward just on the model's completion. I could give more details if there's a specific tool setup you're envisioning!

ag8··on Launch HN: RunRL (YC X25) – Reinforcement learning as a service
Yeah, for better or worse, the way the median startup interfaces with AI these days is through an LLM API, and that's what all the workflows are built around, so that's what we're targeting. Though, depending on what you're trying to do, I wouldn't discount the use of starting with a pretrained model—there was that famous result from 2022 that showed that pretraining a model on _Wikipedia_ made training on Atari games more than twice as efficient [0]; these days, LLMs have huge amounts of priors about the real world that make them great starting points for a surprisingly diverse set of tasks (e.g. see the chemistry example in our video!)

[0]: https://arxiv.org/abs/2201.12122

ag8··on Launch HN: RunRL (YC X25) – Reinforcement learning as a service
It's for any task that has an "eval", which is often verifiable tasks or ones that can be judged by LLMs (e.g. see [0]). There's also been recent work such as BRPO [1] and similar approaches to make more and more "non-verifiable" tasks have verifiable rewards!

[0]: https://runrl.com/blog/funniest-joke

[1]: https://arxiv.org/abs/2506.00103

ag8··on Launch HN: RunRL (YC X25) – Reinforcement learning as a service
prompt optimization is very cool, and we use it for certain problems! The main goal with this launch is to democratize access to "the real thing"; in many cases, full RL allows you to get the last few percent in reliability for things like complex agentic workflows where prompt optimization doesn't quite get you far enough.

There's also lots of interesting possibilities such as RLing a model on a bunch of environments and then prompt optimizing it on each specific one, which seems way better than, like, training and hot-swapping many LoRAs. In any case, _someone's_ ought to provide a full RL api, and we're here to do that well!

ag8··on Revisiting Image Maps
wow, I used to make so many games with image maps back when I first learned HTML. One still survives: https://andrew.fi/beowulf/game/
ag8··on Just redesigned my personal site with a TTY-style interface
Reminds me of https://szge.ca, which comes with fake keyboard and fan noises:)
ag8··on Humans are unreliable models of mouse disease
Yeah, I'm a bit surprised that this has so many upvotes without any easily accessible text
ag8··on Show HN: HN Update – Hourly news broadcast of top HN stories
I started listening to it after this very submission became #1 on HN, so it was very meta to listen to it talk about how it might be making stuff up...

Great project!

ag8··on Ask HN: How to Improve Memory?
Have you looked into Anki/Supermemo? e.g., https://jmcavanagh.neocities.org/ankiforscience/ankiforscien... provides a great example of using it to remember scientific facts, derivations, algorithms, etc
ag8··on Seeing America by train
Apparently the California to Denver part is especially nice; I'd recommend starting with that!

https://derikk.com/blog/zephyr.html

ag8··on Ask HN: Struggling with poor memory and executive function. What to do?
anki?
ag8··on Japanese words and names sound African (2022)
Yeah, this felt like a Gell-Mann moment
ag8··on Line-Item Vetos in Wisconsin
Tony Evers using this to make annual per-pupil funding increases permanent: https://archive.is/0hu0T
ag8··on DMV tells Cruise to reduce its driverless vehicle fleet in SF by 50%
Also discussed here: https://news.ycombinator.com/item?id=37184904
ag8··on Ask HN: Could you share your personal blog here?
https://andrew.gr/stories/
ag8··on The mid in fake midcentury modern
Regardless of my thoughts on the content, this article is very beautifully written. The word choice is impeccable, and reads like a thoughtfully shot short film of the author's day.
ag8··on Make-A-Video: AI system that generates videos from text
Another video generation model from today: https://news.ycombinator.com/item?id=33025189
ag8··on Phenaki: A model for generating minutes-long, changing-prompt videos from text
Wow--this qualitatively feels a lot more impressive than the Meta model. The two-minute video is better than anything I've seen in video generation on that scale.
ag8··on Show HN: A US Evolution Simulator
Yeah! It was also cool that this let me discover new secession movements; I had never heard/thought about the Maryland/West Virginia one[1], or the New York/Pennsylvania one[2], until the simulation showed how stark the voting differences are.

[1]: https://www.wusa9.com/article/news/verify/maryland-counties-...

[2]: https://web.archive.org/web/20150219175611/http://www.wbng.c...

Page 1 of 2Next →