HNHacker News
TopNewBestAskShowJobs

drubs

96 karma · joined March 2, 2025

Making models go brr

https://github.com/drubinstein https://x.com/dsrubinstein

submissionscomments
drubs··on Beam: Reflection's 501B open-weight model
Outside of ML metrics, you're monitoring the health of every piece of hardware in the system. You need to make sure that you have every GPU, every CPU, the PCIe buses, the networking fabric are all working without any errors. You need to ensure that you can respond as fast as possible to any possible error. One bad component can bottleneck the entire job.
drubs··on Beam: Reflection's 501B open-weight model
I remember being in the room with pretraining day 1 to help monitor the training job launch. Watching this model train from day 1 has been an amazing experience!
drubs··on Less is more: When agents learn not because but despite my doings
Fantastic work! This was a really fun collaboration.
drubs··on Scaffolding to Superhuman: How Curriculum Learning Solved 2048 and Tetris
Star the puffer https://github.com/PufferAI/PufferLib
drubs··on Show HN: Beating Pokemon Red with RL and <10M Parameters
Wouldn't make much sense. We generally train with 288 environments simultaneously. I've been thinking about ways to nicely stream all 288 environments though.
drubs··on Reflection – AlphaGo / Gemini team building superintelligent coding agents
Really excited to be a part of the team!
drubs··on Show HN: Beating Pokemon Red with RL and <10M Parameters
Sounds cool to me.
drubs··on Show HN: Beating Pokemon Red with RL and <10M Parameters
Yup!
drubs··on Show HN: Beating Pokemon Red with RL and <10M Parameters
It's silly, but signs were a way to incentivize the agent to explore deeper into the Safari Zone among other areas.
drubs··on Show HN: Beating Pokemon Red with RL and <10M Parameters
My first version of this project 5 years ago involved a python-lua named pipe using Bizhawk actually. No clue where that code went
drubs··on Show HN: Beating Pokemon Red with RL and <10M Parameters
There's a ton of applications for AI. Back when I was at Spotify, I co-authored Basic Pitch (https://basicpitch.spotify.com/), an audio-to-midi library. There are a ton of uses for AI outside of what's heavily publicized.
drubs··on Show HN: Beating Pokemon Red with RL and <10M Parameters
There's an entire section on how the decompilations were used :)
drubs··on Show HN: Beating Pokemon Red with RL and <10M Parameters
Wrote about this in the results section. I think there is a way to mix the two and simplify the rewards in the process. A lot of the magic behind getting the agent to teach and use cut probably could have been handled by an LLM.
drubs··on Show HN: Beating Pokemon Red with RL and <10M Parameters
The environments wouldn't concentrate enough in the Rocket Hideout beneath Celadon Game Corner. The agent would have the player wander the world reward hacking. With wild battles enabled, the environments would end up in Lavender Tower fighting Gastly.

> (and how on earth did you port Pokémon red to a RL environment? O.o)

Read and find out :)

drubs··on Show HN: Beating Pokemon Red with RL and <10M Parameters
and...fixed!
drubs··on Show HN: Beating Pokemon Red with RL and <10M Parameters
Thanks for the heads up. I just pushed a fix.