HNHacker News
TopNewBestAskShowJobs

t55

896 karma · joined August 18, 2023

ML researcher
submissionscomments
t55··on Knowledge factories and the industrialisation of knowledge work
what timelines do you have in mind
t55··on Context Rot: How increasing input tokens impacts LLM performance
that's a standard feature in cursor, windsurf, etc.
t55··on There are no new ideas in AI, only new datasets
this is what deepmind did 10 years ago lol
t55··on ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
For a 100k token context window; all those models are comparable though

gemini 2.5 pro shines for 200k+ tokens

t55··on ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
i didn't say they invented everything; in science you always stand on the shoulders of giants

i still think my original statement is fair

t55··on ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
yeah, RLVR is still nascent and hence there's lots of noise.

> How can these spurious rewards possibly work? Can we get similar gains on other models with broken rewards?

it's because in those cases, RLVR merely elicits the reasoning strategies already contained in the model through pre-training

this paper, which uses Reasoning gym, shows that you need to train for way longer than those papers you mentioned to actually uncover novel reasoning strategies: https://arxiv.org/abs/2505.24864

t55··on ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
so you think it's fake news? another example of a paper with strong claims without much evidence?
t55··on ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
agree, the RG evals feel like a fresh breeze
t55··on ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
> prolonged RL training can uncover novel reasoning strategies that are inaccessible to base models, even under extensive sampling

does this mean that previous RL papers claiming the opposite were possibly bottlenecked by small datasets?

t55··on ReasoningGym: Reasoning Environments for RL with Verifiable Rewards
> I personally think that Gemini 2.5 Pro's superiority comes from having hundreds or thousands RL tasks (without any proof whatsoever, so rather a feeling).

Given that GDM pioneered RL, that's a reasonable assumption

t55··on The Bitter Lesson (2019)
it aged well!
t55··on Show HN: Samurai Interview – a mock interview simulator
Very cool, reminds me of https://rehearsal.so/

which sort of interviews will you support?

t55··on Reservoir Sampling
great article!
t55··on Show HN: Debate Uncle Bob – Is SQL Dead? (Voice RPG)
Other debate topics can be found on the frontpage
t55··on Show HN: Skylight – Run computer-use agents on Windows VMs
Congrats! obviously massive potential for saas automations what made you choose windows over other OSs?
t55··on Superintelligence startup Reflection AI launches with $130M in funding
Well, it depends. on some tasks, they surely are already super-intelligent.
t55··on Intro to DeepSeek's open-source week and why it's a big deal
If they don't use some of the released kernels, then they are leaving money on the table
t55··on Intro to DeepSeek's open-source week and why it's a big deal
Never bet against a cracked team
t55··on Intro to DeepSeek's open-source week and why it's a big deal
in a parallel universe...
t55··on Intro to DeepSeek's open-source week and why it's a big deal
agreed, quite ambitious to cover this entire week in one post
t55··on Intro to DeepSeek's open-source week and why it's a big deal
The level of DS's openness has been really useful to actually understand what LLM folks are working on day in day out
t55··on Intro to DeepSeek's open-source week and why it's a big deal
Indeed. Modern LLM research is more about large-scale systems design than deriving gradients by hand. Math matters, but systems win.
t55··on Claude 3.7 Sonnet and Claude Code
Anthropic doubling down on code makes sense, that has been their strong suit compared to all other models

Curious how their Devin competitor will pan out given Devin's challenges

t55··on Introduction to CUDA programming for Python developers
thank you!
t55··on Introduction to CUDA programming for Python developers
hot take: i don't think you even need to understand much linear algebra/calculus to understand what a transformer does. like the math for that could probably be learned within a week of focused effort.
t55··on Introduction to CUDA programming for Python developers
you're welcome!
t55··on Introduction to CUDA programming for Python developers
Yes, it is great for key concepts but a bit outdated. Hence we added an LLM/FA section in the linked post!
t55··on Introduction to CUDA programming for Python developers
I should have been more precise, sorry. Didn't want to imply they entirely ditched CUDA but basically circumvented it in a few areas like you said.
t55··on Introduction to CUDA programming for Python developers
Agreed, not sure how much math is really needed.
t55··on Introduction to CUDA programming for Python developers
lol. i guess this tutorial is about cutting out guido ;)
Page 1 of 3Next →