HNHacker News
TopNewBestAskShowJobs

t55

896 karma · joined August 18, 2023

ML researcher
submissionscomments

Sokoban Speedrun for RL

github.com·6 pts·t55·
0

RL Speedrun

github.com·2 pts·t55·
0

Target Policy Optimization

arxiv.org·1 pts·t55·
0

Show HN: Kilroy – Knowledge base for teams using Claude Code

github.com·5 pts·t55·
0

Procedural Reasoning Datasets

github.com·1 pts·t55·
0

In Defence of Gary Marcus

reubenadams.substack.com·3 pts·t55·
0

Reasoning Gym – Procedural RL reasoning datasets

github.com·1 pts·t55·
0

ChatGPT Agent [video]

youtube.com·3 pts·t55·
0

ReasoningGym: Reasoning Environments for RL with Verifiable Rewards

arxiv.org·105 pts·t55·
28

Show HN: Rehearsal.so, Duolingo for Public Speaking

rehearsal.so·3 pts·t55·
1

End-to-End Vision Tokenizer Tuning

arxiv.org·3 pts·t55·
0

YC Interview Mock Practice

rehearsal.so·2 pts·t55·
0

D1: Scaling Reasoning in Diffusion LLMs via Reinforcement Learning

dllm-reasoning.github.io·4 pts·t55·
0

Are LLMs more than autocomplete? AI Debate

rehearsal.so·1 pts·t55·
0

Block Diffusion: Interpolating Autoregressive and Diffusion Language Models

m-arriola.com·72 pts·t55·
16

How to stay in flow while using Cursor or Windsurf

rehearsal.so·2 pts·t55·
0

Generative Modelling in Latent Space

sander.ai·2 pts·t55·
0

Show HN: Debate Uncle Bob – Is SQL Dead? (Voice RPG)

rehearsal.so·6 pts·t55·
1

OpenAI O3 and O4-Mini

openai.com·1 pts·t55·
0

Memory in ChatGPT

twitter.com·10 pts·t55·
0

Superintelligence startup Reflection AI launches with $130M in funding

siliconangle.com·38 pts·t55·
26

Intro to DeepSeek's open-source week and why it's a big deal

pyspur.dev·24 pts·t55·
13

Introduction to CUDA programming for Python developers

pyspur.dev·365 pts·t55·
95

Novelty Left on the Table

ansatz.blog·2 pts·t55·
0

Competitive Programming with Large Reasoning Models

arxiv.org·16 pts·t55·
1

The Differences Between Direct Alignment Algorithms Are a Blur

arxiv.org·8 pts·t55·
0

The Octalysis Framework for Gamification and Behavioral Design

yukaichou.com·3 pts·t55·
0

S1: Simple Test-Time Scaling

github.com·40 pts·t55·
3

A Malloc Tutorial [pdf]

github.com·1 pts·t55·
0

Reinforcement Learning: An Overview

arxiv.org·82 pts·t55·
12
Page 1 of 3Next →