6,257 karma · joined March 6, 2013
Currently I'm working on openpipe.ai. Previously worked at YC and Google.
personal site: corbt.com email: kyle@ above. I respond to emails.
We could measure order bias really easily though; we just need to look at the average score by rollout position across many runs. I'll add that to my list of experiments!
Also, how realistic would it be to share the KV cache across vllm nodes within a data center? It would be really nice to be able to freely distribute requests to a pool of vLLM workers without worrying about prefix-aware routing, but maybe that isn't the right approach because moving the KV cache around would be too slow?
https://chatgpt.com/share/685dea79-26ec-8002-bd62-7ed83aedf4...
And of course the effect on throughput at larger batch sizes, which they allude to at the end.
Overall a very interesting result!
The problem comes if the number of years of experience you need to outperform the frontier AI models advances at more than 1 per year, which is not out of the question.
Experimenters in the open source tinkering community have done the opposite (copy/pasting layers in existing models to make them deeper) and it seems to work... fine, with minimal post-training on the new, deeper model required to exceed the performance of the original model. So it's not a crazy idea.
Most likely they built this as a post-train of an open model that is already strong on coding like Qwen 2.5.
By fine-tuning in this context I assume you mean "supervised fine-tuning", or SFT. SFT trains a model to produce a specific string of output tokens, given an input. With SFT, if you were trying to train an assistant to solve math problems using a code interpreter, you might train it on a dataset that looks like:
input: 'What is 934+1208'
output: `print(934+1208)`
input: 'how many "r"s in strawberry'
output: `print(len([l for l in "strawberry" if l == 'r'])`
etc, etc.RL, on the other hand, just means training a model not to produce a concrete string of output tokens, but rather to create an output that maximizes some reward function (you get to decide on the reward).
For the example above, you might create the following dataset for RL training:
input: 'What is 934+1208'
ground_truth: 2142
input: 'how many "r"s in strawberry'
ground_truth: 3
You would then train the model to write python code that produces the ground_truth output. Your training code would take the model's output, run the python it produced, and then check whether the output matches the expected ground_truth. Importantly, this doesn't require you actually writing the code to solve the problem (you don't even have to know if it's solvable, technically!). Over time, the training loop would make the model more likely to produce outputs that get high rewards, which hopefully means it gets better at producing valid and applicable python.This is useful in lots of domains where it's easier to check the answer than actually produce it. In the blog post[1] linked above, we train the agent to effectively use keyword search to try to find the correct emails in an inbox. As the model trainer, I didn't actually know what the right strategy was to choose keywords that would most quickly find the relevant email, but through training with RL, the model was able to figure it out on its own!
[1]: https://openpipe.ai/blog/art-e-mail-agent?refresh=1746030513...
I think there's a lot of benefit to discovering a training regime that allows small specialized models to do extremely well in one narrow task; if we can figure out how to make small models that beat SOTA on a specific task and are cheap to train and run, that's in some ways a more useful outcome than a very large model that is good at many tasks (but is more expensive to run for each of them).
Our customers range from fast-growing startups to Fortune 500s. We're growing 40% MoM and have achieved this with a team of just 5 engineers, including the founders.
Seeking:
- ML Engineers: Help advance our fine-tuning capabilities. (We have thousands of datasets and evals; testing new ideas is really easy!)
- Systems Engineers: Scale our platform serving hundreds of millions of daily requests
- Full-stack Engineers: Build end-to-end features in TypeScript/Python
Ideal candidates:
- Strong programming skills (TypeScript/Python)
- Systems architecture expertise
- Self-starters (founder/founding engineer experience valued)
- Willing to relocate to Seattle
Highly competitive salary + equity. Well-funded with large customers. Perfect for future founders wanting startup experience or engineers breaking into AI/ML with immediate impact.
Email: kyle@openpipe.ai Include something impressive you've built.
I also recently wrote a blog explaining how reinforcement fine-tuning works, which is likely at least part of the pipeline used to train o1: https://openpipe.ai/blog/openai-rft
We've built the world's best fine-tuning platform in just over a year. First to launch self-service preference tuning, integrated evals/data prep/fine-tuning, and self-service learning from human feedback.
Our customers range from fast-growing startups to Fortune 500s. We're growing 40% MoM and have achieved this with a team of just 5 engineers, including the founders.
Seeking:
- ML Engineers: Help advance our fine-tuning capabilities. (We have thousands of datasets and evals; testing new ideas is really easy!)
- Systems Engineers: Scale our platform serving hundreds of millions of daily requests
- Full-stack Engineers: Build end-to-end features in TypeScript/Python
Ideal candidates:
- Strong programming skills (TypeScript/Python)
- Systems architecture expertise
- Self-starters (founder/founding engineer experience valued)
- Willing to relocate to Seattle
Highly competitive salary + equity. Well-funded with large customers. Perfect for future founders wanting startup experience or engineers breaking into AI/ML with immediate impact.
Email: kyle@openpipe.ai Include something impressive you've built.
Of course, it isn't your IP free and clear either, because the base model isn't open so your fine-tuned model will always live inside OpenAI's walled garden.
If you're interested in reinforcement learning on top of truly open models where you own the end product, we're putting a lot of thought into that and are also looking for design partners! Feel free to email me at kyle@openpipe.ai.
Broadly, the main use case for this model (in the RL context) will be to take two different versions of the same post, and predict which of the two is more likely to be upvoted. So what matters isn't that it gets the exact number of upvotes correctly, but that it correctly predicts the relative difference in likely upvote count between two variants.
Now it still doesn't do a great job at that (the correlation is only 0.53 after all) but it still does a good enough job to provide some useful signal.
Everything else in the model before that final layer is exactly identical, architecture-wise.
> it's good to create two models, one for likelihood of zero karma, and another expected karma, conditional on it being non-zero.
Another way to do this is to keep a single model but have it predict two outputs: (1) likelihood of zero karma, and (2) expected karma if non-zero. This would require writing a custom loss function which sounds intimidating but actually isn't too bad.
If I were actually putting a model like this into production at HN I'd likely try modeling the problem in that way.
Yes, this is a fantastic point. I'm curious if there's some other measurable proxy metric for "things I get the most value out of on HN"? Upvotes seems like the most natural but optimizing for it too strongly would definitely take HN down a dark path.
And all the graphs for the blog are from this notebook: https://github.com/OpenPipe/best-hn/blob/main/blog-figures.i...
Lots of other good stuff in that repo, although it's only organized to a "working researcher" standard I'm afraid.