HNHacker News
TopNewBestAskShowJobs

roborovskis

800 karma · joined August 9, 2013

submissionscomments
roborovskis··on Anthropic announces proof of distillation at scale by MiniMax, DeepSeek,Moonshot
What would you define as 'distillation' versus 'learning'? How do you know that what a LLM is doing is 'distillation' vs a process closer to a human reading a book?

From my perspective, pretraining is pretty clearly not 'distilling', as the goal is not to replicate the pretraining data but to generalize. But what these companies are doing is clearly 'distilling' in that they want their models to exactly emulate Claude's behavior.

roborovskis··on DeepSeek-R1
Where are you seeing this? On https://github.com/deepseek-ai/DeepSeek-R1/tree/main?tab=rea... I only see the paper and related figures.
roborovskis··on OpenAI Reinforcement Fine-Tuning Research Program
https://stable-baselines3.readthedocs.io/en/master/ is a great resource for hacking on implementations for RL - many good RL courses out there but https://www.youtube.com/playlist?list=PLwRJQ4m4UJjNymuBM9Rdm... is my personal favorite.

For LLMs / RLHF it's a little more difficult but https://github.com/huggingface/alignment-handbook and the Zephyr project is a good collection of model / dataset / script that is easy to follow.

I would suggest studying the basics of RL first before diving into LLM RLHF, which is much harder to learn on a single GPU.

roborovskis··on SuperPrompt: Better Text to Image Prompts in 77M Parameters
You could definitely use this for upsampling negative prompts, though I haven't tested that much. In theory, future T2I models shouldn't need to be negatively prompted as much; I find it's better to focus on really high quality positive prompts, as that is closer to the captions the model was trained on.

You can take a look at the dataset here: https://huggingface.co/datasets/roborovski/upsampled-prompts... Roughly 5k samples were needed for the smaller ones at a minimum, filtered from the 95k total generated.

roborovskis··on SuperPrompt: Better Text to Image Prompts in 77M Parameters
Yup, the model will still forget details sometimes. This is a common issue with prompt upsampling methods, but I'm hoping to improve this with the next version.
roborovskis··on SuperPrompt: Better Text to Image Prompts in 77M Parameters
Thanks for the kind words! I started with the 780M param flan-t5-large model, and kept trying smaller and smaller base models - I was shocked at how good the output was at 77M. As you go smaller, though, it's much easier to accidentally overfit or collapse the model and produce gibberish. Had to be very careful with hyperparams and sanitizing / filtering the dataset.
roborovskis··on SuperPrompt: Better Text to Image Prompts in 77M Parameters
As Invoke is open-source and already has transformers as a dependency, it should be pretty easy to add.
roborovskis··on SuperPrompt: Better Text to Image Prompts in 77M Parameters
I haven't tested extensively with non SDXL based checkpoints but there's nothing really SDXL specific about the model; if you're using a fine-tune that's trained on booru-style tags, it will probably not work as well - but otherwise it should work just fine. And in that case, just fork the project and tune it on however your fine-tune prompts best :)
roborovskis··on SuperPrompt: Better Text to Image Prompts in 77M Parameters
will fix these, thanks for the heads up!
roborovskis··on Stable Diffusion XL 1.0
https://dreamstudio.ai/
roborovskis··on Panopticlick – How Unique, and Trackable, Is Your Browser?
The fact that they can see what plugins I have makes uTorrent wanting a browser plugin make total sense. Just thinking.