yeah for the next training run im planning on testing that out extensively before putting into training
9 karma · joined February 23, 2026
yeah for the next training run im planning on testing that out extensively before putting into training
the only LLM's input into this was as a peer reviewer
I guess I have spent a bit too much time with the models these days
(for the record I mostly use Fable 5)
totally agree with your points on benchmarking and research
also while the website ui was supported by an LLM, all the actual writing was by me(a human) whether it looks somewhat AI generated(well I guess I'm a little bias considering I wrote it) I think I've been speding too much time with the models these days
in all genuinely appreciate the thoughtful insights on my researcH!
The honest headline results: 14% infill recovery where autoregressive models score ~0 (they can't condition on text after the blank), 7.5% repetition-loop rate vs 37.5% for the teacher, and a genuinely negative result I think is the most useful part: six different self-correction methods all failed at this scale, while a 300k-param external critic head detects errors far above chance. Small models don't doubt; they rationalize.
Weights are open: https://huggingface.co/devnull37/hr-diffuse-1-nano. Happy to answer anything about the architecture, the failed runs.