HNHacker News
TopNewBestAskShowJobs

dippogriff

2 karma · joined October 27, 2025

submissionscomments
dippogriff··on 60% Fable cost cut by converting code to images and having the model OCR it
I want to see more text-free foundation models
dippogriff··on The AI backlash is only getting started
If the labs weren't so aggressive with building datacenters in people's backyards, this could've been a different story. People don't like it when pipelines are built in their backyard either.
dippogriff··on The AI backlash is only getting started
They tried that a few times and the mistakes have had consequences.
dippogriff··on KinetIQ Ascend: Toward 100% Reliable Manipulation and Superhuman Speed
This is excellent! Very useful takeaways. Being able to properly do continuous training in production is key with robotics data being so hard to come by.
dippogriff··on Fixing Failures in Browser-Use Models: Why More Data Isn't Enough
Great work showing on how brittle these GUI benchmarks can be! Love the visuals.

I wonder if SFT is the problem here as opposed to the coordinate discretization; what happens with continuous action space?

dippogriff··on Autodata: An agentic data scientist to create high quality synthetic data
This is cool. Creative ways to do external verification is the only path to solving training on LLM slop
dippogriff··on Every match of the 2026 World Cup as a generative poster
Neat! minor nit - would be nice if the esc button took you back to the list, instead of having the click the X button
dippogriff··on Why eval startups fail (2025)
The current way benchmarks are done and are accepted by the community makes for really uninspired work. Until we're willing to break out of this rigid evaluation format prone to crazy overfitting and gaming, talent will move elsewhere. It is kind of a chicken and egg problem though.
dippogriff··on For Most of the World, Open-Source AI Is the Only Way Forward
Edge models will get much better after the current insane capex and organic data for pre-training is dried out. But hard to see how the best open source models will ever come close to the best closed ones.
dippogriff··on The worthlessness of Vitamin D is mildly exaggerated
Vice versa, the exaggeration of vitamin D is mildly worthless. Some need supplements, some don't.
dippogriff··on Qwen-AgentWorld: Language World Models for General Agents
I'm a fan of this direction. For me the most interesting use case for these world models isn't even training, it's verification. If this thing or some idealized version of it can actually reliably simulate state transitions, could you use it to verify an agent's execution path against hard constraints and replace/eclipse LLMs-as-a-judge?