12,148 karma · joined August 13, 2010
Working on diffui.ai - diffusion for UI design.
Formerly Figma, Atlassian, and Microsoft.
AMA about diffusion, design systems, and webcomponents!
+1 808 366 1708 j@diffui.ai x.com/pwnies
They are correct though that these days you can navigate away from the slop pages that gpt 5 put out pretty effectively, and land on something that really is quite decent. The page I made is a great starting point for 60s of effort on my part.
I am trying to make these webpage buildouts more legitimate benchmarks though, as I do think image->html flows are going to become more popular as people realize how good images as a starting point are. Hit me up if you have any feedback!
I had gemini describe this webpage as prompts, then fed it into an image -> code pipeline. Note that this is with 0 input from me:
https://html.non.io/railcode-animated
This is with Opus 5.5 building it out on low effort, and yes the spec was already made from the original site, but 0->1 that doesn't result in purple gradients or claude brown is very much possible now.
They've since taken down the 20x usage mention on their pricing page. https://archive.is/bCZav
The way I distributed cloud agents for this https://news.ycombinator.com/item?id=49687032 was via grok bot setting up Fable cloud instances.
GPT 6.1 Sol: https://html.non.io/lcars-gpt-6.1-sol
Opus 5.5: https://html.non.io/lcars-opus-5.5
Overall, opus executes a bit better than 6.1 sol, which surprises me. Astra has been the best model for this flow so far, so the fact that Sol missed some alignment / vision pieces here is interesting. It's not bad by any means, but I think where Opus really wins is the motion animation of the svgs / final polish (scroll down to the "customize every detail" section on the homepage, the svg animation is beautiful for that).
Still, it executed quick and was quite cheap to run.
1. Collaboration between always-on agents is a really, really powerful thing. It allows for domain-specific expertise that doesn't overload the context window, while still allowing for access to knowledge if they need it.
2. Domain-specific always on agents creates a good barrier of trust. One of the things I dislike about Claude is sometimes it's memory is all-encompassing. It's weird that it brings up things about my personal life when I'm talking about something related to my business. I've never had that happen with Grok Bot bots because I have one for my biz admin and one for my personal admin. They don't intertwine, which is quite nice.
3. Combined with cloud agents / cloud builds, things become really powerful for development. It was the first time that I felt there was a solution to the git worktrees / multiple streams at once issue. Each bot has its own computer and can spin up additional cloud agents. It comes at the cost of end to end speed - doing something via a grok bot often takes an hour end to end, whereas with a synchronous local prompt it'll take like 10min. The difference is I have to babysit one whereas the other "just works".
On the flip side, since using Grok Bots my inference spend has 2-3x'd. It's worth knowing that tradeoff. Nonetheless I think Luna is a fantastic driver for these, and OAI has very good pricing overall. I'd give these a shot - I think a lot of people would be surprised how helpful they are.
The side effect was I was fully cut off from my AI tools for those two weeks. I was coding "manually" during that time, and I think I accompished in two weeks what I previously had been able to do in a day. I'm not gonna lie, it was very, very stressful as a solo founder.
The industry moves so fast these days, that the only way to keep up with the speed is to leverage them. While I can appreciate the push of this to help your brain think independently/critically, the opportunity cost of a month of development without LLMs is too high a price to pay.
Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
All 3 were given the same prompt to dynamically light these and to create the designs as a SPA with page transitions.
Astra: https://html.non.io/annui-astra
Sol: https://html.non.io/annui-sol
Luna: https://html.non.io/annui-luna
Luna gets the button wrong, and in the same way Grok/MiMo did. Looking into it more, it's because Luna actually searched my computer for similar builds, found the ones that I did for grok/mimo, and referenced their files. Astra is still the best by a significant margin in my eyes. Far more polish, better page transitions, effects that aren't overcooked and take into account the page. Better contrast.
Worth noting though that GLM 5.3 isn't multi-modal, so it doesn't have a vision layer. It is quite clever and hacks around it pretty effectively however. I'm running a deepseek 4 build now and will reply shortly with that.
The gist of it though is I take a prompt, expand it into a json blob specifying structure/palette/positioning of elements/etc, feed that into a diffusion model to output a few choices. Once I lock in a choice I take the pixel output + json blob and use it as input into followup pages. The json helps preserve the brand across multiple pages.
Once I have all the inputs I take their corresponding image+json blobs and feed them into an agent to create a web implementation.
For image models, diffui currently uses gpt-image-2.5, mai-image-2.6, and very, very rarely a post-trained version of flux 2 dev I've made for web design, though that one will be deprecated soon.
Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
Opus 5.5's output: https://html.non.io/annui-opus/
Overall it follows image designs quite well, but it did ignore asks to animate page transitions. Additionally it's the least performant of the ones I've built with Astra/Grok/MiMo, despite using a lot of the same code. I'd rate it just below Astra in capability, but still solidly second place.
For comparison with other drops this week + current #1:
Astra: https://html.non.io/annui/
MiMo: https://html.non.io/annui-mimo/
Grok 4.7: https://html.non.io/Annui-grok/
Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
MiMo 2.6 Pro Ultraspeed (36min): https://html.non.io/annui-mimo/
Grok 4.7 (25min): https://html.non.io/Annui-grok/
Astra (19min): https://html.non.io/annui/
Overall this felt like the weakest of the three. Ultraspeed was fast as far as tokens per second goes, but it overthought quite a lot of things resulting it in having one of the longest build times. That overthinking didn't lead to better results either - note the statue with the cropped off arm. It's also the worst implementation of the dynamic lighting effect / displacement effect - the background especially has some significant distortion. Astra was the only one that seemed to understand that displacement should happen less the further something is in the distance.
Here's a vid of all 3 side by side with the source design: https://non.io/video/annui-comparison.mp4
Grok 4.7: $12.60
GPT Astra: $35.00
Designs: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
Astra's build: https://html.non.io/annui/
Grok's build: https://html.non.io/Annui-grok/
Additional prompt instructions: "Add scrolling clouds behind the statues. Dynamically light the statues based on mouse position. Use diffui to generate the normal maps/depth maps/roughness maps of the objects, and to separate out the assets on to different layers."
Overall I find these models are getting good at following image as a source of instructions, but their refinement of the output varies heavily between the models. Astra's final output feels more polished, has better visual contrast, and the animations between the pages are smoother. Grok also chose to light all of the background elements, which imo overcooks it a bit.
Still though, for the price it's a great starting point.
That internal json backing helps significantly when you want to maintain consistent design system components/patterns across multiple pages. The aligned layout is it working as intended.
https://html.non.io/qwen-comparison/
The text rendering definitely is much, much better than anything else on the open weights market right now. Small text fidelity is quite good. It seems like the text encoder however gets a little bit overloaded with larger prompts - note the presence of hex codes in the design output, those were inputs from the expanded prompt.
I'll be trying a post-training run on this for web design, it has some serious potential.
[1] diffui.ai
I could see it maybe as a one-time purchase, but $90/yr/user is somthing I'd never grab.
Don't get me wrong, I'm excited for a StarCraft shooter despite it not being StarCraft at all (I enjoy the universe). I wanted StarCraft: Ghost back in the day, and this feels like a taste of that dream.
That said, targeting a release of 2030 is wild to me. Look at the Astra game dev hype that's hitting twitter right now. Yes it's all filtered to the best possible examples / yes people are not one shotting these things, but still it's very clear that LLMs are beginning to be very capable at game dev in a way that doesn't look/feel like ass.
LLMs are going to become more capable over time. I think most of us can agree that eventually they'll be competitive with AAA developers given enough GPU cycles. I would argue that given rate of improvement, we're looking at that point coming in the next year or two.
To emphasize this point, a year ago GPT-5 got released. This was the type of game it would make: https://youtu.be/yTHo7tMborY?t=324
This is the type of game Astra is making: https://youtu.be/GuO_Eo34C8E?t=348
It's just starting to be capable of making assets in blender / unreal. Given this rate of advancement, what is the state going to be like 4 years from now? It really feels a bit like the "travelling to distant stars" problem, where at some point it's faster to wait for better engines than to leave now. I worry that by the time this gets released, the market is going to be flooded with fully custom AAA games that are hyperniche. 4 years out is a long, long time to wait, and it's a longer time to bet on success / market dynamics.
I had a bunch of extra Fable credits, so I spent about $10k running autoresearch loops on 200 of the top github repos with frontends to try and speed up the frontend performance. I distilled them down into a leaderboard of the most common wins, and built out an autoresearch loop that takes those learnings and applies them to your repo.
Works with your existing claude/cursor/codex sub in a cute custom TUI.
I just submitted a Show HN for it: https://news.ycombinator.com/item?id=49687032
Tailwind at this point is a coding language and a brand with no direct product. It's a great acquisition when you want mind share and developer love, and are OK with there being no revenue involved. Shopify is a solid match for that.