GPT‑6 and Intelligent UI for everyone
openai.com
openai.com
Given that OpenAI is making noises about merging work with chat (a horrible idea imo), and Work is very similar to Codex... I dearly hope things like these won't have any meaningful cross-polination into the actual work tools.
Seeing that the chat is based on 6.0 and not 6.1 is disappointing. The "Visual and interactive explanations" seems genuinely useful, but 6.1 is just so much better. I wouldn't truly trust the 6.1 with the explanations, but I'd trust them a fair bit more than 6.0. I understand that compute isn't infinite, but tons of people only interact with the chat and having your "things-explainer" be as good as it can be is important when people increasingly treat AI models as the source of truth, or even use them for academic learning and whatnot.
Still, the models will improve, so the visual explainer seems pretty good as an idea / mvp.
Prepare to lose essentially unlimited chat mode.
I imagine many will move to claude, as I will (return), unless anthropic makes more blunders.
Random link I found looking for the twitter post
https://pasqualepillitteri.it/en/news/21024/openai-merge-cha...
The only things that OAI still have over A/ are better coding agent GUI, no 5hr limit on >$100, much more reasonable cybersec guardrails (that allow most RE work) and... that's it.
The rest is basically unlimited unless you're doing something very weird.
Meanwhile I cant ask a simple question to Claude, not even to Haiku, because I used my claude code 5h limit.
Regular chat, even on extra high thinking mode, is sort of unlimited. Idk at what point you hit the abuse gaurdrail but its really high whatever it is.
E.g. codex reviews come from codex quota but also has its separate quota they don’t tell you about. Same with security reviews which has another quota still.
So the chat messages can be sent at several effort levels and from various models including 5.6 and 6 pro. The 6 pro messages are limited and they don’t tell you how many you’ve sent or how many are left. But the 5.6 messages seem to be unlimited, however there could just be a much higher unstated quota.
Kinda just feels like google search results in AI, which imo is a step down from distilled information. Chatgpt is already able to generate charts and visuals upon request.
That's never, ever the case with the AI generated ones and makes them useless.
Curious why that feels condescending? Like, the average person who is seeking a recipe needs a photo to know what to shoot for
The menu: Rosemary and garlic roast lamb, Extra-crispy roast potatoes, Honey-roasted carrots and parsnips, broccoli.
The photo: [1] Roast turkey, mashed potatoes, baby carrots, broccoli, brussels sprouts. No lamb, no parsnips.
If you shoot for what's in the photo, you're going to have a bad time.
You're also going to have a bad time when you try to make an apple crumble with no flour, no sugar, and no butter, because they're not on the shopping list.
[1] https://images.openai.com/static-rsc-4/gyzrX8zp3O2KLEPm0FAqh...
I feel like it can only be successful by luck. Either it's reproducing a recipe verbatim that was tried and validated and tasted by a human, in which case, we didn't need AI for that, just a searchable cookbook. Or, it's making one up. I understand that using RLHF has helped improve the quality of questions about history, programming, or TV show recommendations, but I do not believe that there has been some kind of training regime where people prompt the models for "a recipe that uses X, Y, Z random ingredients," follow the recipe, and then score it. And even if they do, I don't see how a model can learn enough from that besides "this exact recipe is good/bad." '1/2 tsp cumin' may be a great addition to one recipe and not enough for another, so improving the output based on a bunch of scored recipes... I just don't believe cooking is an LLM job.
Maybe some other kind of model that I don't know about.
I feel like the technique recommendations are probably the most valuable element, but I'm getting rave reviews like 85% of the time now?
There have been a couple times where I was out of an ingredient and it proposed an adjustment that sounded a little wild to me, but it almost always is well received.
Unfortunately, one does not exist.
The Internet, however, is full of garbage cooking advice, and while AI is quite happy to parrot that garbage, it's not significantly worse than it's sources.
The physical ones have indices, and the digital ones usually do as well (plus string search).
I often find their contents to be both incredibly broad, and incredibly shallow, and often lead to dishes that don't quite suit my tastes.
I have generally had more success with the internet than I have with them.
I usually just find recipes that are close to something I like and then modify them until I like them
I feel like cooking recipes should always be full of little notes from past attempts, to tweak them to your personal taste
And we aren't talking about something esoteric. How To Cook Everything. (4.5 rating on Amazon, 4.0 on Goodreads), one of the most popular cookbooks in the world...
Has its first recipe for chicken cutlets produce undercooked chicken (15 minutes at 325F has not resulted in 165F internally).
Worse yet, its chile recipe for some weird reason involves boiling and simmering a whole onion with the beans before throwing it out (WTF). It then has you drain the beans only to add that drained water right back (WTF?), unless you optionally replace it with tap water (WTF was the point of boiling the onion then?), so that you could bring them to boil again. The recipe barely has any spice besides actual chile peppers.
This... Is not good or helpful, but is incredibly opinionated with a bunch of bullshit steps. Like, yes, I could tweak that into something that's not insane, but why would I choose that as a starting point..?
I get it. Compiling a thousand-page cookbook is really hard. But this... Ain't great.
Fuck. Does nothing good survive on the internet anymore.
I mostly don’t even need a recipe from AI, just the title and a 1-2 sentence description.
For me, the text-only recipes I already get out of ChatGPT today are good enough
With an image of sour grapes
The user is probably in a grocery store. They need to buy the stuff first. Showing them a picture of a finished product is an irrelevant distraction.
Then, if there is some arguably relevant content you can tack it on afterwards.
“My made up scenario is definitely more realistic than your made up scenario.”
Who goes shopping for the meal before they even have a number of guests?
I'm more worried about picture accuracy: usually the benefit of pictures in recipes is to see what you're aiming at (eg, how finely chopped something is), but I don't know if the image model is up to that level of detail.
As a highly literate person, it is easy to overestimate the share of the population that is highly literate.
Yes, pictures OF THE FOOD. Not unrelated pictures of another lamb-roast made with another recipe.
Good cookbooks do proper food photography, of the food made with the recipe. That's what sets them apart from slop, be it human or AI slop, cook books.
For almost everything else we can create them - just give a reasonable cook a kitchen and the ingredients and they will do it. For a picture we then have an artist do preparation to make it look good, but they should always start with the real dish.
My most used prompt in the last week is probably "Explain this concisely and simply, like I am a child".
I spent years dumbing things down and creating visuals to support it for decision makers. I am frequently asking ChatGPT to do the same for me.
mine is “explain this 2 me in simple terms”
All those complaints about Claude language use make me think we're finally seeing the consequences of a generation growing up on instant messaging.
If I say Claude talks like a meringue it's patently obvious what I'm trying to convey, sweet but full of empty space. If you hear that same metaphor day-in day-out then after a while it doesn't call up a thought at all, it's just an annoying verbal habit.
My go to phrase that I’ve found to work well is “please restate in high school English with direct declarative sentences and no parentheticals or references that require recalling previous turns of this conversation”
Which I have as a hotkey.
Explain it to me like I'm a golden retriever.
> ASD-STE100 Simplified Technical English (STE) is a controlled natural language that is designed to simplify and clarify technical documentation. It was originally developed in the 1980s by the European Association of Aerospace Industries (AECMA) at the request of the European airline industry, which wanted a standardized form of English for aircraft maintenance documentation that could be easily understood by non-native English-speakers.
I've taken a liking to pointing agents at the Simple English Wikipedia editorial guidelines...but so much of that is for making up for shortcomings of modern Anthropic models. Codex running 6.1-sol explains things so much more clearly that I don't feel I need to assert a style guide on its output.
GPT stays too much in text imho.
- Gemini 3.8 has a beautiful answer. Typography, use of lists, brevity, style, even the font choice on the web harness -- all of it is A tier or even S tier. Notice I didn't say answer quality. The answer is always extremely mid, and the harder the question (or more effort needed to answer it), Google just quits. I give it a "C" on answer quality
- ChatGPT (Chat, pre 6, I assume 5.6 although they commonly hide the model picker): Ugly answer. Runs on printing page after page of unnecessary side notes. Mostly just walls of texts in paragraphs with little thought to information design. However, inside of that mess of text is almost always the answer I'm looking for, and it's almost always incredibly better than the truncated, desgined Gemini answer.
For a while I ran all of my question queries on Gemini and ChatGPT, but I noticed that I basically always picked ChatGPT's answer head to head.
I even found it was capable of stepping back from code implementation and bug hunting to “hang on, it’s not the algorithm that is incorrectly implemented, it is the wrong algorithm for the goal.” GPT Sol (5.6,6 and 6.1) just kept hammering in the problem.
That's a very high bar.
And you see this in the other examples too. The teach the CLT piece was funny. It's and error-ridden mess that isn't even an explanation at all!
The best part is, someone looked at that and thought yes, that's right. Shows you clearly why programmers will be needed in the future. The LLM might be able to run circles around that person in math, but that person also has no idea when they're being bullshitted.
Sometimes one needs a top comment like this to remind them how out of touch the average HN voter is.
The fact we've come so far with textual models & the image stuff is still producing stuff that harks back to early-stage tripophobic AI produce (the garlic with roast potatoes here) is interesting.
Don't get me wrong, I'm certainly starting to see some AI-produced imagery today that I can't tell isn't real, but that's largely because a lot of real photography is overproduced & ugly. I have yet to see anything that's aesthetically good.
In all seriousness, his lovingly and expertly crafted explainers are still going to age like a handcrafted heirloom clock in a world of plastic-clad quartz movements. But it’s absolutely incredible that we are now in an age where a computer can manufacture a serviceable interactive explainer on whatever niche topic you desire.
> So too, is his turn to be relegated to a relic of his time.
No, it'll be tasteful artifact, not a relic. Records, fountain pens and automatic watches did not die. They are used by people who discern things, and no, none of these things have to be expensive (i.e. Neither Seiko 5, nor Lamy Safari are expensive, yet they are as dependable as their 100x expensive brethren).
Human touch still has that finesse and warmth.
On the front page right now - https://news.ycombinator.com/item?id=49980626
OTOH, I still believe the exploded view on https://ciechanow.ski/mechanical-watch/ is something else.
For one, it has real physics on the weight, and second it always shows the correct/current time.
FWIW, his all animations has proper physics to begin with.
The entry you posted is nice, but Ciechanowski is still peerless.
It’s all just surface level complexity with no intention behind it. A clumsy approximation at best.
You call it serviceable, but what purpose does this service? What do you now know about 7-speed bicycles that you didn't know before?
If the same explanation works regardless of whether the bicycle has 1, 7, or 21 gears, then the model probably hasn't understood what needs explaining.
Amazing work. Always excited for the next update. This was a human driven and created success.
- Audio volume has 2 levels: on and off
- Play button worked exactly once for me: it played and looped the video. Couldn't be stopped afterwards.
- The video progress bar has no visual indication of where it starts and where it ends..
Is this what Phind died for?
Even TV channels like CNN and FoxNews don't have a decent one.
It's like we need ASI to have a proper embedded video player. The ultimate software challenge.
It's just most of our industry pretends they don't exist, and instead implement toy-like video players to distinguish themselves, I think.
Sure as hell beats stupid Instagram style videos where you have no way to skip ahead. I think this is what future video is going be like, buckle up.
> all audio content ever has its loudness perfectly normalized to a global standard everyone adheres to.
A volume slider on a video player is a really basic feature that every one should have.
Maybe you and TeMPOraL do.
Not to mention, maybe a long-running process will use some audio cue as a notification while you're listening.
(Actually, given the rest of the comment, I'm pretty sure TeMPOraL was being sarcastic.)
Most things these days, good or bad, are not designed for power users.
Ironically many here I'm sure are running MacOS.
System card linked in the blog post.
> regression on the extremism vision evaluation.
> Relative to their respective GPT-5.6 counterparts, GPT-6 Sol (October) shows a statistically significant regression on standard self-harm, while GPT-6 Luna (October) shows statistically significant regressions on standard self-harm, gore, and sexual content
> "GPT-6 Sol (October) and GPT-6 Luna (October) show an improvement on helpfulness on legitimate requests relative to prior models, though it scores lower on some safety requests"
Concerning how there's significant regressions on so many critical benchmarks, but it is newer and creates UI, so must be good.
It’s crazy that a company can document these safety drops in a PDF, ship the model anyway, and focus the announcement entire on shiny new UI
Good. Maybe I'm the only one, but I feel like these have always been stupid measures to waste the finite time of "AI safety" research on anyway. They amount to whether a user can, if determined, manage to make a chatbot say, or depict visually, some taboo thing. Frankly I'd rather just have some cheap classifier judge each chat response after it's generated, with a limited context "Is this response encouraging self-harm?" and skip all the rest of these. Whether some random pervert can make GPT-6 spit out an erotic fanfic story or image has zero impact on the rest of the world. If this particular model won't do it, other models already exist that will, and the determined thoughtcriminal could always just write the forbidden words themselves, or photoshop something taboo.
Maybe the term "safety" has been usurped by those who feel that 'safe from the possibility of being offended' is the most important kind of safety, but it's like worrying about the wallpaper on the Titanic, compared to actual AI safety concerns.
Most "bad outputs" in areas like 'self harm' or 'violence' or 'sexual' whatever are a result of people deliberately trying to elicit those responses. As such, it's meaningless whether an LLM writes the Bad Thing or if the user writes it himself.
In general though you do have to figure that vulnerable and naive people will use it because of the degree of market penetration they're aiming for, including minors etc. and they have some responsibility around that.
> To mitigate the risk of producing disallowed responses for teens, we apply an additional classifier-based block to responses that may contain self-harm, sexual content, and gore; this mitigation is not captured in the evaluation results above and improves safe responses.
Right... women will oppose AI because it's too hot and not because gross chuds have been using it relentlessly to create deep fakes of their likenesses and CSAM of their kids.
This is actually a defensible position. It's not either-or - the "gross chuds" doing things you mention are worrying, too. But providing unrealistically attractive depictions of a human body is also a problem. It's been a problem for ages - all the girls and women who try to compete in beauty with celebrities and models (without a team of specialists supporting them), only to ruin their bodies through anorexia or bulimia, or their minds with depression, deep insecurity, and an inferiority complex, exist and deserve mention. AI(-generated images) is not the reason, but fits "well" into the preexisting social problem, and has a chance of intensifying it on a global scale.
I've been learning music lately and it kept re-pasting the same one chord visualization throughout many conversations, almost randomly and often barely related to the question. So I at least hope this won't be as aggressive so I can prompt it away!
I used to edit my prompts and undo messages if it misunderstood something but that's been failing me recently.
That's my biggest gripe with it. If I ask a quick throwaway question, I'd really rather not wait for the LLM to build a test framework for its composable principles-driven components framework and WebGL/WebGPU abstraction first.
AI: compiling C code...
What's funny about "Intelligent UI" is I said like 2 or more years ago, that these AI companies need to start thinking outside of these basic chat UIs, they do some things here and there, but its really depressing how little they do to innovate in these spaces. Same with the coding harnesses, the UI for all these things could be drastically superior.
Also reasonable: it's possible to hire a world-class ___ to do that job better than AI.
Anyone can having blog posts as good as this by using AI. Just not by zero-shotting it with a 2 line prompt. It still takes a lot of work.
There aren't going to be enough experts to go around. They'll start demanding millions in salary and all of a sudden the advantage is gone.
They use their product to do what it’s best at, and use humans where it isn’t yet good enough.
But this is just a marketing page for some tech company, right? If only it stopped there. I have to deal with this nonsense in news "articles" as well on occasion, when some web designer intern is allowed to larp as a journalist for a day.
sigh Just give me text to read.
This is why I use an IDE instead of a TUI agent... I'm on a computer with a multi-megapixel display, I want to use it. If I could be driving the same process with my Macintosh SE as a serial terminal, what the hell is the point of my recent Macbook?
I should have been more precise than saying "just give me text", but it was what came in to mind when I wrote it, as I was thinking about what an article is meant to contain as its base element.
What changes could they make to really improve the usability of their products?
Fairly likely: https://medium.com/the-engineering-brief/openai-hired-400-ap...
I would prefer if the AI companies stuck to just creating better models and making them as cheap and accessible as possible. Let others build the products. I don't want 1-2 companies to own every product in the world.
That seems to be the plan...
Buckle up boys
Look at the history of Microsoft.
So this is a taste of the slop that's going to invade everywhere in a few months.
If the main user paradigm becomes a chat interface conjuring whatever UI elements needed to best accomplish the task at hand, the entire paradigm of OS, UI frameworks, apps from an App Store to accomplish specific tasks, all come into question.
I keep thinking that Steve Jobs would have been all over the UX ramifications of LLMs and demanding that Apple lead in that area.
Oh give it a rest. Lol who are you compared with the management of Apple?
This preachy stuff is ridiculous. So far Apple have been correct to stay out the LLM space.
Apple absolutely messed-up by abandoning OpenCL's early ML research efforts to oppose a CUDA monopoly. They messed up a second time shipping a raster GPU with Apple Silicon when CUDA and Tegra had proven that GPGPU was mobile-ready. Then when the ARM datacenter had it's moment, Nvidia's Grace ARM CPU displaced billions of dollars in sales that would have been Apple's if they didn't mess up the fastest CPU in the world with macOS. Then they depreciated the Mac Pro, which seems like a mistake since it had the potential to outsell Nvidia's ARM datacenter chips if it ran Linux. According to the rumor mill, Apple's M8 chip will finally be the one that takes GPGPU seriously, after a decade of Apple's innovative ship-second mentality.
Apple's management isn't infallible whatsoever. Many of them are petty, blinded by politics and/or obsessed with their legacy more than they care about profits or quality products.
OS on Demand
I have explicit instructions to tone it down. Less taglines, eyebrow text, subheadings, decorative spacing, pills, cards.
Hopefully this doesn't bleed into the chat...
It's a pretty clever way to sell more tokens, I have to admit
It's SEP-1865: https://modelcontextprotocol.io/seps/1865-mcp-apps-interacti... and https://github.com/modelcontextprotocol/ext-apps
But if you have any ambition at all---if you want to map the local school board, or get a quick understanding of your finances, or set up a robot in your backyard---then, quickly proves useful.
My long-term hope (and call me another Zitron if you wish) is that the labs' hype dissipates if there's no serious improvements beyond stringing together agents to ram through brick wall Millennium Prize style problems and we all just use on-device AI for 99% of cases where it's useful, researchers, governments, militaries, and universities can pay for the more complex models, and image/video gen dies a slow death (if Congress will actually legislate and/or SCOTUS decides that training those specific models does not qualify as fair use unlike training LLMs) and cost for compute rises over time as public interest in the tech sours.
LLMs are such marvelous and incredible technology squandered by genuinely deranged Silicon Valley cultists who are trying to use them as a Trojan horse to force their antisocial visions of the future upon an unsuspecting populace. We could all have just invested in on-device compute and nobody would have lost their jobs in pointless layoffs, productivity would have increased, and maybe we would have forced a conversation about when and when not to use AI and it wouldn't be so ever-present in use cases where it actually does harm to the consumer and society. But sama, PT, and dario just wouldn't have it that way, would they? Because if that were the case, there's no prospective hope where they can assuage their deep-seated insecurities over being antisocial and off-putting by reassuring themselves that there's an imminent realignment of society where they will hold all the cards when the dust settles.
All the current-generation models can oneshot HTML/Javascript frontends of comparable complexity (though maybe not as polished). The only difference is that you have to specifically ask them, and the result is not embedded in the chat.
It feels like it should be easy to have a harness provide a "show HTML widget in the chat" tool and a skill and/or system prompt that instructs the model to generate such widgets on-the-fly if the response could benefit from interactivity.
What am I missing here?
I eventually just had it make a skill for work mode that would build a mini checklist app, but it was slow and felt janky needing to remember to switch to work mode.
This is already such a huge improvement.
I actually do think generating UI has some future, but I remain sceptical it can realistically go as far as promoters claim. It is especially bizarre to me when this is done by ostensibly UX people.
What I actually doubt is that people as whole are this adaptable to continuous novelty, especially when it is not directed by them.
As a developer, it already feels like this with internal company tooling. Anything I have the source code to and a build pipeline setup for, I can just pop open a new tab and tell it what I want changed, and it just gets done for me.
Even my home assistant setup at home now, i hooked up codex to it via an MCP, and when we put out the blow up haloween lawn ornaments I was able to just pop open codex on my phone and tell it to throw together an automation to run them for me, and 5 minutes later it was done, with an override button on my dashboard. It's a shockingly nice workflow!
Wouldn't it be nice if you could just tell your app to reconfigure/rewrite itself to remove those "what the app will think I need at particular moment" features that you dislike? Maybe have it remove a bunch of the whitespace, getting rid of the lazy loading infinite scroll, whatever you want. It's the GPL dream, except you don't need any coding ability, just a subscription to a megacorp.
Agents cut through all that crap and allow me to make changes without learning a new overwrought configuration language that thrashes syntax every few months. The fact that it is open and fully modifiable redeems it, because I can ignore all the annoying parts while achieving what all that was envisioned to allow.
Just like Codex failed so would that.
And I don't need to build or debug complex automations any more! I can just tell some AI to make it so it auto-locks the house door when my phone isn't tracked at home, but also to check the wifi and if there are people using the guest wifi then to instead send me a notification asking if I want to lock the house because guests might be there.
1) I want a quick answer, and I don't care for the boilerplate UI. For example, if I ask how to make pancakes, I make them all the time and just want to a quick reminder on the ratios, but it might trigger a full UI that I need to sort through to find information.
2) If I ask a to me unrelated to UI question and it triggers a big UI build that is completely off topic for my question (meaning I'm desperately pressing the stop button and prepping rewriting my query)
Or the other one I see, for example if I look up a unix command like:
"ls all hidden files in the /xxx directory"
And I get back:
"Sorry, I am am unable to find /xxx in my current environment"
At least seems a fist step out of the serial interface of chats.
People make simple text websites → Google comes along and indexes all websites → People can freely and easily find information! → Google slowly perverts the incentives with ads → Websites replace simple text with complex UIs and paywalls → People can no longer easily find information → OpenAI indexes all these complex websites and turns them into simple text answers → People can freely and easily find information again! → OpenAI perverts the incentives with ads → OpenAI replaces simple text with complex UIs and paywalls → People can no longer easily find information.
And so on…
Getting a UI that explains the parts of a bike is ok I guess, but isn't it simpler to get an actual breakdown? Google "parts of bicycle breakout" gets tons of useful images instantly.
Getting a specialized app to split a bill? It was already trivial to put in a calculator if we cared to go item by item on the bill. Having to provide names and tag every item as I go is just more work. Usually real people just go $total divide by 5, I had more, let me chip in an extra $10.
Same with booking travel and wedding plans, these aren't things people would even delegate to a trusted friend usually, much less a one-off request to an AI bot or custom UI.
Most of the useful tasks it can do right now are research, technical question/answer, coding. In terms of "build a flexible ui that solves real-world problem", if existing mobile app isn't useful in this arena, then its unlikely a completely custom UI will do the job.
That said, I've done a few small ones like a quick one to practice alphabet of a foreign language, or prototyping a web game, and the like. But ultimately its nothing that is worth trillions of dollars.
"ChatGPT is where people start with AI, with more than 900M weekly active users, and we now have more than 50 million consumer subscribers."
Feb '26 https://openai.com/index/scaling-ai-for-everyone/I've always found his work impressive, but for the most part it was all prototypes and proofs of concept.
Having such interactivity available to everyone now is nothing short of groundbreaking and another indication AI works as the great equalizer by democratising capabilities previously available only to a select few.
URL-->user request--> answer & visualization.
No more static menus, buttons, lists, etc.
There's probably some good work out there helping small businesses adapt to this change? If done right, the real-time token use to generate dynamic pages can be minimized.
And now we are asking LLMs to solve it.
Is this supposed to be humor that I'm not getting, or are they really saying, "look at this solution to a problem you would have if you did normal things in the worst possible way"? It almost feels like the video is trolling us.
This is now trying to steal YouTube repair and cooking videos. Here is news for you: People prefer YouTube repair and cooking videos.
But getting actual service manuals cheap/free these days can be tricky, and their quality can leave a good bit to be desired.
Not sure about the cooking one but repair videos are horrible. People with thick accent, the worst camera, shittiest lightning and bad angles where you can't see what you have to do.
Gettin closer to fully dynamic interfaces for a lot of software. Hell, give me a mode in Google Docs that takes every pixel of chrome away then vibe the rest as I need it. Persist across new documents going forward.
Make it open or it is not for everyone.
To that front - model providers generating visualization elements (i,e taking over the tools for interaction and visualization) is just a natural next step in value capture.
Anthropic launched docs, presentations, sites and am sure they have sheets next up their alley. Apps are artifacts. Tool forming is a natural model adjacency.
Currently, we have highly specialized apps which do exactly what they need to do. They were created by someone who cares about making the program which works right and reliably. Most programs have years of bug fixes and deterministic algorithms. Many of them run locally and waste a minimum amount of energy to do the calculation.
But this is different - this encourages you to make one-off apps for each little thing you want to do. Imagine someone making a "bill splitter" app every time they go out with friends and how much energy that would waste... Just to do "14+16" that you could do in your head.
Also imagine it coming up with new UIs every time (depending on the prompt/context), and even potentially gaslighting you and doing calculations wrong because no one ever tested it before.
Like I said, it’s too subtle a deliberate self-own.
seems a bit like a snake eating its own tail
Unsure if this has been done already but the best thing I did was get it to create a list of next steps as quick buttons that it keeps updated in a docked card. It works well but occasionally gets stuck on some un related rabbit hole.
why not just stream html?
For elements that manage layout and drawing things like these explainers, though, it's hard for me to imagine how that would work. My immediate thought would be to an abstraction over it, which is what it sounds like they did.
What is it supposed to demonstrate? That the model knows some kind of folk mereology?
I suspect it will struggle with anything that isn't simple enough for demo-mode or lifestyle fluff.
Remember when Steve said ‘the computer for the rest of us?’ We’re seeing that here.
Today i wanted to reencode and downscale a screen capture video to make it small enough to send to other people. I had a chatbot generate me the ffmpeg command line, but when will i be able to tell "Apple Intelligence" or some other "Intelligence" just "resize this to 720p, make it mp4 and drop the audio"? You know, like in Blade Runner.
Edit: oh wait, Samsung already did it like in Blade Runner, when they "enhanced" the moon photos until they had details that were impossible to capture with those optics...
"Here is a library of [svelte/react/whatever] components, use them to construct a helpful visual to demonstrate your point."
The deconstructed bike at the beginning was in a class of its own, however.
Imagine a world where you get paint chips with no color! Then you could use our product (you could also just look up the name of the colors).
- People don’t have to stick to a single provider. They can all claim the same 20%.
This is false.
Edit: be sure to check Claude settings -> Capabilities -> Visuals.
I guess the above is mostly mobile focused.
Obviously if your model is stuck inside a CLI terminal, then not so much. But in a GUI harness (shameless plug for my own one: https://juggler.studio, but I assume others can do this too), you just ask them to answer in HTML and they'll happily draw pretty pictures inline in the conversation. I've been doing this for ages with claude, GPT, Deepseek and others.
Supposedly, people were struggling to follow written manuals and they needed an interactive explanations.
I'm not a mathematician and I would certainly love having a tool that would do ELI5 on some complex stuff, but I'm really worrying about using this too often and outsourcing my ability to do stuff to some mega corp.
I also believe that a small talented group of developers just out of college can do just as good a job. That is why there is no moat around AI, it still comes down to original thinking and talent.
> The Check, Please.
> TABLE OF FIVE · ITEMIZED BILL
This is one of my _least_ favorite AI behaviors, I am surprised they left it in the example. This behavior where it insists on putting text everywhere that over-explains the context.
> you stupid silly human, instructions are hard! Who can read "fold this" and "fold that", when you can instead have interactive short-form content with music and smileys to keep you entertained while you offload all the thinking to AI?
... as a service?
> Also pick up garlic, rosemary, thyme, lemons, honey, almonds, olive oil, gravy ingredients, mint sauce, crumble topping and vanilla ice cream.
what??? how the hell are you supposed to remember all that. Even the before example tells you what kind of potatoes to get. tbh it would be cool if you could just click a button and get an order pre-filled out on a grocery delivery/pickup service. I really wonder who reviewed this post and if they cook, because imagining yourself in that scenario and reading those instructions falls apart very fast
With computer use Both Astra and Sol have been able to make sense of my markdown recipes, and based on those place prepare a shopping cart for me to review and finish buying
Probably a bit much for the average user though, even if it's trivial to setup
It seems like every week OpenAI adds some new feature to the chat UX that wasn't tested, doesn't work at certain resolutions, doesn't work on some platform/browser combos and hogs all the memory/CPU.
Intelligent UI my ass.
Edit: I should add, features that no one asked for, too.
First thing to ask is a design teardown of AI Slop images into individual artistic choices.
Jokes aside: always having the right UI on a dataset is a very cool promise and this looks like a leap in that direction.
Now they extracted the value from interoperable open systems, they would love to replace it with closed proprietary systems.
Or, more probably IMHO, most people will keep using the standard UIs because they are ready and they need no work to build.
yes you and everyone else.
How do sites like Wikipedia get updated? What if I need to upload important documents to different portals? It all has to go through the AI companies? They have access to medical records too? My taxes, social security, etc?
Is this how Microsoft imagines Office 365 and Sharepoint will be merged into? Disney, NY Times, Youtube, Netflix will all be fine with chatbots handling their content?
Ok so this is how you build a moat. Models are not a commodity and visualization isnt a purely harness problem.
> We expanded our training methods to help the model make thoughtful decisions about content, layout, visuals, and interaction. This included evaluating the interfaces it creates for clarity, usefulness, and completeness. GPT‑6 learned to use the component library and make good design decisions, including how to organize information clearly, when to use interactivity, and when a simple text response is enough.
Curious to know how they trained it produce appropriate visuals. And why that cant be done with propmt engineering.
extraordinary AGI claims require extraordinary evidence
Most apps that people use are really very similar and don't have anything original. It would make more sense for it to be more personalized. For example someone might prefer not to interact with interface at all, and just access services just by chatting with a bot, while other person might want only last steps to be provided as UI. For example to see summary of his cart before paying. While another person might be insterested in just browsing all options. Some might prefer to have filters, while other would want AI to filter everything for them.
I think the real result would be when all services would be automated by AI. Like if you want to be a small business, you no longer need to build anything. You just describe what real life services can you provide: delivery, barber, baker, cleaning, repair. Then all of that would be in some agent network, and available to other agents as an option. While humans on both sides would just get a bridge between them in whatever form is most comfortable for them. Buyer might get a list of bakeries as normal ecommerce website, while the baker might just be someone who receives phone calls explaining him what his next order is in human voice.
It's time to promote an .htmd HTML document standard so that we can actually have nice looking documents again. Remember columns? Colors? Typography? Layouts? Graphic design? Interactivity?
Just think - we could actually have a true rich text standard! Imagine if it were adopted everywhere and you could send WYSIWYG bold, italicized or underlined text as easily as you can a custom skin-colored emoji?
It would almost be like were living in the 21st century again!
I used to see it a lot in China, like the idea of pinching something on one phone and transferring it to another, or identification by showing your palm.
It's also the type of UI you would see in Hollywood movies, both because it's stuff that laymen scriptwriters and directors would find interesteng, but also the audiences would. Like the motion based UI in Minority Report, whatever was in the movie hackers, or that laptop suitcase thing that allows you to deploy nuclear codes or send wires.
I mean, I don't want to be a snob, clearly people like it, at least initially, I think it's outsider art of sorts, of course it will have problems, but it's not less legitimate because of that.
Anyways, this Intelligent UI thing where you ask it how a bike is made and it unravels a hollywood visual report as if it were showing the interiors of the target in a war room conference, it reminds me of that type of futurisitic outsider UI.