4o Image Generation
openai.com
openai.com
Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on.
You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff like "change day to night", or "put a hat on him", and so forth.
I get the feeling these models are quite restricted in resolution, and that more work in this space will let us do really wild things such as ask a model to create an app step by step first completely in images, essentially designing the whole app with text and all, then writing the code to reproduce it. And it also means that a model can take over from a really good diffusion model, so even if the original generations are not good, it can continue "reasoning" on an external image.
Finally, once these models become faster, you can imagine a truly generative UI, where the model produces the next frame of the app you are using based on events sent to the LLM (which can do all the normal things like using tools, thinking, etc). However, I also believe that diffusion models can do some of this, in a much faster way.
I do not think that this is correct. Prior to this release, 4o would generate images by calling out to a fully external model (DALL-E). After this release, 4o generates images by calling out to a multi-modal model that was trained alongside it.
You can ask 4o about this yourself. Here's what it said to me:
"So while I’m deeply multimodal in cognition (understanding and coordinating text + image), image generation is handled by a linked latent diffusion model, not an end-to-end token-unified architecture."
>"So while I’m deeply multimodal in cognition (understanding and coordinating text + image), image generation is handled by a linked latent diffusion model, not an end-to-end token-unified architecture."
Models don't know anything about themselves. I have no idea why people keep doing this and expecting it to know anything more than a random con artist on the street.
See this chat for example:
https://chatgpt.com/share/67e355df-9f60-8000-8f36-874f8c9a08...
While LLM code generation is very much still a mixed bag, it has been a significant accelerator in my own productivity, and for the most part all I am using is o1 (via the openAI website), deepseek, and jetbrains' AI service (Copilot clone). I'm eager to play with some of the other tooling available to VS Code users (such as cline)
I don't know why everyone is so eager to "get to the fun stuff". Dev is supposed to be boring. If you don't like it maybe you should be doing something else.
Please sir step away from the keyboard now!
That is an absurd proposition and I hope I never get to use an app that dreams of the next frame. Apps are buggy as they are, I don't need every single action to be interpreted by LLM.
An existing example of this is that AI Minecraft demo and it's a literal nightmare.
While I think current AI can’t come close to anything remotely usable, this is a plausible direction for the future. Like you, I shudder.
> “DLSS Multi Frame Generation generates up to three additional frames per traditionally rendered frame, working in unison with the complete suite of DLSS technologies to multiply frame rates by up to 8X over traditional brute-force rendering. This massive performance improvement on GeForce RTX 5090 graphics cards unlocks stunning 4K 240 FPS fully ray-traced gaming.”
"Draw a picture of a full glass of wine, ie a wine glass which is full to the brim with red wine and almost at the point of spilling over... Zoom out to show the full wine glass, and add a caption to the top which says "HELL YEAH". Keep the wine level of the glass exactly the same."
Using Dall-e / old model without too much effort (I'd call this "full".)
For Gemini it seems to me there's some kind of "retain old pixels" support in these models since simple image edits just look like a passthrough, in which case they do maintain your identity.
Got it in two requests, https://chatgpt.com/share/67e41576-8840-8006-836b-f7358af494... for the prompts.
That sounds really interesting. Are there any write-ups how exactly this works?
> The system uses an autoregressive approach — generating images sequentially from left to right and top to bottom, similar to how text is written — rather than the diffusion model technique used by most image generators (like DALL-E) that create the entire image at once. Goh speculates that this technical difference could be what gives Images in ChatGPT better text rendering and binding capabilities.
https://www.theverge.com/openai/635118/chatgpt-sora-ai-image...
The general gist is that you have some kind of adapter layers/model that can take an image and encode it into tokens. You then train the model on a dataset that has interleaved text and images. Could be webpages, where images occur in-between blocks of text, chat logs where people send text messages and images back and forth, etc.
The LLM gets trained more-or-less like normal, predicting next token probabilities with minor adjustments for the image tokens depending on the exact architecture. Some approaches have the image generation be a separate "path" through the LLM, where a lot of weights are shared but some image token specific weights are activated. Some approaches do just next token prediction, others have the LLM predict the entire image at once.
As for encoding-decoding, some research has used things as simple as Stable Diffusion's VAE to encode the image, split up the output, and do a simple projection into token space. Others have used raw pixels. But I think the more common approach is to have a dedicated model trained at the same time that learns to encode and decode images to and from token space.
For the latter approach, this can be a simple model, or it can be a diffusion model. For encoding you do something like a ViT. For decoding you train a diffusion model conditioned on the tokens, throughout the training of the LLM.
For the diffusion approach, you'd usually do post-training on the diffusion decoder to shrink down the number of diffusion steps needed.
The real crutch of these models is the dataset. Pretraining on the internet is not bad, since there's often good correlation between the text and the images. But there's not really good instruction datasets for this. Like, "here's an image, draw it like a comic book" type stuff. Given OpenAI's approach in the past, they may have just bruteforced the dataset using lots of human workers. That seems to be the most likely approach anyway, since no public vision models are quite good enough to do extensive RL against.
And as for OpenAI's architecture here, we can only speculate. The "loading from top to be from a blurry image" is either a direct result of their architecture or a gimmick to slow down requests. If the former, it means they are able to get a low resolution version of the image quickly, and then slowly generate the higher resolution "in order." Since it's top-to-bottom that implies token-by-token decoding. My _guess_ is that the LLM's image token predictions are only "good enough." So they have a small, quick decoder take those and generate a very low resolution base image. Then they run a stronger decoding model, likely a token-by-token diffusion model. It takes as condition the image tokens and the low resolution image, and diffuses the first patch of the image. Then it takes as condition the same plus the decoded patch, and diffuses the next patch. And so forth.
A mixture of approaches like that allows the LLM to be truly multi-modal without the image tokens being too expensive, and the token-by-token diffusion approach helps offset memory cost of diffusing the whole image.
I don't recall if I've seen token-by-token diffusion in a published paper, but it's feasible and is the best guess I have given the information we can see.
EDIT: I should note, I've been "fooled" in the past by OpenAI's API. When o* models first came out, they all behaved as if the output were generated "all at once." There was no streaming, and in the chat client the response would just show up once reasoning was done. This led me to believe they were doing an approach where the reasoning model would generate a response and refine it as it reasoned. But that's clearly not the case, since they enabled streaming :P So take my guesses with a huge grain of salt.
I built this exact thing last month, demo: https://universal.oroborus.org (not viable on phone for this demo, fine on tablet or computer)
Also see discussion and code at: http://github.com/snickell/universal
I wasn't really planning to share/release it today, but, heck, why not.
I started with bitmap-style generative image models, but because they are still pretty bad at text (even this, although it’s dramatically better), for early-2025 it’s generating vector graphics instead. Each frame is an LLM response, either as an svg or static html/css. But all computation and transformation is done by the LLM. No code/js as an intermediary. You click, it tells the LLM where you clicked, the LLM hallucinates the next frame as another svg/static-html.
If it ran 50x faster it’d be an absolutely jaw dropping demo. Unlike "LLMs write code", this has depth. Like all programming, the "LLMs write code" model requires the programmer or LLM to anticipate every condition in advance. This makes LLM written "vibe coded" apps either gigantic (and the llm falls apart) or shallow.
In contrast, as you use universal, you can add or invent features ranging from small to big, and it will fill in the blanks on demand, fairly intelligently. If you don't like what it did, you can critique it, and the next frame improves.
Its agonizingly slow in 2025, but much smarter and in weird ways less error prone than using the LLM to generate code that you then run: just run computation via the LLM itself.
You can build pretty unbelievable things (with hallucinated state, granted) with a few descriptive sentences, far exceeding the capabilities you can “vibe code” with the description. And it never gets lost in its rats nest of self generated garbage code because… there is no code to in.
Code is medium with a surprisingly strong grain. This demo is slow, but SO much more flexible and personally adaptable than anything I’ve used where the logic is implemented cia a programming language.
I don’t love this as a programmer, but my own use of the demo makes me confident that programming languages as a category will have a shelf life if LLM hardware gets fast, cheap and energy efficient.
I suspect LLMs will generate not programming language code, but direct wasm or just machine code on the fly for things that need faster traction than they can draw a frame, but core logic will move out of programming languages (not even llm written code). Maybe similar to the way we bind to low level fast languages but a huge percentage of “business” logic is written in relatively slower languages.
FYI, I may not be able to afford the credits if too many people visit, I put a a $1000 of credits on this, we'll see if that lasts. This is claude 3.7, I tried everything else, a claude had the visual intelligence today. IMO this is a much more compelling glance at the future than coding models. Unfortunately, generating an SVG per click is pricey, each click/frame costs me about $0.05. I’ll fund this as far as I can so folks can play with it.
Anthropic? You there? Wanna throw some credits at an open source project doing something that literally only works on claude today? Not just better, but “only Claude 3.7 can show this future today?”. I’d love for lots more people to see the demo, but I really could use an in-kind credit donation to make this viable. If anyone at anthropic is inspired and wants to hook me up: snickell@alumni.stanford.edu. Very happy to rep Claude 3.7 even more than I already do.
I think it’s great advertising for Claude. I believe the reason Claude seems to do SO much better at this task is, one it shows far greater spatial intelligence, and two, I distract they are the only state of the art model intentionally training on SVG.
If you end up taking this further and self hosting a model you might actually achieve a way faster “frame rate” with speculative decoding since I imagine many frames will reuse content from the last. Or maybe a DSL that allows big operations with little text. E.g. if it generates HTML/SVG today then use HAML/Slim/Pug: https://chatgpt.com/share/67e3a633-e834-8003-b301-7776f76e09...
Nobody has really decided on a name.
Also chain of thought is somewhat different from chain of thought reasoning so mb throw in multimodal chain of thought reasoning
You can do that with diffusion, too. Just lock the parameters in ComfyUi.
With current GPU technology, this system would need its own Dyson sphere.
I'm super excited for all the free money and data our new AI written apps will be giving away.
https://chatgpt.com/share/67e32d47-eac0-8011-9118-51b81756ec...
https://mordenstar.com/blog/chatgpt-4o-images
It's definitely impressive though once again fell flat on the ability to render a 9-pointed star.
[1] https://techcrunch.com/wp-content/uploads/2024/03/pasted-ima...
Then I asked for some changes:
> That's almost perfect! Retain this style and the elements, but adjust the text to read:
> [refined text]
> And then below it should add the location and date details:
> [location details]
Then google:
> Gemini 2.5: Our most intelligent AI model
> Introducing Gemini 2.0 | Our most capable AI model yet
I could go on forever. I hope this trend dies and apple starts using something effective so all the other companies can start copying a new lexicon.
Apple is more of a hardware company. Still, Cook does have a few big wins under his belt: M-series ARM chips on Macs, Airpods, Apple watch, Apple pay.
Hotwheels: Fast. Furious. Spectacular.
Which is especially relevant when it's not obvious which product is the latest and best just looking at the names. Lots of tech naming fails this test from Xbox (Series X vs S) to OpenAI model names (4o vs o1-pro).
Here they claim 4o is their most capable image generator which is useful info. Especially when multiple models in their dropdown list will generate images for you.
<Product name>: Our most <superlative> <thing> yet|ever.
No API yet, and given the slowness I imagine it will cost much more than the $0.03+/image of competitors.
Gemini "integrates" Imagen 3 (a diffusion model) only via a tool that Gemini calls internally with the relevant prompt. So it's not a true multimodal integration, as it doesn't benefit from the advanced prompt understanding of the LLM.
Edit: Apparently Gemini also has an experimental native image generation ability.
The results are ground breaking in my opinion. How much longer until an AI can generate 30 successive images together and make an ultra realistic movie?
im not going to get super hyperbolic and histrionic about “entitlement” and stuff like that, but… literally this technology did not exist until like two years ago, and yet i hear this all the time. “oh this codegen is pretty accurate but it’s slow”, “oh this model is faster and cheaper (oh yeah by the way the results are bad, but hey it’s the cheapest so it’s better)”. like, are we collectively forgetting that the whole point of any of this is correctness and accuracy? am i off-base here?
the value to me of a demonstrably wrong chat completion is essentially zero, and the value of a correct one that anticipates things i hadn’t considered myself is nearly infinite. or, at least, worth much, much more than they are charging, and even _could_ reasonably charge. it’s like people collectively grouse about low quality ai-generated junk out of one side of their mouths, and then complain about how expensive the slop is out of the other side.
hand this tech to someone from 2020 and i guarantee you the last thing you’d hear is that it’s too slow. and how could it be? yeah, everyone should find the best deals / price-value frontier tradeoff for their use case, but, like… what? we are all collectively devaluing that which we lament is being devalued by ai by setting such low standards: ourselves. the crazy thing is that the quickly-generated slop is so bad as to be practically useless, and yet it serves as the basis of comparison for… anything at all. it feels like that “web-scale /dev/null” meme all over again, but for all of human cognition.
The animation is a lie. The new 4o with "native" image generating capabilities is a multi-modal model that is connected to a diffusion model. It's not generating images one token at a time, it's calling out to a multi-stage diffusion model that has upscalers.
You can ask 4o about this yourself, it seems to have a strong understanding of how the process works.
This option is not exposed in ChatGPT, it only uses vivid.
Was anyone else surprised how slow the images were to generate in the livestream? This seems notably slower than DALLE.
I ran stable diffusion for a couple of years (maybe?, time really hasn't made sense since 2020) on my Dual 3090 rendering server. I built the server originally for crypto heating my office in my 1820s colonial in upstate NY then when I was planning to go back to college (got accepted into a university in England), I switched it's focus to Blender/UE4 (then 5), then eventually to AI image gen. So I've never minded 20 seconds for an image. If I needed dozens of options to pick the best, I was going to click start and grab a cup of coffee, come back and maybe it was done. Even if it took 2 hours, it is still faster than when I used to have to commission art for a project.
I grew out of Stable Diffusion, though, because the learning curve beyond grabbing a decent checkpoint and clicking start was actually really high (especially compared to LLMs that seamed to "just work"), after going through failed training after failed fine-tuning using tutorials that were a couple days out of date, I eventually said, fuck it, I'm paying for this instead.
All that to say - if you are using GenAI commercially, even if an image or a block of code took 30 minutes, it's still WAY cheaper than a human. That said, eventually a professional will be involved, and all the AI slop you generated will be redone, which will still cost a lot, but you get to skip the back and forth figuring out style/etc.
Currently, my prompts seem to be going to the latter still, based on e.g. my source image being very obviously looped through a verbal image description and back to an image, compared to gemini-2.0-flash-exp-image-generation. A friend with a Plus plan has been getting responses from either.
The long-term plan seems to be to move to 4o completely and move Dall-E to its own tab, though, so maybe that problem will resolve itself before too long.
the native just.. works
I'm not saying that it's not true, it's just "wait and see" before you take their word as gold.
I think MS's claim on their quantum computing breakthrough is the latest form of this.
just tried it, prompt adherence and quality is... exactly what they said, it extremely impressive
I was blown away when they showed this many months ago, and found it strange that more people weren't talking about it.
This is much more precise than the Gemini one that just came out recently.
Some simply dislike everything OpenAI. Just like everything Musk or Trump.
If that's best of 8, I'd love to see the outtakes.
How much longer until an AI that can generate 30 frames with this quality and make a movie?
About 1.5 years ago, I thought AI would eventually allow anyone with an idea to make a Hollywood quality movie. Seems like we're not too far off. Maybe 2-3 more years?
Other image generators I've used lately often produced pretty good images of humans, as well [0]. It was DALLE that consistently generated incredibly awful images. Glad they're finally fixing it. I think what most AI image generators lack the most is good instruction following.
[0] YandexArt for the first prompt from the post: https://imgur.com/a/VvNbL7d The woman looks okay, but the text is garbled, and it didn't fully follow the instruction.
For drawings, NovelAI's models are way beyond the uncanny valley now.
To think that a few years ago we had dreamy pictures with eyes everywhere. And not long ago we were always identifying the AI images by the 6 fingered people.
I wonder how well the physics is modeled internally. E.g. if you prompt it to model some difficult ray tracing scenario (a box with a separating wall and a light in one of the chambers which leaks through to the other chamber etc)?
Or if you have a reflective chrome ball in your scene, how well does it understand that the image reflected must be an exact projection of the visible environment?
EDIT: Ok it works in Sora, and my jaw dropped
For example, I asked it to render a few lines of text on a medieval scroll, and it basically looked like a picture of a gothic font written onto a background image of a scroll
It is incredibly difficult to develop an art style, then get the model to generate a collection of different images in that unique art style. I couldn't work out how to do it.
I also couldn't work out how to illustrate the same characters or objects in different contexts.
AI seems great for one off images you don't care much about, but when you need images to communicate specific things, I think we are still a long way away.
Asking it to draw the Balkans map in Tolkien style, this is actually really impressive, geography is more or less completely correct, borders and country locations are wrong, but it feels like something I could get it to fix.
> I wasn't able to generate the map because the request didn't follow content policy guidelines. Let me know if you'd like me to adjust the request or suggest an alternative way to achieve a similar result.
Are you in the US?
...why are we living in such a retarded sci-fi age
Generate a photo of a lake taken by a mobile phone camera. No hands or phones in the photo, just the lake.
The hand holding a phone is always there :D
The general idea of indistinguishable real/fake images; yeah
You don't even need deepfakes. https://www.newsweek.com/doug-mastriano-pennsylvania-senator...
The disaster scenario is already here.
Theme: Educational Scientific Visualization – Ultra Realistic Cutaways Color: Naturalistic palettes that reflect real-world materials (e.g., rocky grays, soil browns, fiery reds, translucent biological tones) with high contrast between layers for clarity Camera: High-resolution macro and sectional views using a tilt-shift camera for extreme detail; fixed side angles or dynamic isometric perspective to maximize spatial understanding Film Stock: Hyper-realistic digital rendering with photogrammetry textures and 8K fidelity, simulating studio-grade scientific documentation Lighting: Studio-quality three-point lighting with soft shadows and controlled specular highlights to reveal texture and depth without visual noise Vibe: Immersive and precise, evoking awe and fascination with the inner workings of complex systems; blends realism with didactic clarity Content Transformation: The input is transformed into a hyper-detailed, realistically textured cutaway model of a physical or biological structure—faithful to material properties and scale—enhanced for educational use with visual emphasis on internal mechanics, fluid systems, and spatial orientation
Examples: 1. A photorealistic geological cutaway of Earth showing crust, tectonic plates, mantle convection currents, and the liquid iron core with temperature gradients and seismic wave paths. 2. An ultra-detailed anatomical cross-section of the human torso revealing realistic organs, vasculature, muscular layers, and tissue textures in lifelike coloration. 3. A high-resolution cutaway of a jet engine mid-operation, displaying fuel flow, turbine rotation, air compression zones, and combustion chamber intricacies. 4. A hyper-realistic underground slice of a city showing subway lines, sewage systems, electrical conduits, geological strata, and building foundations. 5. A realistic cutaway of a honeybee hive with detailed comb structures, developing larvae, worker bee behavior zones, and active pollen storage processes.
One area where it does not work well at all is modifying photographs of people's faces.* Completely fumbles if you take a selfie and ask it to modify your shirt, for example.
* = unless the people are in the training set
Sounds like it may be a safety thing that's still getting figured out
Might take a day or two before it's available in general.
It seems like an odd way to name/announce it, there's nothing obvious to distinguish it from what was already there (i.e. 4o making images) so I have no idea if there is a UI change to look for, or just keep trying stuff until it seems better?
Truly infuriating, especially when it's something like this that makes it tough to tell if the feature is even enabled.
The glaring issue for the older image generators is how it would proudly proclaim to have presented an image with a description that has almost no relation to the image it actually provided.
I'm not sure if this update improves on this aspect. It may create the illusion of awareness of the picture by having better prompt adherence.
https://news.ycombinator.com/item?id=42628742
The new one can.
https://chatgpt.com/share/67e36dee-6694-8010-b337-04f37eeb5c...
It's much better than prior models, but still generates hands with too many fingers, bodies with too many arms, etc.
I see errors like this in the console:
ewwsdwx05evtcc3e.js:96 Error: Could not fetch file with ID file_0000000028185230aa1870740fa3887b?shared_conversation_id=67e30f62-12f0-800f-b1d7-b3a9c61e99d6 from file service at iehdyv0kxtwne4ww.js:1:671 at async w (iehdyv0kxtwne4ww.js:1:600) at async queryFn (iehdyv0kxtwne4ww.js:1:458)Caused by: ClientRequestMismatchedAuthError: No access token when trying to use AuthHeader
I'm excited to see what a Flux 2 can do if it can actually use a modern text encoder.
Sora is one of the worst video generators. The Chinese have really taken the lead in video with Kling, Hailuo, and the open source Wan and Hunyuan.
Wan with LoRAs will enable real creative work. Motion control, character consistency. There's no place for an OpenAI Sora type product other than as a cheap LLM add-in.
> Developers will soon be able to generate images with GPT‑4o via the API, with access rolling out in the next few weeks.
That's it folks. Tens of thousands of so-called "AI" image generator startups have been obliterated and taking digital artists with them all reduced to near zero.
Now you have a widely accessible meme generator with the name "ChatGPT".
The last task is for an open weight model that competes against this and is faster and all for free.
ChatGPT has already had a that via Dall-E. If it didn't kill those startups when that happened this doesn't fundamentally change anything. Now its got a new image gen model, which — like Dall-E 3 when it came out — is competitive or ahead of other SotA base models using just text prompts, the simplest generation workflow, but both more expensive and less adaptable to more involved workflows than the tools anyone more than a casual user (whether using local tools or hosted services) is using. This is station-keeping for OpenAI, not a meaningful change in the landscape.
Trying out 4o image generation... It doesn't seem to support this use-case at all? I gave it an image of myself and asked to turn me into a wizard, and it generate something that doesn't look like me in the slightest. A second attempt, I asked to add a wizard hat and it just used python to add a triangle in the middle of my image. I looked at the examples and saw they had a direct image modification where they say "Give this cat a detective hat and a monocle", so I tried that with my own image "Give this human a detective hat and a monocle" and it just gave me this error:
> I wasn't able to generate the modified image because the request didn't follow our content policy. However, I can try another approach—either by applying a filter to stylize the image or guiding you on how to edit it using software like Photoshop or GIMP. Let me know what you'd like to do!
Overall, a very disappointing experience. As another point of comparison, Grok also added image generation capabilities and while the ability to edit existing images is a bit limited and janky, it still manages to overlay the requested transformation on top of the existing image.
Even when I told it to transform it into a text description, then draw that text description, my earlier attempt at a cat picture meant that the description was too close to a banned image...
I can't help but feel like openAI and grok are on unhelpful polar opposites when it comes to moderation.
In the coming days, people will Anime all sorts of images, for example historical images: https://x.com/keysmashbandit/status/1904764224636592188
Iterations are the missing link. With ChatGPT, you can iteratively improve text (e.g., "make it shorter," "mention xyz"). However, for pictures (and video), this functionality is not yet available. If you could prompt iteratively (e.g., "generate a red car in the sunset," "make it a muscle car," "place it on a hill," "show it from the side so the sun shines through the windshield"), the tools would become exponentially more useful.
I‘m looking forward to try this out and see if I was right. Unfortunately it’s not yet available for me.
Ditto Instruct Pix2Pix https://www.timothybrooks.com/instruct-pix2pix
For example, https://news.ycombinator.com/item?id=43388114
Otherwise impressive.
Am I the only one immediately looking past the amazing text generation, the excellent direction following, the wonderful reflection, and screaming inside my head, "That's not how reflection works!"
I know it's super nitpicky when it's so obviously a leap forward on multiple other metrics, but still, that reflection just ain't right.
Edit: are we talking about the first or second image? I meant to say the image with only the woman seems normal. Image with the two people does seem a bit odd.
In games they did it by creating a duplicate then reversing it, I wonder if this is the same idea.
I think it is too biased to use heuristics discovered in the first response to apply the same level of compute to subsequent requests.
It makes me kind of want to rewrite an interface that builds appropriate context and starts new chats for every request issued..
As to why they don't automatically detect when reasoning could be appropriate and then switch to o3, I don't know, but I'd assume it's about cost (and for most users the output quality is negligible). 4o can do everything, it's just not great at "logic".
--
Comparison with Leonardo.Ai.
ChatGPT: https://chatgpt.com/share/67e2fb21-a06c-8008-b297-07681dddee...
ChatGPT again (direct one shot): https://chatgpt.com/share/67e2fc44-ecc8-8008-a40f-e1368d306e...
ChatGPT again (using word "photorealistic instead of "photo"): https://chatgpt.com/share/67e2fce4-369c-8008-b69e-c2cbe0dd61...
Leonardo.Ai Phoenix 1.0 model: https://cdn.leonardo.ai/users/1f263899-3b36-4336-b2a5-d8bc25...
I'm curious if you said 2d animation style for both or just for chatgpt.
Edit: Your second version of chatgpt doesn't say photorealistic. Can you share the Leonard.ai prompt?
It also misses the arrow between "[diffusion]" and "pixels" in the first image.
nah. i pass and stick with midjourney.
EDIT: Seems not, "The smallest image size I can generate is 1024x1024. Would you like me to proceed with that, or would you like a different approach?"
How easy is this to remove? Is it just like exif data that can be easily stripped out, or is it baked in more permanently somehow
I couldn't find anything on the pricing page.
This dynamic happens on Twitter every day. Tomorrow it'll be a different craze.
If the subject matter is paywalled, I feel that the post should include some explanation of what is newsworthy behind the link.
It's more pragmatic to pipeline the results to a background removal model.
EDIT: It appears GPT-4o is different as there is a video demo dedicated to transparancy.
Sorry, but how are these useful? None of the examples demonstrate any use beyond being cool to look at.
The article vaguely mentions 'providing inspiration' as possible definition of 'useful'. I suppose.
And I hope that people who worked on this know this. They are pure evil.
May 7, 2024 - The “Let Loose” event, focusing on new iPads, including the iPad Pro with the M4 chip and the iPad Air with the M2 chip, along with the Apple Pencil Pro.
June 10, 2024 - The Worldwide Developers Conference (WWDC) keynote, where Apple introduced iOS 18, macOS Sequoia, and other software updates, including Apple Intelligence.
September 9, 2024 - The “It’s Glowtime” event, where Apple unveiled the iPhone 16 series, Apple Watch Series 10, and AirPods 4.
Via Press releases: MacBook Air with M3 on March 4, the iPad mini on October 15, and various M4-series Macs (MacBook Pro, iMac, and Mac mini) in late October.
so much fun.
...Once the wait time is up, I can generate the corrected version with exactly eight characters: five mice, one elephant, one polar bear, and one giraffe in a green turtleneck. Let me know if you'd like me to try again later!
ofc 4.5 is best, but its slow and I am afraid I'm going to hit limits.
Was it public information when Google was going to launch their new models? Interesting timing.