First time native image gen was introduced in Gemini 1.5 Flash if I'm not wrong, and then OpenAI was released for 4o which took over the internet by Ghibli Art.
We have been getting good quality images from almost all image generators like Midjourney, OpenAI and other providers, but the thing that made it special was true "multimodal" nature of it. Here's what I mean
When you used to ask chatgpt to create an image, it will rephrase that prompt and internally send that prompt to Dalle, similarly gemini would send it to Imagen which were diffusion models and they had little to know context in your next response about what's there in the previous image
In native image generation, it understands Audio, Text and even Image tokens in the same model and need not to rely on diffusion models internally, I don't think both Openai and google has released how they've trained it but my guess is that it's partially auto-regressive and diffusion but not sure about it