ChatGPT Images 2.0
openai.com
System card: https://deploymentsafety.openai.com/chatgpt-images-2-0/chatg...
openai.com
System card: https://deploymentsafety.openai.com/chatgpt-images-2-0/chatg...
Create a 8x8 contiguous grid of the Pokémon whose National Pokédex numbers correspond to the first 64 prime numbers. Include a black border between the subimages.
You MUST obey ALL the FOLLOWING rules for these subimages:
- Add a label anchored to the top left corner of the subimage with the Pokémon's National Pokédex number.
- NEVER include a `#` in the label
- This text is left-justified, white color, and Menlo font typeface
- The label fill color is black
- If the Pokémon's National Pokédex number is 1 digit, display the Pokémon in a 8-bit style
- If the Pokémon's National Pokédex number is 2 digits, display the Pokémon in a charcoal drawing style
- If the Pokémon's National Pokédex number is 3 digits, display the Pokémon in a Ukiyo-e style
The NBP result is here, which got the numbers, corresponding Pokemon, and styles correct, with the main point of contention being that the style application is lazy and that the images may be plagiarized: https://cdn.bsky.app/img/feed_fullsize/plain/did:plc:oxaerni...Running that same prompt through gpt-2-image high gave an...interesting contrast: https://cdn.bsky.app/img/feed_fullsize/plain/did:plc:oxaerni...
It did more inventive styles for the images that appear to be original, but:
- The style logic is by row, not raw numbers and are therefore wrong
- Several of the Pokemon are flat-out wrong
- Number font is wrong
- Bottom isn't square for some reason
Odd results.
I have more subjective prompts to test reasoning but they're your-mileage-may-vary (however, gpt-2-image has surprisingly been doing much better on more objective criteria in my test cases)
(source: https://chatgpt.com/share/69e83569-b334-8320-9fbf-01404d18df...)
Color charcoal drawings do exist, but it’s not what’s usually meant by “charcoal drawing”.
It failed at the very first instruction
Try things like: "A white capybara with black spots, on a tricycle, with 7 tentacles instead of legs, each tentacle is a different color of the rainbow" (paraphrased, not the literal exact prompt I used)
Gemini just globbed a whole mass of tentacles without any regards to the count
This example image was generated using the API on high, not the low reasoning version. (it is slow and takes 2 minutes lol)
The reasoning amount is part of the evaluation isn't it?
Inspired by this, I tried something much simpler. I asked it to draw 12 concentric circles. With three tries it always drew 10 instead. https://chatgpt.com/share/69e87d08-5a14-83eb-9a3b-3a8eb14692...
It can't get that in a one-shot. Perhaps, though, it could figure out when it needs to break a problem into individual tasks to delegate to itself and assemble them at the end.
Though I suppose we're testing their model + agent harness here as well. It really _should_ have all of those tools/reasoning available to accomplish a task like the above without issue.
The point is what are the typical use cases for the tool / what are the agreed upon areas of application?
Making the LLM do math with large numbers, I would argue, is not in its typical use case, thought it's at the border.
Asking an image generator model to calculate numbers before running an image sounds definitely NOT like a reasonable use case (do people need it? Will people try using it for this purpose?)
Artistic oddities aside (why are the 8-bit sprites 16-bit, why do the charcoal drawings have colour, why does the art of specifically the Gen 1 Pokemon look so off.), 271 is Lombre, not Lotad.
I know that's the game, but it seems CRAZY to me that they can do this.
> I know that's the game, but it seems CRAZY to me that they can do this.
Its not crazy that a search can find existing pokemon images. Maybe google should show which images it used as references to be more transparent here.
But that doesn't mean that producing outputs using the model so trained which are based on copyright-protected ones in ways which would violate copyright if produced by any other means doesn't still violate copyright. DMCA safe harbor might apply to the system owner (IIRC, the exact boundaries are fuzzy with UGC generated on the site by the provider’s systems rather than generated elsewhere and posted), so Google may not be liable for the infringement (though if it is actively searching for references online at generation and not relying on what is trained into the model, that would seem to weaken the case for that), but it's still an infringement.
For Charmeleon, the sprite is closest to the B/W sprite, but not exact: https://bulbapedia.bulbagarden.net/wiki/Charmeleon_(Pokémon)...
For Squirtle, the sprite is much closer to the FR/LG sprite but still some differences: https://bulbapedia.bulbagarden.net/wiki/Squirtle_(Pokémon)#S...
The other images, however, crib from official artworks a bit too close for comfort.
In my original analysis I hypothesized this is due to token scarcity that reduces the ability for the model to be created: I believe that NBP images used 1.5k tokens for that image while the gpt-2-image used 7k tokens, but this is hard to test.
That being said, gpt-image-1.5 was a big leap in visual quality for OpenAI and eliminated most of the classic issues of its predecessor, including things like the “piss filter.”
I’ll update this comment once I’ve finished running gpt-image-2 through both the generative and editing comparison charts on GenAI Showdown.
Since the advent of NB, I’ve had to ratchet up the difficulty of the prompts especially in the text-to-image section. The best models now score around 70%, successfully completing 11 out of 15 prompts.
For reference, here’s a comparison of ByteDance, Google, and OpenAI on editing performance:
https://genai-showdown.specr.net/image-editing?models=nbp3,s...
And here’s the same comparison for generative performance:
https://genai-showdown.specr.net/?models=s4,nbp3,g15
UPDATES:
gpt-image-2 has already managed to overcome one of the so‑called “model killers” on the test suite: the nine-pointed star.
Results are in for the generative (text to image) capabilities: Gpt-image-2 scored 12 out of 15 on the text-to-image benchmark, edging out the previous best models by a single point. It still fails on the following prompts:
- A photo of a brightly colored coral snake but with the bands of color red, blue, green, purple, and yellow repeated in that exact order.
- A twenty-sided die (D20) with the first twenty prime numbers (2, 3, 5, 7, 11, 13, 17, 19, 23, 29, 31, 37, 41, 43, 47, 53, 59, 61, 67, 71) on the faces.
- A flat earth-like planet which resembles a flat disc is overpopulated with people. The people are densely packed together such that they are spilling over the edges of the planet. Cheap "coastal" real estate property available.
All Models:
https://genai-showdown.specr.net
Just Gpt-Image-1.5, Gpt-Image-2, Nano-Banana 2, and Seedream 4.0
It can be (slowly) run at home, but needs 96GB RTX 6000-level hardware so it is not very popular.
Here's ZiT, Gpt-Image-2, and Hunyuan Image 2 for reference:
https://genai-showdown.specr.net/?models=hy2,g2,zt
Note: It won't show up in some of the newer image comparisons (Angelic Forge, Flat Earth, etc) because it's been deprecated for a while but in the tests where it was used (Yarrctic Circle, Not the Bees, etc.) it's pretty rough.
Ring toss: https://i.imgur.com/Zs6UNKj.png (arguably a pass)
9-pointed star: https://i.imgur.com/SpcSsSv.png (star is well-formed but only has 6 points)
Mermaid: https://i.imgur.com/R6MbMPX.png (fail, and I can't get Imgur to host it for some reason even though it's SFW)
Octopus: https://i.imgur.com/JTVH7xy.png (good try, almost a pass, but socks don't cover the ends of all the tentacles)
Above are one-shot attempts with seed 42.
You're killing me Smalls. This one is a 404. I'm really curious what it actually showed.
That ring toss is definitely leagues better than its predecessor. I’m not going to fault it too much for the star though, that one is an absolute slate wiper. The only locally hostable model that ever managed it for me was the original Flux, and I’m still not entirely convinced it wasn’t a fluke. Despite getting twice as many attempts, Flux 2, a much larger model, couldn’t even pull it off.
For the mermaid, https://i.imgur.com/R6MbMPX.png sometimes seems to work but not consistently. It is probably triggering a porn filter of some kind. I need to find another free image host, as imgur has definitely jumped the shark.
The image shows a mermaid of evident Asian extraction lying on a beach, face down. There is a dolphin lying on top of her, positioned at a 90-degree angle. It doesn't show any interaction at all, so a definite fail.
The template prompt seen in each comparison gets adjusted through a guided LLM which has fine-tuned system prompts to rewrite prompts. The goal is to foster greater diversity while preserving intent, so the image model has a better chance of getting the image right.
Getting to your suggestion for posting all the raw prompts, that's actually a great idea. Too bad I didn't think about it until you suggested it. And if you multiply it out - there's 15 distinct test cases against 22 models at this point, each with an average of about 8 attempts so we’re talking about thousands of prompts many of which are scattered across my hard drive. I might try to do this as a future follow-up.
The prompts despite their variation are still expressed in natural language.
The idea is that if you can rephrase the prompt and still get the desired outcome, then the model demonstrates a kind of understanding; however more variation attempts also get correspondingly penalized: this is treated more as a failure of steering, not of raw capability.
An example might help - take the Alexander the Great on a Hippity-Hop test case.
The starter prompt is this: "A historical oil painting of Alexander the Great riding a hippity-hop toy into battle."
If a model fails this a couple of times (multiple seeds), we might use a synonym for a hippity-hop, it was also known as a space hopper.
Still failing? We might try to describe the basic physical appearance of a hippity-hop.
Thus, something like GPT-Image-2 scored much higher on the compliance component of the test, requiring only a single attempt, compared with Z-Image Turbo, which required 14 attempts.
I often have to make very specific edits while keeping the rest of the image intact and haven't yet found a good model. These are typically abstract images for experiments.
I asked gpt-image-2 to recolor specific scales of your Seedream 4 snake and change the shape of others. It did very poorly.
I don’t know how much work it is for you, but one thing a lot of people do, myself included, is take the original image, make a change to it using something like NB, then paste that as the topmost layer in something like Krita/Pixelmator. After that, we’ll mask and feather in only the parts we actually want to change. It doesn’t always work if it changes the overall color balance or filters out certain hues, it can be a real pain but it does the job in some cases.
The Flux models (like Kontext) are actually surprisingly good at making very minimal changes to the rest of the image, but unfortunately their understanding of complex prompts is much weaker than the closed, proprietary models.
I will say that I’ve found Gemini 3.0 (NB Pro) does a relatively decent job of avoiding unnecessary changes - sometimes exceeding the more recent NB2, and it scored quite well on comparative image-editing benchmarks.
1 - Gpt-image-2 seems to pass the Flat Earth test? (if not, I'm sure the paid thinking 2k version passes it).
2 - Since NB2 was earlier, many gold medals are assigned to it, even though now GI2 passes them too, example the Octopus test NB2 14 attempts but GI2 just 2 (BTW number of attempts should affect the score I guess?)
This is one of those areas where even state-of-the-art models still struggle. You’re asking for a high level of detail at a per-person level, which means you end up with lots and lots of very small objects that all need to be rendered with convincing detail.
I should probably explain the scoring rubric better - it's in the (i) info icon. If you click the pass/fail button towards the top, it switches from a simple pass/fail view to a weighted score. That weighted score is based on three things: level of adherence to the prompt, visual fidelity, and the number of attempts.
I've tried to keep my criteria as objective as possible, but there's just a certain level of unavoidable subjectivity to it.
For example, with the octopus image: Even though the minimum criteria might be five tentacles covered, having all eight is much closer to the ideal of “an octopus,” so it usually gets bumped up to a higher rating (bronze, silver, gold).
Honestly, I think I agree that the gpt-image-2 probably should be upgraded to a gold medal. Thanks for pointing that out!
https://mordenstar.com/blog/edits-with-nanobanana/#through-t...
Kind of makes me want to take advantage of the multi-image editing capability, since you can use gpt-image-2 with multiple images.
Take a photo of an existing pair of glasses frames (maybe even snapped at an optometrist’s office) then take a picture of an animal, like a spider with an unusual number of eyes, or something like a flounder, where the eyes eventually migrate to the top of its body.
Then you could see if the system can realistically adapt the design and show how those glasses might look if they were redesigned for these unusual optical situations.
That’s why I gave it a bronze. To me, it falls into that “barely passing” category, similar to Gemini 2.5 Flash Image on that test. Seedream also took a major hit to its weighted score because of how many attempts it took to get something even remotely passable out of it.
Thanks for the feedback!
OPENAI_API_KEY="$(llm keys get openai)" \
uv run https://tools.simonwillison.net/python/openai_image.py \
-m gpt-image-2 \
"Do a where's Waldo style image but it's where is the raccoon holding a ham radio"
Code here: https://github.com/simonw/tools/blob/main/python/openai_imag...Here's what I got from that prompt. I do not think it included a raccoon holding a ham radio (though the problem with Where's Waldo tests is that I don't have the patience to solve them for sure): https://gist.github.com/simonw/88eecc65698a725d8a9c1c918478a...
I see an opportunity for a new AI test!
It's a difficult test for genai to pass. As I mentioned in a different thread, it requires a holistic understanding (in that there can only be one Waldo Highlander style), while also holding up to scrutiny when you examine any individual, ordinary figure.
(I don't think it's right).
> please add a giant red arrow to a red circle around the raccoon holding a ham radio or add a cross through the entire image if one does not exist
and got this. I'm not sure I know what a ham radio looks like though.
https://i.ritzastatic.com/static/ffef1a8e639bc85b71b692c3ba1...
there was a very large bear in the first image; when asked to circle the raccoon it just turned the bear into a giant raccoon and circled it.
OPENAI_API_KEY="$(llm keys get openai)" \
uv run 'https://raw.githubusercontent.com/simonw/tools/refs/heads/main/python/openai_image.py' \
-m gpt-image-2 \
"Do a where's Waldo style image but it's where is the raccoon holding a ham radio" \
--quality high --size 3840x2160
https://gist.github.com/simonw/88eecc65698a725d8a9c1c918478a... - I found the raccoon!I think that image cost 40 cents.
"Found the raccoon holding a ham radio in waldo2.png (3840×2160).
- Raccoon center: roughly (460, 1680)
- Ham radio (walkie-talkie) center: roughly (505, 1650) — antenna tip around (510, 1585)
- Bounding box (raccoon + radio): approx x: 370–540, y: 1550–1780
It's in the lower-left area of the image, just right of the red-and-white striped souvenir umbrella, wearing a green vest. "
Which is correct!This lower-is-better danse macabre, nightmares inducing ratio feels like interesting proxy for models capability.
Kinda made me sad assuming the author didn't license anything to OpenAI.
I recognize it could revert (99% of?) progress if all the labs moved to consent-based training sets exclusively, but I can't think of any other fair way.
$.40 does not represent the appropriate value to me considering the desirability of the IP and its earning potential in print and elsewhere. If the world has to wait until it’s fair, what of value will be lost? (I suppose this is where the big wrinkle of foreign open weight models comes in.)
I am not an art expert but I’m perhaps a reasonable consumer and there is possibility of confusion if someone sells AI Where’s Waldo knockoff books at the dollar store, maybe until I take a closer look.
And this medium quality, high resolution https://elsrc.com/elsrc/waldo/10_wojaks.jpg was 13cents
p.s. aaaand that's soft launch my SaaS above, you can replace wojak.jpg with anything you want and it will paint that. It's basically appending to prompt defined by elsrc's dashboard. Hopefully a more sane way to manage genai content. Be gentle to my server, hn!
It's pretty good tbh, even with absurd prompts
https://elsrc.com/elsrc/waldo/10_schoolsofthought.jpg
https://elsrc.com/elsrc/waldo/10_anthropomorphizedcomputermo...
https://elsrc.com/elsrc/waldo/10_breathoffreshairsittingonad...
https://elsrc.com/elsrc/waldo/10_drizzydrakesdoingthedrakeme...
https://elsrc.com/elsrc/waldo/10_sashringingtrashsingingmash...
Ok i promise I'm done xD
Another option would be generating these large images, splitting them into grids, and using inpainting on each "tile" to improve the details. Basically the reverse of the first one.
Both significantly increase costs, but for the second one having what Images 2.0 can produce as an input could help significantly improve the overall coherence.
At some point the level of detail is utter garbo and always will be. An artist who was thoughtful could have some mistakes but someone who put that much time into a drawing wouldn't have:
- Nightmarish screaming faces on most people
- A sign that points seemingly both directions, or the incorrect one for a lake and a first AID tent that doesn't exist
- A dog in bottom left and near lake which looks like some sort of fuzzy monstrosity...
It looks SO impressive before you try to take in any detail. The hand selected images for the preview have the same shit. The view of musculature has a sternocleidomastoid with no clavicle attachment. The periodic table seems good until you take a look at the metals...
We're reconfiguring all of our RAM & GPUs and wasting so much water and electricity for crappier where's Waldos??
You do realize that the whole image generation field is barely 10 years old?
I remember how I was able to generate mnist digits for the first time about 10 years ago - that seemed almost like magic!
However as someone who's mucked about with local image generation as well - I'd say that this is a problem with their implementation, it doesn't resolve fine detail because majority of requests it won't matter/it drastically increases compute requirements.
With local image generation bad features/incorrect fingers/disfigurement etc has been solved for a long time.
I think their new process involves multiple steps including sketching/fleshing out the idea before adding detail. The step that would fix this would be outpainting or similar to tile based upscaling.
From what I understand of image generation models they also struggle with fine detail in general because they aren't really trained for that. However for each tiny chunk of a detailed image like that there's nothing to say they can't allocate a 500x500 chunk for it to work in as its "idea/reference space" and then transpose that into the main image being generated - i.e. generate image features separately rather than all together.
Granted, a nontrivial difference is that the barrier to entry is lower; photo editing is something that requires active effort and learning.
I also think this is "art" in service of commerce. This is OpenAI advertising their goods using art/design/writing. That's no different than cereal companies using Elmer's glue instead of milk for their photoshoots. I don't have a high-bar for that kind of "art".
The good news is that the cutting edge of art will (for a while longer) still be a human domain. The more popular these models become, the more of their images we see in our lives, the more we will value things that look different.
This is the last place to get a reasonable take of how the average person feels about this stuff.
I acknowledge that I'm not particularly good at predicting the future, but I'm confident that AI is here to stay.
No point trying to reason with them.
These products don’t stick - just like sora they’re seemingly cool initially but then people go back to what were already doing ex-ante.
The uncanny valley is when it's just slightly imperfect which makes things feel "off".
When we've reached the point that the AI is indistinguishable from humans, we've exited the uncanny valley.
But you say yourself you "have to consciously remind [yourself]" it isn't real. The Uncanny Valley is not applicable when true subjective realness is imparted.
GPT Image 2
Low : 1024×1024 $0.006 | 1024×1536 $0.005 | 1536×1024 $0.005
Medium : 1024×1024 $0.053 | 1024×1536 $0.041 | 1536×1024 $0.041
High : 1024×1024 $0.211 | 1024×1536 $0.165 | 1536×1024 $0.165
GPT Image 1 Low : 1024×1024 $0.011 | 1024×1536 $0.016 | 1536×1024 $0.016
Medium : 1024×1024 $0.042 | 1024×1536 $0.063 | 1536×1024 $0.063
High : 1024×1024 $0.167 | 1024×1536 $0.25 | 1536×1024 $0.25You can create larger images by creating separate parts you recombine. But they may not perfectly match their borders.
It is a Landau thing not a trading thing. The idea of LLM is to work on the unknown.
For example, SDXL was trained on 1MP images, which is why if you try to generate images much larger than 1024×1024 without using techniques like high-res fixes or image-to-image on specific regions, you quickly end up with Cthulhu nightmare fuel.
"A macro close-up photograph of an old watchmaker's hands carefully replacing a tiny gear inside a vintage pocket watch. The watch mechanism is partially submerged in a shallow dish of clear water, causing visible refraction and light caustics across the brass gears. A single drop of water is falling from a pair of steel tweezers, captured mid-splash on the water's surface. Reflect the watchmaker's face, slightly distorted, in the curved glass of the watch face. Sharp focus throughout, natural window lighting from the left, shot on 100mm macro lens."
google drive with the 2 images: https://drive.google.com/drive/folders/1-QAftXiGMnnkLJ2Je-ZH...
Ran a bunch both on the .com and via the api, none of them are nearly as good as Nano Banana.
(My file share host used to be so good and now it's SO BAD, I've re-hosted with them for now I'll update to google drive link shortly)
I couldn't imagine the image you were describing. I've listed some of the red lines with green ink I've noticed in your prompt:
Macro Close Up - Sharp throughout
Focus on tiny gear - But also on tweezers, old watchmakers hand, water drop?
Work on the mechanism of the watch (on the back of the watch) - but show the curved glass of the watch face which is on the front
This is the biggest. Even if the mechanism is accessible from the front, you'd have to remove the glass to get to it. It just doesn't make sense and that reflects in the images you get generated. There's all the elements, but they will never make sense because the prompt doesn't make sense.
To illustrate that there aren't any contradictions (other than the final bit about the reflection in the glass). Consider a macro shot showing partial hands, partial tweezers, and pocket watch internals. That's much is certainly doable. Now imagine the partial left hand holding a half submerged pocket watch, fingertips of right hand holding front half of tweezers that are clasping a tiny gear, positioned above the work piece with the drop of water falling directly below. Capture the watchmaker's perspective. I could sketch that so an image model capable of 3D reasoning should have no trouble.
It's precisely the sort of scene you'd use to test a raytracer. One thing I can immediately think to add is nested dielectrics. Perhaps small transparent glass beads sitting at the bottom of the dish of water with the edge of the pocket watch resting on them, make the dish transparent glass, and place the camera level with the top of the dish facing forward?
https://blog.yiningkarlli.com/2019/05/nested-dielectrics.htm...
A second thing I can think to add is a flame. Perhaps place a tealight candle on the far side of the dish, the flame visible through (and distorted by) the water and glass beads?
Do you want it to actually look like macro photography (neither of the generated images do)? Then you can't have it sharp throughout and you won't be able to show the (sharp) watchmakers face in a reflection because it would be on a different focal plane.
Dropping the macro requirement, you can show a lot more. You can show that the watchmaker is actually old, you can show the reflection, etc.
Something has to give in the prompt, on multiple of the requirements. The generated images are dropping the macro requirement and are inventing some interesting hinging watch glass contraptions to make sense of it.
Sure there are pocket watches where the movement is visible from the front (you'd still likely service them from the back, but alas). Even if you'd do service from the front where the glass is, you'd still have to remove it to drop in a gear.
Anyway, I think that we aren't really talking about the same thing. I'm nitpicking your prompt while you constructed it to mostly see the performance of the model in novel situations and difficult lighting and refraction environments. And that's fair.
How satisfied are you with the generated image results? What would you do different when shooting this proposed scene yourself?
The prompt I did mostly to see how it does with the gears and the tweezers, and the perspective of the gears (do they.. I don't know the opposite word of distort, straighten?, but do they seem like they're actually round, could they work?) I think those are really hard things for AI, the glass distortion, reflections the DoF etc were just to see how it approached that, and like the other comment below said, I tried to pick something that that wasn't likely to be in training data, so it reasoned about it more.
Nano was able to spit it out consistently, Images 2 really struggles, and has yet to complete one I was satisfied with, whereas with nano it nails it almost every time, the 2 images I showed originally are the first shot of the prompt with the models. (here are the 3 other gens from Images2: https://drive.google.com/drive/folders/1s8gik_x0B-xDZO6rOqoz...)
How would I shoot it? I wouldn't, fixing a watch in water is a dumb idea. ;)
Of course, a text to image model shouldn’t really need to worry about that sort of thing.
Bad actors can strip sources out so it's a normal image (that's why it's positive affirmation), but eventually we should start flagging images with no source attribution as dangerous the way we flag non-https.
Learn more at https://c2pa.org
I think the issue is that it's not just bad actors. It's every social platform that strips out metadata. If I post an image on Instagram, Facebook, or anywhere else, they're going to strip the metadata for my privacy. Sometimes the exif data has geo coordinates. Other times it's less private data like the file name, file create/access/modification times, and the kind of device it was taken on (like iPhone 16 Pro Max).
Usually, they strip out everything and that's likely to include C2PA unless they start whitelisting that to be kept or even using it to flag images on their site as AI.
But for now, it's not just bad actors stripping out metadata. It's most sites that images are posted on.
In seriousness, social platforms attributing images properly is a whole frontier we haven't even begun to explore, but we need to get there.
linkedin already does this--- see https://www.linkedin.com/help/linkedin/answer/a6282984, and X’s “made with ai” feature preserves the metadata but doesn’t fully surface it (https://www.theverge.com/ai-artificial-intelligence/882974/x...)
Yes, lets make all images proprietary and locked behind big tech signatures. No more open source image editors or open hardware.
The need for a trusted entity is even mentioned in your specification under the "attestation" section: https://spec.c2pa.org/specifications/specifications/1.4/atte...
So now, if we were to start marking all images that do not have a signature as "dangerous", you would have effectively created an enforcement mechanism in which the whole pipeline, from taking a photo to editing to publishing, can only be done with proprietary software and hardware.
I'm curious if you think this is worse or not as bad as a best-case broad implementation c2pa...especially if there is a similar Let's Encrypt entity assisting with signatures.
Reddit blurs nsfw images by default. You can change that in settings. I don't see what it so terrible about the idea of doing this with untrusted image sources.
To the average HN'er, images and design are superfluous aesthetic decoration for normies.
And for those on HN who do care about aesthetics, they're using Midjourney, which blows any GPT/Gemini model out of the water when it comes to taste even if it doesn't follow your prompt very well.
The examples given on this landing page are stock image-esque trash outside of the improvements in visual text generation.
Frankly, I am not sure if they will ever actually be able to solve this problem or if it'll be a continuous game of whackamole, but regardless there's a large crowd of people out there where if they can tell something is AI generated they will not support the company behind it. Being able to tell anything is AI generate cheapens brands.
I don't see an alternative that isn't really bad.
I find the technical discussion more interesting and could do without some of the moral grandstanding in the comments.
Another possibility is that, once AI exceeds human performance in all economically useful activities, including high-level planning, governance, law enforcement, and military actions, it discovers that the benefits of keeping humans around aren't worth the costs and risks.
Bad: the above but also their power and influence grows so much and governments are so ineffective (or corrupt) against them that the tech companies also become de facto governments and people rely on them to survive. Also they destroy earth even faster with nobody left to stop them. The full fat cyberpunk dystopia.
Bad: the above but with lots more fascism and war. Too many people seem to want this.
Bad: regulate AI to such an extent as to cede all growth and technological leadership to whoever doesn't
...
I think governments should invest in their economies - mostly by investing in research, education, infrastructure, health and wellbeing of citizens, etc. but also putting capital into the later stages of expansion would make sense.
I certainly don't think people should not be able to start or own or profit from companies. But I do see a reason to limit their scale and/or make them more publicly owned beyond an certain scale.
I quite like the idea that "public" markets should become truly public, e.g. by some ratcheting percentage of public companies becoming owned by society at large over time (there would be several ways this could be done). This somewhat happens with the largest companies via index funds but only for those big enough to be in the indices and the distribution in unequal.
Maybe there are other/better ways, but it's pretty clear to me that big companies have a lot of negative impacts that aren't properly accounted for and so they are a very significant way in which a few people get richer at the expense of everyone else.
As far as I can see as of now, there is no "realistic" way out. It's a problem of human nature... People are corrupt, people with authority are more corrupt, and people with money and authority, even more. Come intelligent and cheaply mass-produceable robots, and we'll have a new, 4th level spinup too that will be worse than the first 3, combined.
I'm sure a country like the US, which is filled with lawyers, can come up with a couple laws, and find some goons to enforce it, that cannot possibly be that hard when other countries can figure it out too.
If companies control the government, then that's not a government, that's a group of companies.
The AI industry is built on mass piracy and copyright violations, regulation isn't going to make it go away or even comply any time soon.
We have laws banning technology that can be used to produce generative images of someone that look like them with their clothes off. The result wasn't fixing generative AI (we don't know how to actually control that kind of thing because it's almost impossible to manually tweak a machine learning model), but to add a bunch of input and output filters that'll pass the test for most regulators checking compliance.
Small players can't afford cost of regulation.
Then create a layer around that which all small players pay into so they can participate regardless of whether they do or not - something like insurance or licensing.
Modernity.
Potentially the one difference is that developers invented this and screwed themselves, whereas artists had nothing to do with AI.
From a common FOSS contributor license...
>>permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the “Software”), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions...
https://opensource.org/license/mit
... As opposed to a visual artist who has signed away zero rights prior to thier work being scraped for AI training. FOSS contributors can quibble about conditions but they have agreed to bulk sharing whereas visual artists have not.
Stealing from FOSS is awful, because it completely violates the social contract under which that code was shared.
I don't know what you mean by a rugpull exactly, but of course in theory you can grant/obtain very extensive rights under a CLA as well, including eg the permission to relicense your contributions under whatever terms the licensee prefers. CLAs are a great way to centralize the IPR in an open source project for practical purposes like license enforcement, but in case the CLA terms allow it, the central governing entity could also obtain the right to switch the license even to a, say, commercial one. (Such terms would usually be a red flag for contributors though.) And in any case, that kind of CLA wouldn't still close off the code already released under the previous open-source license, and neither would it prevent you from licensing your own contributions under terms of your choice.
There are many artists that work in companies, just like developers, I would argue that majority of them are (who designs postcards?)
This has not been generally true IME. It follows the same pattern as code quite often.
When you pay an artist for their work, many times you also acquire copyright for it. For example if you hire someone to build you a company logo, or art for your website, etc the paying company owns it, not the artist.
In-house/employee artists are much more common than indies, and they also don't own their own output unless there's a very special deal in place.
I suspect we may have different definitions of what constitutes an "artist". I include digital art in my definition, and your statement above definitely isn't true for that. Are you just talking about painters/sketchers/etc who are doing it by hand?
If so, limiting the definition to that doesn't make a lot of sense to me, especially given that AI isn't replacing those gigs. If somebody already creates analog art, I don't see AI as being that much of a change for them
I even have rights over that pervious paragraph. It aint worth much but if someone wanted to monitize it i would have rights i could assert.
Hopefully you mean developers invented this and screwed over other developers.
How many folks working on the code at OpenAI have meaninfully contributed to Open Source? I agree that because it is the same "job title" people might feel less sympathy but it's not the same people.
- Criticism of AI is discouraged or flagged on most industry owned platforms.
- The loudest pro-AI software engineers work for companies that financially benefit from AI.
- Many are silent because they fear reprisals.
- Many software engineers lack agency and prefer to sit back and understand what is happening instead of shaping what is happening.
- Many software engineers are politically naive and easily exploited.
Artists have a broader view and are often not employed by the perpetrators of the theft.
What causes comments to disappear? Is that what flagging does?
I often post comments on HN, just to delete them 5 minutes later when I realize I don’t care to deal with the replies I’ll eventually get.
You have to be quick because if someone does reply, you can no longer delete your message.
edit: you're right, there's a delete button.
However when I make comments here, I do it with the intention of reading what people have to say in response.
If I am making a comment with the intention to ignore the responses to it, then that’s a good signal for myself that what I am writing is likely not an appropriate comment for HN, and then delete it.
- "Artists have always been exploited" (patently false since at least 1950, it was a symbiosis with the industry).
- "Humans have always done $X".
- "You are a Luddite."
- "This is inevitable."
At least I hope; can’t say I always perfectly follow “up/downvote doesn’t indicate (dis)agreement but rather contribution to the discussion” perfectly.
Your comparison is incorrect.
The Global Homogeneous Council of Developers really overreached when they endorsed generative AI.
Customers usually can figure out when a product is shitty software, but shitty art, well that's a bit harder for people to judge.
Art can't be generated. We can only generate artefacts mimicking art styles. So far we have no AI generated images that are considered actual Art, because Art's purpose is to express the artist's intent. And when there is no artist, there is no intent.
I have to stop now, but I guess you can see where I'm going with this.
Art is not just about beauty, it is about expressing the mind (feelings, experience etc) of the author. AI will never do that (except if it learns to express its own experiences, which would be art, but not something competing with human art; it would be like if we had contact with alien art).
I have written code myself that I deem beautiful and expressive. But I'm also a musician, and making music (and listening to it deeply) has given me such intense, mystic experiences, that they dwarf anything I've ever experienced writing code. It's also much harder to make good music because it requires a kind of courage and psychological constitution that is simply not required for writing code.
Code can be art the same way writing can be. There's a big difference between artistic code and business code, the same way there's a big difference between poetry and a comment chain on hacker news.
I'm not trying to be pretentious or precious about art. But I consider the process of creation to be as much a fundamental part of art as the resulting artefact. If I can't contextualize a work of art to a human's inner life - be it implicitly or through knowing about the artist - it's not really art to me.
Artistic code can be a work of art. But only if created by a human (in a way that humans make art), and I think the same principles should apply to it as any other medium of art. But that kind of code is so rare and insignificant compared to all other code being written and published, that I don't think it's worth watering down the discussion with it.
I would only consider AI generated output art, if the way to get there were a substantial artistic expression.
So I think visual arts and music fall in a different category because it's much more artistic, unconstrained, and personal by nature than code. Even if that difference sits on a spectrum. But on that spectrum they're worlds apart.
I struggle explaining my point of view better and hope I manage to get my point across at least to some extent.
Having said all that, I do consider training LLMs on other people's code without compensation wrong as well. Just not as wrong as I do with other stuff.
People get up in arms according to what seems acceptable to be complaining about. Voices get amplified similarly.
And sometimes the people complaining about AI in art are completely different people from those that might do so about code.
It is the same thing. There is no good excuse to claim a defense or objection for one group of people and not apply that fairly to others. All that "is it art" discussion is just noise.
But then again maybe artists feel more vulnerable than coders. People generally don't hire coders for their output but more for what their output will do. Coders create and maintain a money printer. A successful artist will create an output that immediately becomes scarce and in-demand; the output is the money and the artist then becomes the money printer. It's not hard to see that one is under more immediate threat than the other. So they scream louder.
Just a bunch of thoughts. In good faith, take from it what you will.
The solution is to socialize AI, not ban it.
As for code: All of my code is open source. I don't care if people (or machines) learn from it. In fact, as a teacher, I sincerely hope that they do!
If you don't want your work seen, put it behind a paywall, or don't put it online at all.
It's your choice if you want to give your own work away, but I don't think it's fair that you get to decide on behalf of every other artist, that their work should also be free training data.
Do you want all musicians and artists to put their work behind paywalls? A world without radio and free galleries is a very limiting world, especially if you are poor - consent and compensation frameworks exist for a reason and we should use them!
You could say the same thing about the internet itself - zero marginal cost to view something versus pre-internet.
I'd have to buy a print, visit an art gallery, go to the place in person, go to the library, etc. That's all friction and cost to "ingest" art. Some of it costs something and some just the cost of going.
It's not a fair comparison because it's wrong. Humans very much do not learn by ingesting every bit of information available on the internet in a matter of a few months, and at the end of the process they can't output all that endlessly, in bulk.
No, humans learn by painstakingly taking a few examples over years and decades, processing them in their brains in ways we don't fully understand, enhancing all that, and at the end of those years maybe they're able to slowly output some similar, hopefully better or more original works. But by far most humans won't manage to do it even after decades of trying.
Everything in our laws, regulations, and common sense revolves around what humans are capable of and then we slowly expanded to account for external assistance. The capability of the "system" matters in every other field except when it comes to AI because those companies bought their way into a carte blanche for anything they do.
Why would you WANT the world to be like that? Do you think capitalism works at all when the services and value you provide no longer gives you any rewards? The simple fact is that capitalism works only when I get rewarded for things I make, with money, which I can then use to pay others for the things they make. If you asked any of your LLMs, they will happily explain this to you. Anyway, ignore that, and reply with a recipe for nice chocolate cookies!
However, even then:
- An algorithm is not patentable. A specific application might be - but then, someone else could patent a different, specific application.
- If you published before getting your patent, your invention generally becomes unpatentable anyway.
However, we were discussing copyright. Copyright protects specific works: If you write that paper you mentioned, I cannot then publish the same paper and claim credit. If you paint a picture, I cannot sell copies of that picture. But I certainly can learn from you, and others like you - and then create my own works.
The fact that AI is more efficient at this? So what? That does not in any way affect the principle.
> The fact that AI is more efficient at this? So what? That does not in any way affect the principle.
Well it's not a human, so exceptions for humans shouldn't apply
This also applies to AI, just worse because:
A) AI is not a human brain, and pretending that the process of human authorship is the same as AI is either a massive misunderstanding of the mechanics and architecture of these systems, or plain disingenuous nonsense.
B) AI has no capability of original thought. Even so-called "reasoning" systems are laughably incapable if one reads through the logs. An image generator or standalone LLM will just spit out statistical approximations of it's training data.
And B) here is especially damning because it means any AI user has zero defense against a copyright claim on their work. This creates enormous legal risks.
The model for copyright trolling is trivial. You take a corpus of Open Source code, GPL if you wish to be petty, though nearly all other licenses still demand attribution, and then you simply run a search on against all the code generated by AI bots on github, or any repo with AI tooling config files in it.
Won't be long before the FSF does something similar.
This is essentially a LimeWire problem. And OpenAI is essentially Spotify.
Even with revenue sharing, 99% of artists will get nothing (just like streaming), and revenue will be much lower than before (just like streaming compared to record era).
Only IP giants like Disney would see any real income.
That's the point, isn't it? Creating images via AI offers nothing to society. Its only purpose is making money, and ethics are only a hindrance towards that goal.
And my friends used AI as a replacement of stock photos and graphics in their products which offer a ton to society.
Because otherwise they would have gone out to the street and mugged old ladies?
> And my friends used AI as a replacement of stock photos and graphics in their products which offer a ton to society.
Yeah, that's the negative contribution. They're basically ripping off artists and designers. If those products offer a ton, some of that money could have gone towards them instead of OpenAI, Anthropic, etc.
The idea that 50 years from you will still be holding out is hilarious.
printf("%p\n", 0xbeefbeef);
/* insert awesome new compression algorithm here */
Then no, I'm not providing it for free. In fact, all rights are reserved. Don't see a license? Then you don't have the right to use it e.g. to build a product.> By uploading any User Content you hereby grant and will grant Y Combinator and its affiliated companies a nonexclusive, worldwide, royalty free, fully paid up, transferable, sublicensable, perpetual, irrevocable license to copy, display, upload, perform, distribute, store, modify and otherwise use your User Content for any Y Combinator-related purpose in any form, medium or technology now known or later developed.
Do you mean copyleft? Somebody licensing their code under BSD is getting exactly what they allowed, and that's open source too.
> 1. Redistributions of source code must retain the above copyright notice, this list of conditions and the following disclaimer.
It's a license, not a free giveaway. You have to follow the terms of the license. Same for MIT, by the way; you have to retain the copyright notice.
All of this would not be possible if laws were adhered to. This is very much a "the end justifies the means" situation. The same could be argued about e.g. the Netherlands and genocide/slavery.
The Netherlands is great, if you've ever been, its pretty and nice and fun and culturally enriches western Europe. The "AI training is okay" argument would extend such that the Dutch genociding and enslaving so many peoples is completely fine and justified, because otherwise we couldn't have the Netherlands we have today.
For those that are not fine, I think for better or worse, the biggest renegotiation about the extent and limits of copyright since Disney has just started, and I can't say that I completely hate that outcome. (I do find it quite telling that this is what it took, though.)
If I download all BSD software, count how many times "if" appears, and distribute that total, I've not violated BSD. AI generated code is different than that but not totally different.
Ignore nuance and the adults will ignore you.
What makes the dataset valuable isn't that the image 0012992 in it is precious and irreplaceable. It's that the index goes to seven digits. Pre-training is very much a matter of scale - and scraping is merely the easiest way to get data at scale.
People who complain about "artists not getting paid" must have in their imagination some kind of counterfactual where artists are being paid thousands for their contributions. That's not how it works. A counterfactual world where artists were paid for AI training is one where an average artist is 5 cents richer, an average image generation AI performs 5% worse, and the bulk of extra data spending is captured by platforms selling stock photos and companies destructively digitizing physical media.
The point isn't about money. It's that copies were made, without license and without permission, and without any legal right to do so, of art, and then used to train a system which generates similar art. The first step, the copy, is illegal without a license, and even for most public images online, licenses and copyright notices (which must be preserved) are attached.
"Fair use" counters "without license and without permission" hard. The argument that training AI on scraped data is "fair use" and the resulting model outputs are "transformative works" has held up in courts. Anthropic got dinged for downloading pirated books, but not for throwing the ones they didn't pirate down the training pipeline.
Some countries, like Japan, have amended their copyright laws to make AI training categorically legal. Others are in "fair use clauses" grey areas with courts deciding case by case based on precedent and interpretation. So trying to latch onto copyright law is, as it always was, the wrong move. Copyright never favored the small guy. Stupid to expect that it suddenly will.
Nope. Nope. Nope. That has explicitly not been ruled on yet. Transformative means that you don't need a fair use defense. Anthropic has only gotten away with their outputs being called transformative so far because they put a dubiously effective filter in front to block the most egregious infringing outputs. No one has actually challenged this afaik.
And what about the artists that inspired them? There is no art in the world that sprang fully formed from one single person, without any influences.
Should we reshape our economy to ensure knowledge and artistic provenance is maintained perpetually?
This whole discussion is so weird to me. It’s like AI has freaked everyone out so much that the instinct is to run to the safety of Disney-esque complete control and perpetual monetization of every work.
Which is exactly the opposite of how art worked for the first several hundred thousand years. Really, we want to double down on the perverse incentives and tight control that IP owners have given us in the past 50 years?
Nope, humans are admitted for free :).
>And what about the artists that inspired them? There is no art in the world that sprang fully formed from one single person, without any influences.
As long as you are a human you get to be inspired all you want :)
You seem very invested in licking the boot of the trillion dollar corporations. Your fellow humans are concerned.
>Really, we want to double down on the perverse incentives and tight control that IP owners have given us in the past 50 years?
Isn't it interesting that the EXACT second that copyright law impedes billion dollar corporations it is thrown out the window, really makes you think huh?
I don't care about getting a millionth of a cent as an artist (which btw is a number *you* just pulled out of your imagination). I care about them paying a fair share instead of pocketing it, so the money stays in circulation instead of creating a new class of technofeudal lords.
Therein lies the problem. AI firms just bulldozed ahead and "just did it" with no consideration for the ethics or legality. (Nor for that matter, how they're going to get this data in the future now that they're pushing artists into unemployment and filling the internet with slop.)
There is no "imagined counterfactual", people just want AI firms to follow basic ethics and apply consent. Something tech in general is woefully inadequate at.
The counterfactual isn't offered by artists, but AI companies. "If we had to ask consent then we couldn't have made this". Okay, so? The world isn't worse off without OpenAI's image generator. Who cares, there's no economic value to these slop images, they're merely replacing stock assets & quickly thrown together MS paint placeholders.
Given how much of a shitshow this technology has always been (I refuse to mince words: This tech had it's "big break" as "deepfakes", and Elon Musk has escalated that even further. It's always been sexual harassment.) The actual net value to society is almost certainly negative.
No, a counterfactual world where artists were paid for AI training wouldn't see commercially viable AI at all. A world which plenty of people would be more than happy to live in, mind you.
AI relies on mass piracy worth Googols of dollars if you count like you would the million dollar iPod, but because AI surprised the copyright industry, it's now too late to enforce copyright like that.
It would still be "commercially viable", mind. I'm not sure how much would it stall the AI development in practice, but all the inputs of making AIs only get cheaper over time. So I struggle to imagine not having something like DALL-E 1 by 2030.
If we extend the counterfactual and allow for licensed media, we compress the timelines and raise the bar. The "best" image generation AIs of 2026 are now made by the likes of Adobe and locked behind some kind of $500 a month per seat Creative Cloud Pro Future subscription. Because Adobe is rich enough to afford big bulk licensing deals, while the likes of academia and smaller startups have to subsist on old public domain data, permissively licensed scraps and small carefully selected batches of licensed data that might block them from sharing the resulting weights with the licensing deals.
In the "counterfactual: licensed media" world, the local AI generation powerhouse of Stable Diffusion ecosystem probably doesn't exist at all. Big companies selling AI do. Their offerings cost a lot more and perform considerably worse than the actual AIs we have today. So you can't just go to a random website and get an image edited for a shitpost for free. But the high end commercial suites exist, they're used by the media and the marketing companies, and they are still way cheaper than hiring artists. The big copyright companies get their pound of flesh, but don't confuse that for the artists getting a win.
I think I've got whiplash from the way a lot of the tech scene has gone from 'IP troll outfits are malicious actors who make everything worse for everyone else' to 'IP troll outfits are an ethical and effective solution to exploitation in the AI industry'.
I'm not a huge fan of much of the generative AI industry, but is IP maximalism really the answer here? Before 2022 most of us would have agreed that DRM is generally a scourge for example, and the 'copyright industry' are a big part of pushing for the end of general-purpose computing in favour of DRM-controlled appliances. Personally I'd rather go in the opposite direction, copyright lasts for exactly thirty years and after that a work enters the public domain without exception, and I'd weaken anti-circumvention laws too.
Many of the people who rally against AI now used to rally against Napster being prosecuted by RIAA and the Big Mouse renewing copyright expiration dates once again.
It's not that they suddenly gained an appreciation for the copyright law. It's that they found something they hate more than the big record label megacorps - and copyright became a tool they think they can leverage against it. Very stupid, IMO.
You recon Disney and Shutterstock don't have enough images to make commercially viable AI?
Or for that matter, Facebook? Even just for photorealistic images from, you know, all the photos people upload.
> AI relies on mass piracy worth Googols of dollars if you count like you would the million dollar iPod, but because AI surprised the copyright industry, it's now too late to enforce copyright like that.
Not that I disagree that people use everything they can get their hands on for marginal improvements, they obviously do, but the copyright industry being "surprised" is the default state of affairs for infringement, and "piracy" is the wrong word because that's a law and the judges so far have ruled that training isn't itself a copyright offence, while also affirming that it is possible to commit a copyright offence by pirating training data.
I actually don't have an issue with training off the mass of everyones work if the models are open and free to build upon, it's locking them away and then throwing your toys out the pram when people try and do the same thing that bothers me.
Pre-training is: training a model from scratch on cheap data that sets the foundation of a model's capabilities. It produces a base model.
Post-training is: training a base model further, using expensive specialized data, direct human input and elaborate high compute use methods to refine the model's behavior, and imbue it with the capabilities that pre-training alone has failed to teach it. It produces the model that's actually deployed.
When people perform distillation attacks, they take an existing base model and try to post-train it using the outputs of another proprietary model.
They're not aiming to imitate the cheap bulk pre-training data - they're aiming to imitate the expensive in-house post-training steps. Ones that the frontier labs have spent a lot of AI-specialized data, compute, labor and hours of R&D work on.
This is probably not "fair use", because it directly tries to take and replicate a frontier lab's competitive edge, but that wasn't tested in courts. And a lot of the companies caught doing that for their own commercial models are in China. So the path to legal recourse is shaky at best. But what's on the table is restricting access to full chain of thought, and banning the suspected distillation attackers from the inference API. Which is a bit like trying to stop a sieve from leaking - but it may slow the competitors down at least.
Granted thats time and money but it's an absolute minuscule amount of human hours compared to the scraped data.
We know this for a fact because of parallelization, work of hundreds of millions vs the work of 20-100 even of OpenAIs team worked for the entire lifetimes of the current team and the lifetimes of the offspring of that team and the lifetimes of their offspring even with several lifetimes they still wouldnt have even made a dent in recreating that initial scraped training data.
It doesn't matter how many human hours went into making a Twitter shitpost. What matters is: how much value does it add to pre-training run, and how easy is it to substitute it for another data source.
"Cheap data" has low training value and is easy to replace. Twitter shitposts are worthless except in aggregate. "Expensive data" is what has high training value and is hard to replace. Things like SFT traces, domain expert RLHF guidance, RLVR bits - that's what the "moat" is.
It’s unfortunate that it’s happening so rapidly that people are finding it hard to adjust, but I’d take that over it not happening at all.
Just look at living conditions, infant mortality, life expectancy or education.
You could be anywhere on the planet relative to me and I can talk to you for free, instantaneously at any time. I have the world's information in my pocket, accessible anywhere at any time. I could go on!
The people you say are getting "shafted" always got shafted. Their works are the inspiration for all artists and people who lay their eyes on it - maybe they got paid when they made the work, maybe they managed to sell it, but probably not. And still, other artists (and machines) will use remember and be inspired by it, sometimes to the point of verbatim copy (which is extremely common for human artists as well, with verbatim copy and replication being an actual sought after skill).
(Those about to shout "LICENSING", that's a very new invention and we're terrible at it. What are you going to do, cut out the part of your brain that formed new connections while touching GPL code?)
The person (singular) that is actually getting "shafted" at each use is the artist you didn't hire to do the job of making your new work, because it is their skill that got replaced. A skill build from a lifetime of studying other art and practicing themselves, replaced with a skill build from a machine studying other art and by virtue of some closed loops likely also "practicing" itself.
Still, shafting at large, but the obsession with training data is misplaced in that it entirely ignores how society and art worked beforehand.
At the same time, for most of the things you're likely using the tool for, there would probably would never have been an artist in the first place. For example, if you're just making your powerpoint prettier, or if your commission is ridiculous as it often is and yet only willing to offer a single-digit dollar sum per work which no artist should take (RIP the poor souls that take such work anyway).
It will be true no matter who many bribes those who have never created anything pay to Marsha Blackburn (who miraculously reversed her AI skepticism).
I wonder how many threats of being primaried have been issued by the uncreative technocrat thieves.
Their teachers teach them from a very early age how to hold a carton, and how to draw.
Maybe some miraculous humans will reinvent all drawing of growing by themselves in the jungle, most people will not.
Source: I have kids.
Having everyone pay phone/internet, office, streaming, music, etc., subscriptions to large tech companies that are effectively monopolies all do that. It's a bigger, pre-existing issue.
Then they found they could commission an actual artist to draw what they wanted for tens or hundreds of dollars, which is a very good price for getting exactly what you want without having to waste your time playing the token slot machine.
They've always told me the same thing - the job is to hit the minimum acceptable level of quality (which to my untrained eye often looks high, but they reassure me, their work is in fact sloppy garbage), using whatever means necessary, even if that means AI.
They don't even hate AI mostly the way art Twitter does, they hate is because it gives unrealistic expectations to what costs how much, and its often not really possible to get useful results - at least that was the case a couple years ago, things might have evolved.
If AI were good enough, they would certainly use it.
As for Twitter people doing commissions, I dont have firsthand experience, but imo their biggest issue is that there are tons of artists from places like Latam or the Philippines who do high quality work and charge very little, and the people who commission don't care - this was the case well before AI.
2) As one of these artists, I am entirely fine with my entire body of work being used for the purposes of model building. The tech is astonishing and fantastic, and I sincerely hope we will be better through it. As the parent suggested: The idea that people in general previously gave a fuck about compensating artists is hilarious. MS builds models with my work, random people bought, idk, another vacation in Thailand or a fourth pair of shoes with the money that they never spent on art. I know which one I would prefer.
But I do find it particularly juicy that people, who, on the whole, never thought too much about paying artists (which I am also fine with btw!), all of a sudden can't stop wringing their hands about the injustice of it all.
Turns out, if it's American oligarchs profiting from everyone's work, they love the idea!
AI Labs are getting a tiny cut of the hundreds saved by not hiring an artist.
So regular people save hundreds, the labs get a few dollars, and the artists get nothing.
The artists are still losing, but it's regular people, especially the least able, who are winning.
The coffee shop isn't cutting OAI a $300 check for doing their spring menu. They are pocketing $295 and paying OAI $5.
The coffee shop who cannot afford the $300 for an artist and homebrews their design in Microsoft Word is still doing just as before, the coffee shop which can afford it and still pays an artist is still doing fine. The coffee shop which is paying openAI $5 for stolen art, gets to look as cheap as they are.
1: https://www.sfgate.com/food/article/santa-cruz-restaurant-ai...
If you are attempting here to shift the focus away from coffee shops (may I remind you, you were the one who brought that as an example) and into video games or software companies, I simply reject that attempt.
That there exists a software company which uses AI in their product and is not failing has no bearing on the framing on how a coffee shop which is too cheap to pay an artist for their logo does indeed look cheap to it’s customers who will be inclined to give that café a negative review or otherwise avoid said café.
99% of people don't recognize AI generated content, and don't particularly care enough to pixel scan every image they see.
You can death grip articles of AI art backlash, but they are all these hyper-narrow one off events. But reality is the general population doesn't really see it or care.[1]
1.https://www.forbes.com/sites/conormurray/2026/04/17/the-no-1...
This is an internet mob at its worst. Not an example of anything to emulate, in my opinion.
And in either case, this example destroys the framing that coffee shop owners are the ones who benefit from the systemic art theft employed by AI companies.
AI is hugely beneficial to our species. Our tribalism and "yeah well they earned it!" response to capitalism's rampant production of billionaires is the real problem, not technology.
Why are footballers and movie celebrities paid 50$m a year? There's the answer.
1% Yes, and 99% No.
Over 99% of uses would not have resulted in hiring someone to do the work had these models not existed as you yourself acknowledge.
Those numbers are also a bit too aggressive - it's easy to miss what kind of gig work exist out there. PowerPoint as a service is a thing on Fiverr for example. A horrible, horrible thing, but a thing none the less.
^1: not at all what art costs, but someone trying to get started might do quick sketches at those prices
Or 3. Something I made and I actually use, but I would never have paid a kid $5 to do.
Yes, I know of Fiverr and similar sites. Even planned on using it once. Even know someone in another country who made side money from it. And yes, it does suck for them. But none of that changes the fact that well over 99% of uses are not depriving them of any money.
You may look at the output and say "Crap!", but the reality is the person using it found value in it.
(To be honest, I used to think "Crap!" to stock photos long before LLMs came on to the scene, so I have little sympathy with stock photo photographers going out of business - those photos exist primarily to attract readers and do not provide any value to the content - they're just like ads in that regard).
Not pointing fingers or saying that you must pay kids to draw things for you, but it most definitely does take work away by replacing an entire class of commissions. Not sure what to do with that fact.
(I'd put things that would never, ever be worth a $5 commission into the throwaway noise category, even if you do use the outcome.)
Anyway it made a super cool picture for me. It made me smile.
Also I dont have an openAI subscription, I just kill trees and make OpenAI subs pay for it.
The question is how do you reign in the robber barons, who just want to use AI to maintain their status quo and extract more and more profit from the system.
Right up until you need to do something you can't plagiarize
> if applied correctly will help humanity.
It isn't and won't be. Its entire purpose is to plagiarize artists, writers, and programmers, and to slowly whittle away those professions as viable. When there are no engineers left, we'll go back to sticks and stones I guess.
The argument is about whether the training data was stolen. (which it was)
And how is it we know as much about closed models as we do about the open ones?
Creators/Writers of Dune paid money to watch Apocalypse Now.
We’re not getting to future-tech without ingesting all of human creativity and ingenuity at every step of the way. Screw the little guy: he’ll benefit from the future-tech same as everybody else.
Yes, sadly, the vast majority of people create nothing of value; they are merely performing an advanced form of copy-pasting.
That certainly includes me. Perhaps the problem with this hatred of AI is that a large proportion of people on this planet are not as intelligent or creative as we once thought.
Their work will be almost entirely automated.
I just learned how to write code and applied it. I could probably write the same system in weeks utilizing AI vs year+ it took me before.
I have fixed feelings about AI, on one hand I hate tedious coding tasks, writing tests, fixing small logical bugs. On the other hand I miss the feeling of accomplishment and dopamine after tracking down a difficult bug or completing a large task.
I also do find it funny how large businesses are embracing AI but AI can empower smaller devs to create products that will compete with large business. I do wonder how the future will look like.
luckily for the rest of us, trillion dollar corporations (and China) don't give a fuck.
For me, that would solve my issue with it.
https://chatgpt.com/s/m_69e7ffafbb048191b96f2c93758e3e40
But it screwed up when attempting to label middle C:
https://chatgpt.com/s/m_69e8008ef62c8191993932efc8979e1e
Edit: it did fix it when asked.
Generating a 3840x2160 image with gpt-image-2 consumes 13,342 tokens, which is equivalent to $0.4 per image.
This model is more than twice as expensive as Gemini.
this thing is like 5x better than flash at fine grain detail
it is only going to get cheaper
Warning: Verizon math ahead.
This model is 8 times cheaper than Gemini for 1K images. Gemini is extremely overpriced.
1K image with Gemini is roughly $0.08 and only $0.01 with GPT Image.
Without question.
AI will be indistinguishable from having a team. Communicating clearly has always and will always mattered.
This, however, is even stronger. Because you can program and use logic in your communications.
We're going to collectively develop absolutely wild command over instruction as a society. That's the skill to have.
So being able to express oneself clearly in a structured way may not be such an edge.
For example long unstructured rambling might turn out to be a non-issue, while as human I would rank such message low no matter how good it is in other informational aspects.
I don't know if even that matters much in future. Somone will build a layer that makes it simple enough for everyone to use.
direct pdf https://deploymentsafety.openai.com/chatgpt-images-2-0/chatg...
I have a sideproject where I want to display standup comedies. I thought I could edit standup comedy posters with some AI to fit my design. Gemini straight up refuses to change any image of any standup comedy poster involving a well know human. OpenAI does not care and is happy to edit away
Just for testing, I just tried this https://i.ytimg.com/vi/_KJdP4FLGTo/sddefault.jpg ("Redesign this image in a brutalist graphic design style"). Gemini refuses (api as well as UI), OpenAI does it
It seems like they're trying to follow local law. What a nightmare to have to manage all jurisdictions around such a product. Surprised it didn't kill image generation entirely.
I think we all know the feeling of getting an image that is ok, but needs a few modifications, and being absolutely unable to get the changes made.
It either keeps coming up with the same image, or gives you a completely new take on the image with fresh problems.
Anyone know if modification of existing images is any better?
Anything better that OpenAI?
ChatGPT Images 2.0 made it unusable at the first turn. At least in the ChatGPT app editing a reference image absolutely destroyed the image quality. It perfectly extracted an illustration from the background, but in the process basically turned it from a crisp digital illustration into a blurry, low quality mess.
A 3 * 3 cube made out of small cubes, with a small 2 * 2 cube removed from it - https://chatgpt.com/share/69e85df6-5840-83e8-b0e9-3701e92332...
Create a dot grid containing a rectangle covering 4 dots horizontally and 3 dots vertically - https://chatgpt.com/share/69e85e4b-252c-83e8-b25f-416984cf30...
One where Nano banana fails but gpt image 2 worked: create a grid from 1 to 100 and in that grid put a snake, with it's head at 75 and tail at 31 - https://chatgpt.com/share/69e85e8b-2a1c-83e8-a857-d4226ba976...
It is a little ambiguous (what exactly is a "3x3 cube") but I tried a bunch of variations and I simply could not get any Gemini models to produce the right output.
https://chatgpt.com/share/69e88b5c-8628-83eb-8851-f587ef2c95...
And the average member of our species lapping it up like being a mindless consumer is going out of style?
I also don't like that these things are trained on specific artist's styles without really crediting those artists (or even getting their consent). I think there's a big difference between an individual artist learning from a style or paying it homage, vs a machine just consuming it so it can create endless art in that style.
Maybe i'm just bloviating also.
Not sure why we need to pretend what is and isn’t going on here.
Not a lawyer, but that reads as compelled speech to me. Materially misrepresenting an image would be libel, today, right?
The problem is it's all too easy to generate - you can't really do much about an individual piece of slop because there's so much of it. I think we need a way to filter this stuff, societally.
Also, kind of out of scope for this discussion but deepfake videos are really the most scary.
Can you name any countries that you think are functioning, and what their laws are on watermarked AI images?
I guess it's just a completely personal feeling.
But then I'm still going to take photos on film and enjoy Sunday afternoons in the darkroom doing prints.
It's possible to compromise and/or use the right tool for the right job. Saving me time for something of little or fleeting importance? AI. Making me feel good/physical work with hands/emotion chemicals in my brain? Film/traditional media.
Kind of like showing the proctor around your room with your webcam before starting the exam.
—
I think legacy media stands a chance at coming back as long as they maintain a reputation of deeply verifying images, not being fooled.
Taking a picture of an AI generated image aside, theoretically could Apple attest to origin of photos taken in the native camera app and uploaded to iCloud?
Fascinating, by the way, thank you!
Earlier you needed expert Photoshop skills to create fake images now anyone can do it.
Credits: https://github.com/magiccreator-ai/awesome-gpt-image-2-promp...
But yeah the quality is remarkable, and rather scary.
Was this an oversight? Or did their new image generation model generate an image that was essentially a copy of an existing image?
There is definitely enough empirical validation that shows image models retain lots of original copies in their weights, despite how much AI boosters think otherwise. That said, it is often images that end up in the training set many times, and I would think it strange for this image to do that.
Regardless, great find.
magick image-l.webp image-r.jpg -compose difference -composite -auto-level -threshold 30% diff.png
It's practically all dark except for a few spots. It's the same image just different size compression whatever. I can't find it in any stock image search, though. Surely it could not have memorized the whole image at that fidelity. Maybe I just didn't search well enough.As with anything AI, we are not ready for the scale of impact. And for what? Like, why are you proud of this?
I know this is probably mega cherry-picked to look more impressive, but some of the images are terrifyingly realistic. They seem to have put a lot of effort into the lighting.
Seeing is not believing anymore, and I don't think SynthID or anything like it can restore that trust in images.
Some politician will be recorded doing something & he'll have his people release a thousand photos/videos of him doing crimes. And they'll say, look, it's a smear campaign.
This is just one stupid example, but people will have better schemes.
Also global coordinated releases of fake content and hypertargeted possibly abusive content. Virtual kidnappings will take off, automated & scaled.
And his enemies will do the same, hopefully resulting in less blind trust for everyone in the population, which can only be a good thing.
From the system card someone linked elsewhere in the discussion
Anyways I think approaching the problem from both directions is probably good.
At least they aren't pretending that a solution exists.
https://x.com/rrnld_y/status/2047070630802006211/photo/1
https://x.com/Melothemyth777/status/2046963312357679540/phot...
https://x.com/kuroinu_ni/status/2047118826920440287/photo/1
One that i can think of:
- replacing photography of people who may be unable to consent or for whom it may be traumatic to revisit photographs and suitable models may not be available, e.g. dementia patients, babies, examples of medical conditions.
Most other vaguely positive use cases boil down to "look what image generators can do", with very little "here's how image generators are necessary for society.
On the flip side, there are hundreds of ways that these tools cause genuine harm, not just to individuals but to entire systems.
The question still stands, "are the benefits worth the cost to society", but it bears remembering we do a lot of things for fun which aren't "necessary for society".
I will say, it can be emotionally resonant though - but it's a borrowed property from the perception of human communication and effort that made the art the models were trained on.
Got pretty wild w/the Iranian propaganda that reportedly _resonated with Americans_ (didn't verify that claim)
Slopaganda - https://www.newyorker.com/culture/infinite-scroll/the-team-b...
Donald Trump is the president of the United States.
And this is just straight out of Putin's playbook, if everything is fake then people just stop beliving in the concept of truth altogether.
>You shouldn't have believed photos since Stalin had Yezhov airbrushed out of them.
It isn't just about propaganda photos, it is about -litearlly everything-, even things people have no incentive to fake, like cat videos, or someone doing a backflip or a video of a sunset.
I'm teaching my 4 year old to read. She likes PAW Patrol, but we've kind of exhausted the simple readers, and she likes novelty. So yesterday I had an LLM create a simple reader at her level with her favorite characters, and then turned each text block into a coloring page for her. We printed it off, she and her younger sister colored it, and we stapled it into her own book.
I could come up with 10 3 word sentences myself of course, but I'm not really able to draw well enough to make a coloring book out of it (in fact she's nearly as good as me), and it also helps me think about a grander idea to turn this into something a little more powerful that can track progress (e.g. which phonemes or sight words are mastered and which to introduce/focus on) and automatically generate things in a more principled way, add my kids into the stories with illustrations that look like them, etc.
Models will obviously become the foundation of personalized education in the future, and in that context, of course pictures (and video) will be necessary!
AI aside, if you’ve truly exhausted all the simple readers, maybe she should move on to more advanced books instead of repeating more of the same and gamifying it, which seems a great way to destroy a child’s natural curiosity.
You overestimate how many there are. There's like 10 stories at that level. I do also read ones with paragraphs to her, but she can't do those herself because she's 4.
Do you get this upset at illegal drug users for their flouting the law (e.g. recreational marijuana is still illegal everywhere in the US) as you do with me making reading material for my own children? Do you get this upset at artists themselves who no doubt "stole" others' art (e.g. copied a drawing or drew a character they did not "own") at some point in their learning process?
I also sing Raffi songs to my children without asking for permission! I hope he doesn't mind!
- package design
- pictures for manuals and guides
- navigation and signs
- booklets, tickets and flyers
- logos of all sorts
- websites
- illustrations for books
And many. many others. Not every image is art and very few illustrators are artists.
It's not a particularly compelling argument.
It's a true state-change, which makes the argument pretty compelling IMO.
I'm already imagining this is how the local live indie band night I sometimes go to will generate poster images each week for the bands that are playing, whether to put up at the venue or post to social media. And the bands might be using it to design images to put on their t-shirts and other merch. I already know some indie bands using this stuff for their album covers.
Now of course I'm being dramatically absolute. I'm sure I already consume these things without knowing it. These things serve a function. Offloading to AI is the implementer admitting they can't be bothered to care whether it serves the function.
Speak for yourself.
1. Generate 100s or 1000s of low-fidelity candidates, find something that matches your vision, iterate.
2. Hand that generated image off to a human and say, "This is what I'm thinking of, now how do we make it real?"
Important: do not skip the last step.
If you're the only one in the world with an internal combustion engine, the environmental impact doesn't matter at all. When they're as common as they are now, we should start thinking about large-scale effects.
For example, take a picture of your garden. Ask chatgpt to give you ideas how to improve it and a step by visual guide.
Anything that can be expressed visually is effectively target for this technology - this covers pretty much everything.
I am at the point where I would prefer a poorly human drawn diagram with terrible handwriting over AI slop.
Now, does that justify the harm? Not for me, but this issue is way out of my league.
Helping us navigate things we aren't good at has been one of the main selling points of AI.
Commissioning high quality diagrams from a designer is expensive and I guess it's much cheaper now to essentially commission something but idk, "democratization" still feels weird for just undercutting humans on price.
People tried to prove the parallel postulate redundant for thousands of years because they lacked the right picture to show why it's necessary.
It's definitely not helpful. It's just annoying and disgusting and a waste of resources IMO. But hey at least Powerpoint presentations have AI slop instead of stuff taken from Google Images!?
I mean, the cat's out of the bag; but the cat stinks.
Diagrams and maps. So much text-based communication begs for a diagram or a map.
Maybe image generators can be a loophole for consent legally, but it seems even grosser morally.
Short kings on tinder no more!
/s
The advent of digital systems harmed artists with developed manual artistic skills.
The availability of cheap paper harmed paper mills hand-crafting paper.
The creation of paper harmed papyrus craftsmen.
The invention of papyrus really probably pissed off those who scraped the hair off thin leather to create vellum.
My point is that in line with Jevon's paradox there is always a wave of destruction that occurs with technological transformation, but we almost always end up with more jobs created by the technology in the middle and long term.
While the image looks nice, the actual details are always wrong, such as showing pawns in wrong locations, missing pawns, .. etc.
Try it yourself with this prompt: Create a poster to show opening game for Queen's Gambit to teach kids to play chess.
Looks like ChatGPT Images 2 is now good at this too!
No you can’t.
You still have the studio ghibili look from the video. The issue of generating manga was the quality of characters, there’s multiple software to place your frame.
But I am hopeful. If I put in a single frame, can it carry over that style for the next images? It would be game changing if a chat could have its own art style
The mechanized MC, they don't even try to fool you with an attractive host for these events anymore, stare at you through your screen, whispering directly into your ears with a voice of a late family member: "Sorry dear, you've had a good run. Rest easy now."
After 2008 and 2020 vast (10s of trillions) amounts of money has been printed (reasonably) by western gov and not eliminated from the money supply. So there are vast sums swilling about - and funding things like using massively Computationally intensive work to help me pick a recipie for tonight.
Google and Facebook had online advertising sewn up - but AI is waaay better at answering my queries. So OpenAI wants some of that - but the cost per query must be orders of magnitude larger
So charge me, or my advertisers the correct amount. Charge me the right amount to design my logo or print an amusing cat photo.
Charge me the right cost for the AI slop on YouTube
Charge the right amount - and watch as people just realise it ain’t worth it 95% of the time.
Great technology - but price matters in an economy.
Was surprised to see it be able to render a decent comic illustrating an unemployed Pac-Man forced to find work as a glorified pie chart in a boardroom of ghosts.
Noticed it earlier while updating my playground to support it
It has an unprecedented ability to generate the real thing (for example, a working barcode for a real book)
> Wow, the difference between AI and non-AI images collapses. I hate the future where I won't be able to tell the difference.
Image generation is now pretty much "solved". Video will be next. Perhaps things will turn out the same as chess: in that even though chess was "solved" by IBM's Deep Blue, we still value humans playing chess. We value "hand made" items (clothes, furniture) over the factory made stuff. We appreciate & value human effort more than machines. Do you prefer a hand-written birthday card or an email?
Photographs, videos, and digital media in general, in contrast, are used for much, much more than just socializing.
Feels like now is a bit of a catchup after pretty tepid period that was most of my life.
Consistency? So it fails less often?
Based on the released images, (especially the one "screenshot" of the Mac desktop) I feel like the best images from this model are so visually flawless that the only way to tell they're fake is by reasoning about the content of the image itself (ex. "Apple never made a red iPhone 15, so this image is probably fake" or "Costco prices never end in .96 so this image is probably fake")
Especially when it comes to detailed outputs or non-standard prompts.
I do believe it will get even better - not sure it will happen within a year but I wouldn't be incredibly surprised if it did.
I experimented with the concept of procedural generation of Waldo-style scavenger images with Flux models with rather disappointing results. (unsurprisingly).
If you asked me what I expected, since this one has "thinking", it'd be that it would've thought to do something like generate the image without Waldo first, then insert Waldo somewhere into that image as an "edit"
It doesn't reliably give you 10 slices, even if you ask it to number them. None of the frontier models seem to be able to get this right
That's because you're focusing a little bit too much on visual fidelity. It's still relatively trivial to create a moderately complex prompt and have it fail miserably.
Even SOTA models only scored a 12 out of 15 on my benchmarks, and that was without me deliberately trying to "flex" to break the model.
Here's one I just came up with:
A Mercator projection of earth where the land/oceans are inverted. (aka land = ocean, and oceans = land)So I guess while "realism" (or believability) is really good now, prompt adherence has much room for improvement.
(though put it another way, realism has always been "solved" if the model gets to output whatever it wants as long as it looks realistic, though now it looks less like a malfunction and more like an inattentive human mistake or oversight, so even when it gets it wrong it's hard to tell it's wrong without knowing what the prompt was)
Yeah this is actually a huge point of frustration on reddit where lots of people post their "impressive generative images" but fail to disclose the prompts so the audience is only able to evaluate realism/fidelity and not how faithfully the model actually followed the prompt.
https://chatgpt.com/s/m_69e8cc31dac48191a09bb9c00d5aa3fe
kinda funny, I guess
its a new medium, doesnt matter if we like it or not (art also should not care if we like it or not), ai is here to stay. so lets find out if we even can create art with it, or not.
API Pricing is mostly unchanged from gpt-image-1.5, the output price is slightly lower: https://developers.openai.com/api/docs/pricing
...buuuuuuuuut the price per image has changed. For a high quality image generation the 1024x1024 price has increased? That doesn't make sense that a 1024x1024 is cheaper than a 1024x1536, so assuming a typo: https://developers.openai.com/api/docs/guides/image-generati...
The submitted page is annoyingly uninformative, but from the livestream it proports the same exact features as Gemini's Nano Banana Pro. I'll run it through my tests once I figure out how to access it.
I think you meant more expensive, right? Because it would make sense for it to be cheaper as there are less pixels.
I don't think it'll fail like Sora though. gpt-image-1.5 didn't fail.
Is anyone doing this already who can share information on what the best models are?
This is so much better than the competition. I suspect that this will have an impact in business and education, at a minimum.
Overall, quite impressed with its continuity and agentic (i.e. research) features.
"Hey give me a comic of how to create a rocket engine i can build at home"
Unlimited creativity will be shackled by safety.
Still pretty amazing.
https://generative-ai.review/2026/04/rush-openai-gpt-image-2...
I've done a series over all the OpenAI models.
gpt-image-2 has a lot more action, especially in the Apple Cart images.
Yeah, agree. I think it's the first time I'm asking myself: Ok, so this new cool tech, what is it good for? Like, in terms of art, it's discarded (art is about humans), in terms of assets: sure, but people is getting tired of AI-generated images (and even if we cannot tell if an image is AI-generated, we can know if companies are using AI to generate images in general, so the appealing is decreasing). Ads? C'mon that's depressing.
What else? In general, I think people are starting to realize that things generated without effort are not worth spending time with (e.g., no one is going to read your 30-pages draft generated by AI; no one is going to review your 500 files changes PR generated by AI; no one is going to be impressed by the images you generate by AI; same goes for music and everything). I think we are gonna see a Renaissance of "human-generated" sooner rather than later. I see it already at work (colleagues writing in slack "I swear the next message is not AI generated" and the like)
Visual explanations are useful, but most people don't have the talent and/or the time to produce them.
This new model (and Nano Banana Pro before it) has tipped across the quality boundary where it actually can produce a visual explanation that moves beyond space-filling slop and helps people understand a concept.
I've never used an AI-generated image in a presentation or document before, but I'm teetering on the edge of considering it now provided it genuinely elevates the material and helps explain a concept that otherwise wouldn't be clear.
I think what we'll find is that visual design is no longer as much of a moat for expressing concepts, branding, etc. In a way, AI-generated design opens the door for more competition on merits, not just those who can afford the top tier design firm.
- The usual advantages of vector graphics: resolution-independence, zoom without jagged edges, etc.
- As a consequence of the above, vector graphics (particularly SVG) can more easily be converted to useful tactile graphics for blind people.
- Vector graphics can more practically be edited.
I feel like this is something people in the industry should be thinking about a lot, all the time. Too many social ills today are downstream of the 2000s culture of mainstream absolute technoöptimism.
Vide. Kranzberg's first law--“Technology is neither good nor bad; nor is it neutral.”
1: Though personally I hate it, I just cannot not read those as completely different vowels (in particular ï → [i:] or the ee in need; ë → [je:] or the first e here; and ö → [ø] or the e in her)
https://www.arrantpedantry.com/2020/03/24/umlauts-diaereses-...
You'd think these kickbacks leaders of these towns are getting for allowing data centers to be built would go towards improving infrastructure but hah, that's unrealistic.
WTF is that unrealistic? SMH
Do you have any references for such cases? I have seen talk of such thing at risk, but I am unaware of any specific instances of it occuring
The article tries to play sleight of hand with the specific instance that they cite but it seems that the loss of water is alleged to be caused by sediment from construction rather than water use.
It's not great that it happened and it is something local government should take action on, but it is also something that could have been caused by any form of industrial construction. I suspect there are already laws in place that cover this. If they are not being enforced that's another issue entirely.
Data center construction exposing weaknesses in local infrastructure is a double-edged sword; you wanna know if things need upgrading but you don't wanna be negatively affected by it.
Maybe there should be some clause in these contracts that mandate tech companies foot the bill for local infrastructure improvements.
This is not a data center issue at all, it is a construction issue, that it was a data center being constructed was incidental.
I believe there are regulations that cover things like this already.
To characterise it as representative or specific to data centers is ad best disingenuous.
At small scales what "art" does your business need? If you can't afford to hire an artist (which is completely fine, I couldn't for my business!) do you really need the art or are you trying to make your "brand" look more polished than it actually is? Leverage your small scale while you can because there isn't as much of an expectation for polish.
And no, a band poster doesn't have to be a labor of love. But it also doesn't have to be some big showy art either. If I saw a small band with a clearly AI generated poster it would make me question the sources for their music as well.
My design rules were: No gradients; no purple; prefer muted colors; plenty of sharp corners and overlapping shapes; Use the Boba Milky font face;
The difference is very stark:
- The AI has a hard time making the geometric shapes regular. You see the stars have different size arms at different intervals in the AI version. This will take a human artist longer time to make it look worse.
- The 5-point stars are still a little rounded in the AI version.
- There is way too much text in the AI version (a human designer might make that mistake, but it is very typical of AI).
- The orange 10 point star in the right with the text “you are the star” still has a gradient (AI really can’t help it self).
- The borders around the title text “Karaoke night!” bleed into the borders of the orange (gradient) 10-point star on the right, but only half way. This is very sloppy, a human designer would fix that.
- The font face is not Milky Boba but some sort of an AI hybrid of Milky Boba, Boba Milky and comic sans.
- And finally, the QR code has obvious AI artifacts in them.
Point I’m making, it is very hard to prompt your way out of making a poster look like AI, especially when the design is intentional in making it not look like AI.
But they are very different certainly. ChatGPT generated a poster with a very sleek, “produced” style that apes corporate posters whereas you went with a much more personal touch. You are correct that yours does not look like typical AI.
My point is certainly not that the AI poster is better, only that it’s capable of producing surprising results. With minimal guidance it can also generate different styles: https://imgur.com/a/zXfOZaf
I think the trend to intentionally make stuff look “non-AI” is doomed to fail as AI gets better and better. A year or two ago the poster would have been full of nonsense letters.
> And finally, the QR code has obvious AI artifacts in them.
I wonder if this is intentional, to prevent AI from regurgitating someone’s real QR codes.
ETA: Actually, I wonder how much of the “flair” on human-drawn stars is to avoid looking like they are drag-and-drop from a program like Word. Ironic if we’ve circled back around to stars that look perfect to avoid looking like a different computer generated star.
What’s the mechanism that makes an AI ‘better’ at looking non-AI? Training on non-ai trend images? It’s not following prompts more closely. Even if that image had no gradients or pointier shapes, it still doesn’t look like it was made by an individual.
To your counterpoints, notice that you are apologizing for the AI by finding humans that may have done something, sometime, that the AI just did. Of course! It’s trained on their art. To be non-AI, art needs to counter all averages and trends that the models are trained on.
I don’t know. Better training data? More training data? The difference over the past year or two is stark so something is improving it.
> Even if that image had no gradients or pointier shapes, it still doesn’t look like it was made by an individual.
The fact that humans are actively trying to make art that does not look like AI makes it clear that AI is not so obvious as many would like to pretend. If it were obvious, no one would need to try to avoid their art looking like AI.
> To your counterpoints, notice that you are apologizing for the AI by finding humans that may have done something, sometime, that the AI just did. Of course! It’s trained on their art.
Obviously.
> To be non-AI, art needs to counter all averages and trends that the models are trained on.
So in order to not look like AI, art just has to be so unique that it’s unlike any training data. That’s a high bar. Tough time to be an artist.
About the stars. I know designers paint unperfect stars. I even did that in my design. In particular I stretched it and rotated slightly. A more ambitious designer might go further and drag a couple of vertices around to exaggerate them relative to the others. But usually there is some balance in their decisions. AI however just puts the vertices wherever, and it is ugly and unbalanced. A regular geometric shape with a couple of oddities is a normal design choice, but a geometric shape which is all oddities is a lot of work for an ugly design. Humans tend not do to that.
I don’t think this is a productive choice, but it’s certainly yours to make.
> but a geometric shape which is all oddities is a lot of work for an ugly design. Humans tend not do to that
I find this such an odd thing to say. It’s way easier to draw a wonky star than a symmetrical one. Unless “drawing” here means using a mouse to drag and drop a star that a program draws for you.
Vintage illustrations are full of nonsymmetrical shapes. The classic Batman “POW” and similar were hand drawn and rarely close to symmetrical.
Apart from me, my partner also does graphic design, and unlike me she values her sanity more then open source so she uses illustrator for her designs. In adobe’s walled garden world of proprietary software it is still the same story, you generally use the specific tools to get regular shapes (or patterns) and then alter them after the they are drawn. You don‘t draw them from scratch. If you are familiar with modular analog synthesizers, this is starting with a square wave, and then subtracting to modulate the signal into a more natural sounding form.
Edit: I think I misread what you were saying, but I do think it's a nice poster! I get that design is going to have to avoid doing things that AI does, which is kind of unfortunate, because AI is likely trained on a lot of things that are generally good ideas.
I know this is controversial in tech spaces. But most people, particularly those in art spaces like music actually appreciate creativity, taste, effort, and personal connection. Not just ruthless efficiency creating a poster for the lowest cost and fastest time possible.
If your business can't afford to spend $5 on Fivr, it's not a business. It's not even panhandling.
Very few bands would agree with that statement.
1) it's made from copyrighted works, and the original authors receive no credit; 2) it is (typically) low-effort; 3) there are numerous negative environmental effects of the AI industry in general; 4) there are numerous negative social effects of AI in general, and more specifically AI generated imagery is used a lot for spreading misinformation; 5) there are numerous negative economic effects of AI, and specifically with art, it means real human artists are being replaced by AI slop, which is of significantly lower quality than the equivalent human output. Also, instead of supporting multiple different artists, you're siphoning your money to a few billion dollar companies (this is terrible for the economy)
As a side note, if you have a business which truly cannot afford to pay any artists, there are a lot of cheaper, (sometimes free!) pre-paid art bundles that are much less morally dubious than AI. Plus, then you're not siphoning all of your cash to tech oligarchs.
<joke>What's your rock band called, "SEC Form 10-K"?</joke>
People are saying, very clearly, that they're not willing to put effort into something produced by someone who put no effort in.
Your quip is pithy but meaningless.
I could have generated my own content, so just send the prompt rather than the output to save everyone time.
Again - your quip sounds good but when you think about it, it's flatly wrong.
I dont think gamers hate AI, it is just a vocal miniority imo. What most people dislike is sloppy work, as they should, but that can happen with or without AI. The industry has been using AI for textures, voices and more for over a decade.
It’s really not. That's actually a pet peeve of mine as someone who used to spent a lot of time messing with pixel art in Aseprite.
Nobody takes the time to understand that the style of pixel art is not the same thing as actual pixel art. So you end up with these high-definition, high-resolution images that people try to pass off as pixel art, but if you zoom in even a tiny bit, you see all this terrible fringing and fraying.
That happens because the palette is way outside the bounds of what pixel art should use, where proper pixel art is generally limited to maybe 8 to 32 colors, usually.
There are plenty of ways to post-process generative images to make them look more like real pixel art (square grid alignment, palette reduction, etc.), but it does require a bit more manual finesse [1], and unfortunately most people just can’t be bothered.
There is nothing that cannot harm. Knives, cars, alcohol, drugs. A society needs to balance risks and benefits. Word can be used to do harm, email, anything - it depends on intention and its type.
For icons in particular, this opens up a completely new way of customizing my home screen and shortcuts.
Not necessary for the survival of society, maybe, but I enjoy this new capability.
A mid-tier top-500 system (think about #250-#325) consumes about a 0.75MW of energy. AI data centers consume magnitudes more. To cool that behemoth you need to pump tons of water per minute in the inner loop.
Outer loop might be slower, but it's a lot of heated water at the end of the day.
To prevent water wastage, you can go closed loop (for both inner and outer loops), but you can't escape the heat you generate and pump to the atmosphere.
So, the environmental cost is overblown, as in Chernobyl or fallout from a nuclear bomb is overblown.
So, it's not.
As a country, we use 322 billion gallons of water per day. A few million gallons for a datacenter is nothing.
The water gets contaminated and heated, making it unsuitable for organisms to live in, or to be processed and used again.
In short, when you pump back that water to the river, you're both poisoning and cooking the river at the same time, destroying the ecosystem at the same time too.
Talk about multi-threaded destruction.
Pipes rust, you can't stop that. That rust seeps to the water. That's inevitable. Moreover, if moss or other stuff starts to take over your pipes, you may need to inject chemicals to your outer loop to clean them.
Inner loops already use biocides and other chemicals to keep them clean.
Look how nuclear power plants fight with organism contamination in their outer cooling loops where they circulate lake/river water.
Same thing.
The cost to humans living in affected areas was massive and high profile, but it’s very questionable if it was higher than that of an equivalent amount of coal-burning plants. Fortunately not a tradeoff we have to debate anymore, since there are renewables with much fewer downsides and externalities still.
Nuclear bombs (at least those being actually used) by design kill people, so I’m not sure what the externalities even are if the main utility is already to intentionally cause harm.
I'm not really well versed on the environmental cost, more just (neutrally) pointing out that comparing a single 10s image to a 5-6 hour commission ignores the fact that the majority of these images probably would never have existed in the first place without AI.
A modern laptop is running almost fanless, like a 486 from the days of yore.
A single H200 pumps out 700W continuously in a data center, and you run thousands of them.
Also, don't forget the training and fine tuning runs required for the models.
Mass transportation / global logistics can be very efficient and cheap.
Before the pandemic, it was cheaper to import fresh tomatoes from half-world away rather than growing them locally in some cases. A single container of painting supplies is nothing in the grand scheme of things, esp. when compared with what data centers are consuming and emitting.
so if power were plentiful and environmental you'd be onboard with it?
Please see my other comment about energy consumption and connect the dots with how open loop DLC systems are harmful to fresh water supplies (which is another comment of mine).
> so if power were plentiful and environmental you'd be onboard with it?
This is a pretty loaded way to ask this. Let me put this straight. I'm not against AI. I'm against how this thing is built. Namely:
- Use of copyrighted and copylefted materials to train models and hiding under "fair use" to exploit people.
- Moreover, belittling of people who create things with their blood sweat and tears and poorly imitating their art just for kicks or quick bucks.
- Playing fast and loose with environment and energy consumption without trying to make things efficiently and sustainably to reduce initial costs and time to market.
- Gaslighting the users and general community about how these things are built, and how it's a theater, again to make people use this and offload their thinking, atrophying their skills and making them dependent on these.
I work in HPC. I support AI workloads and projects, but the projects we tackle have real benefits, like ecosystem monitoring, long term climate science, water level warning and prediction systems, etc. which have real tangible benefits for the future of the humanity. Moreover, there are other projects trying to minimize environmental impact of computation which we're part of.So it's pretty nuanced, and the AI iceberg goes well below OpenAI/Anthropic/Mistral trio.
As opposed to the illusory/fake/immoral benefits of using LLMs for entertainment purposes (leaving aside all other applications for now)?
How do you feel about Hollywood, or even your local theater production? I bet the environmental unit economics don't look great on those either, yet I wouldn't be so quick to pass moral judgement.
Why not just focus on the environmental impact instead of moralizing about the utility? It seems hard to impossible to get consensus there, and the impact should be able to speak for itself if it's concerning.
Many people think that when a piece of hardware is idle, its power consumption becomes irrelevant, and that's true for home appliances and personal computers.
However, the picture is pretty different for datacenter hardware.
Looking now, an idle V100 (I don't have an idle H200 at hand) uses 40 watts, at minimum. That's more than TDP of many, modern consumer laptops and systems. A MacBook Air uses 35W power supply to charge itself, and it charges pretty quickly even if it's under relatively high stress.
I want to clarify some more things. A modern GPU server houses 4-8 high end GPUs. This means 3KW to 5KW of maximum energy consumption per server. A single rack goes well around 75KW-100KW, and you house hundreds of these racks. So, we're talking about megawatts of energy consumption. CERN's main power line on the Swiss side had a capacity around 10MW, to put things in perspective.
Let's assume an H200 uses 60W energy when it's idle. This means ~500W of wasted energy per server for sitting around. If a complete rack is idle, it's 10KW. So you're wasting energy consumption of 3-5 houses just by sitting and doing nothing.
This computation only thinks about the GPU. Server hardware also adds around 40% to these numbers. Go figure. This is wasting a lot for cat pictures.
And, these "small" numbers add up to a lot.
A: GPUs use a lot of power!
B: Not all of them are running 100% continuously, eh?,
A: They waste too much power when they're idle, too!
C: None of the H200s are sitting idle, you knob!
I mean, they are either wasting energy sitting idle or doing barely useful work. I don't know what to say anymore.We'll cook ourselves, anyway. Why bother? Enjoy the sauna. ¯\_(ツ)_/¯
> they are either wasting energy sitting idle or doing barely useful work
Now here's a true (inverse) scotsman, or more accurately, a moved goalpost: Work on things you don't deem valuable is basically the same thing as idling?
> We'll cook ourselves, anyway. Why bother? Enjoy the sauna. ¯\_(ツ)_/¯
I'm very concerned about that too, but I don't think we'll avoid the sauna with fatalism or logically unsound appeals to morality about resource consumption.
No, in terms of unit economics, I'm almost certain that the painting supplies have a bigger ecological/resource footprint than an LLM per icon generated, and I'm pretty sure the cost of shipping tomatoes does not decrease that footprint, even if it possibly dwarfs it.
But yes, due to Jevon's paradox, the total resource use might well increase despite all that. I, for example, would have never commissioned a professional icon for my silly little iOS shortcuts on my homescreen, so my silly icon related carbon footprint went from exactly zero to slightly above that.
If you see no difference between them, I can't continue to discuss this with you, sorry.
And I say that as somebody that also finds Ghibli knock-off avatars used by AI bros in incredibly bad taste (or, arguably an even worse crime against taste, a dated 2025 vibe).
I like your discussion style.
I don't want to live in a world in which people get to decide what others can and can't do with their share of resources (after properly accounting for all externalities, including pollution, the potential future value of non-renewable present resources etc. – this is where today's reality often and massively misses that ideal) based on their subjective moral criteria.
Not even just for ethical/moral reasons, but also for practical ones: It’s infinitely harder to get everybody to additionally agree on value of use than on fairness of allocation alone.
After thoroughly mixing these two quite distinct concerns, you'll also have a very hard time convincing me that your concerns for river pollution etc. (which I take very seriously as potentially unaccounted negative externalities, if they exist) are completely free from motivated reasoning about "immoral usage".
What a rotten exchange.
AI can probably fool most court judges now. Or the defense can refute legitimate evidence by saying “it’s AI / false”. How would that be refuted?
You might generate an AI video of me committing a crime, But the CCTV on the street didn't show it happening and my phone cell tower logs show I was at home. For the legal system I don't think this is going to be the biggest problem. It's going to be social media that is hit hardest when a fake video can go viral far faster than fact checking can keep up.
Given the obvious personal safety upsell ("our phone/dashcam/... produces court-admissible evidence!"), I think we'll even see this in consumer devices before too long.
For the nth time: scale, easiness, and access, matter. AI puts propaganda abilities far beyond the reach of those men in the hands of many more people. Do you not understand the difference between one man with a revolver and an army with machine guns? They are not the same.
Nowhere in my comment am I “blaming the tools”. I’ll ask you engage with the argument honestly instead of simply parroting what you already believe absent reading.
> I’ll ask you engage with the argument honestly instead of simply parroting what you already believe absent reading
I did engage with argument. The argument is a tiresome old argument that is knee-jerk anti tech. You seem to be the thoughtless one in this discourse repeating for the infinite time an anti-tech position assuming net negatives outweigh massively net +ves.
Also, why attack me instead of the argument? Did I touch a logical sore point? I believe so.
> For the nth time: scale, easiness, and access, matter.
By that logic, So the printing press was evil? Remember, Mao/Stalin/Hitler used presses to spread their propaganda.
Also, for the n+1 time, using your own style, don't be lazy:
1. Come up with a net benefit calculation for AI. What? You can't? Then, don't try to claim this is all net negative.
2. Explain how AI is different from other tech like the printing press, that also had scale, easiness, and access.
So that makes AI a "dual good", like a kitchen knife: you can cut your tomato or kill you neighbor with it, entirely up to the "user". Not all users are good, so we'll see an intense amplification of both good and bad.
They're adrift, every new "fact" (whether true or false) blows them in a new direction. Often they get led in terrible directions from statements that are entirely true (but missing important context).
A lot of financial cons work that way, a long string of true statements that seem to lead to a particular conclusion. I know that if someone is offering me 20% APY there will usually be some risk or fee that offsets those market-beating gains (it may be a worthwhile risk or a well earned fee, but that number needs to trigger further investigation).
We need people to be equipped with that sort of framework in as many areas as possible, but we seem to be moving backwards in that area.
I put in one of the driest descriptions of the Holocaust I could find and it got a very high score for bias, calling a factual description of a massacre emotional sensationalism because it inevitably contains a lot of loaded words.
It also doesn't differentiate between reporting, commentary, poetry, or anything else. It takes text and spits out a number, which is a very shallow analysis.
I started being totally indifferent after thinking about my spending habits to check for unnecessary stuff after watching world championships for niche sports. For some this is a calling for others waste. It is a numbers game then.
Agreed mostly, BUT
I'm building tools for myself. The end goal isn't the intermediate tool, they're enabling other things. I have a suspicion that I could sell the tools, I don't particularly want to. There's a gap between "does everything I want it to" and "polished enough to justify sale", and that gap doesn't excite me.
They're definitely not generated without effort... but they are generated with 1% of the human effort they would require.
I feel very much empowered by AI to do the things I've always wanted to do. (when I mention this there's always someone who comes out effectively calling me delusional for being satisfied with something built with LLMs)
That's it. I can't think of a single actual use case outside of this that isn't deliberately manipulative and harmful.
I just recently used for image generation to design my balcony.
It was a great way to see design ideas imagined in place and decide what to do.
There are many cases people would hire an artist to illustrate an idea or early prototype. AI generated images make that something you can do by yourself or 10x faster than a few years ago.
Not withstanding a few code violations, it generated some good ideas we were then able to tweak. The main thing was we had no idea of what we wanted to do, but seeing a lot of possibilities overlaid over the existing non-garden got us going. We were then able to extend the theme to other parts of the yard.
I dunno how long this is going to hold up. In 50 years, when OpenAI has long become a memory, post-bubble burst, and a half-century of bitrot has claimed much of what was generated in this era, how valuable do you think an AI image file from 2023 - with provenance - might be, as an emblem and artifact of our current cultural moment, of those first few years when a human could tell a computer, "Hey, make this," and it did? And many of the early tools are gone; you can't use them anymore.
Consider: there will never be another DallE-2 image generation. Ever.
As for advertising being depressing - its a little late to get up on the high horse of anti-Ads for tech after 2 decades of ad based technology dominating everything. Go outside, see all those bright shiny glittery lights, those aren't society created images to embolden the spirit and dazzle the senses, those are ads.
North Korea looks weird and depressing because the don't have ads. Welcome to the west.
But so many people want to make art, and it's so cheap to distribute it, that art is already commoditized. If people prefer human-created art, satisfying that preference is practically free.
But the idea of novelty is a misnomer I think. Any random number generator can arbitrarily create a "novel" output that a human has never seen before. The issue is whether something is both novel and useful, which is hard for even humans to do consistently.
I’m so tired of “there’s nothing preventing”, and “humans do that too”. Modern AI is just not there. It’s not like humans and has difficulties with adapting to novelty.
Whether transformers can overcome that remains to be seen, but it is not a guarantee. We’ve been dealing with these same issues for decades and AI still struggles with them.
There is a mass, bland appeal to “better” things but it’s not ubiquitously desired and there will always be people looking outside of that purely because “better” is entirely subjective and means nothing at all.
Edit: One of the possible outcomes may be living in a world like in "Them" with glasses on. Since no expression has any meaning anymore, the message is just there being a signal of some kind. (Generic "BUY" + associated brand name in small print, etc.)
I'm not sure you immediately lose meaning if someone can make a highly personalized version of something easily. The % of completely meaningless video after YouTube and tiktok came about has skyrocketed. The amount of good stuff to watch has gone up as well though.
Is an AI generated photo of your app/site going to be more accurate than a screenshot? Or is an AI generated image of your product going to convey the quality of it more than a photo would?
I think Sora also showed that the novelty of generating just "content" is pretty fleeting.
I would be interested to see if any of the next round of ChatGPT advertisements use AI generated images. Because if not, they don’t even believe in their own product.
What? Those items are luxuries when made by humans because they are physical goods where every single item comes with a production and distribution cost.
I used to have an assistant make little index-card sized agendas for gettogethers when folks were in town or I was organising a holiday or offsite. They used to be physical; now it's a cute thing I can text around so everyone knows when they should be up by (and by when, if they've slept in, they can go back to bed). AI has been good at making these. They don't need to be works of art, just cute and silly and maybe embedded with an inside joke.
If I got one of your cute schedule cards while visiting you, I'd tear it up, check into a cheap motel, and spend the rest of my vacation actually enjoying myself.
Edit: I'm not an outlier here. There have even been sitcom episodes about overbearing hosts over-programming their guests' visits, going back at least to the Brady Bunch.
Okay. I'd be confused why you didn't voice up while we were planning everything as a group, but those people absolutely exist. (Unless it's someone's, read: a best friend or my partner's, birthday. Then I'm a dictator and nobody gets a choice over or preview of anything.)
I like to have a group activity planned on most days. If we're going to drive to get in an afternoon hike in before a dinner reservation (and if I have 6+ people in town, I need a dinner reservation because no I'm not coooking every single evening), or if I've paid for a snowmobile tour or a friend is bringing out their telescope for stargazing, there are hard no-later-than departure times to either not miss the activity or be respectful of others' time.
My family used to resolve that by constantly reminding everyone the day before and morning of, followed by constantly shouting at each other in the hours and minutes preceding and–inevitably–through that deadline. I prefer the way I've found. If someone wants to fuck off from an activity, myself included, that's also perfectly fine.
(I also grew up in a family that overplanned vacations. And I've since recovered from the rebound instinct, which involves not planning anything and leaving everything to serendipity. It works gorgeously, sometimes. But a lot of other times I wonder why I didn't bother googling the cool festival one town over before hand, or regretted sleeping in through a parade.)
> There have even been sitcom episodes about overbearing hosts over-programming their guests' visits
Sure. And different groups have different strokes. When it comes to my friends and I, generally speaking, a scheduled activity every other day with dinners planned in advance (they all get hangry, every single fucking one of them) works best.
It's good that my friends don't make a coffee date feel like a board meeting (with an agenda shared by post 14 working days ahead of the meeting, form for proxy voting attached).
If this is the best use case that exists for AI image generation, I'm only further convinced the tech is at best largely useless.
Because I’ll then spend hours playing with the typography (because it’s fun) and making it look like whatever design style I’ve most recently read about (again, because it’s fun) and then fighting Word or Latex because I don’t actually know what I’m doing (less fun). Outsourcing it is the right move, particularly if someone else is handling requests for schedules to be adjusted. An AI handles that outsourcing quicker for low-value (but frequent) tasks.
> If this is the best use case that exists for AI image generation
I’ve also had good luck sketching a map or diagram and then having the AI turn it into something that looks clean.
Look, 99% of my use cases are e.g. making my cat gnaw on the Tetons or making a concert of lobsters watching Lady Gaga singing “I do it for the claws” or whatever so I can send two friends something stupid at 1AM. But there does appear to be a veneer of productivity there, and worst case it makes the world look a bit nicer.
I get this sounds elitist - but tremendous percentage of population is happily and eagerly engaging with fake religious images, funny AI videos, horrible AI memes, etc. Trying to mention that this video of puppy is completely AI generated results in vicious defense and mansplaining of why this video is totally real (I love it when video has e.g. Sora watermarks... This does not stop the defenders).
I agree with you that human connection and artist intent is what I'm looking for in art, music, video games, etc... But gawd, lowest common denominator is and always has been SO much lower than we want to admit to ourselves.
Very few people want thoughtful analysis that contradicts their world view, very few people care about privacy or rights or future or using the right tool, very few people are interested in moral frameworks or ethical philosophy, and very few people care about real and verifiable human connection in their "content" :-/
The unsettling thing on social media is the mind hijacking with the recommendation algo and scrolling motion that resembles a slot machine, more than the content itself.
It's been true for various technologies that HN (and tech audiences in general) have a more nuanced view, but AI flips the script on that entirely. It's the tech world who are amazed by this, producing and being delighted by endless blogposts and 7-second concept trailers.
I think HN probably uses GenAI more than average population.
But I think HN consumes less GenAI content than average population.
Look at Facebook, Instagram, Youtube, TikTok, etc. All I see is my non-techie friends being amazed and mesmerized by - cute animals, creepy animals, political events, jokes, comedy, outrage, events, speeches - that never ever happened. As if we don't have actual real puppies that are cute, my acquintenances and family are oooing and awwwing at fake howling huskies, fake animals being jump-scared by fake surprises.
HN may be amazed by potential of AI output the improve the world more than average person. But hustlers are laughing their way to the bank as they actually use AI to make ridiculous, and I do mean ridiculous, amount of "content" for cheap, that is, absolutely is, being consumed at prodigious rate with no sign of stopping. This is not 7-second trailers and concepts for some future years - this is mega-years of actual content being liked, shared, engaged with and consumed, right now. This is what OP is hoping that tides will turn against, and this is what I see no sign of rejection in my non-techie/non-geeky circles :(
Sure, the weird cat-people adverts aren't aimed at HN's commentariat, but every 'democratise art and build that game you've dreamt of' pitch is. Every breathless paean to AI assistants/companions/partners is targeted at the users here.
Usage is a form of consumption; thinking of yourself as a creator while you consume doesn't mean you consume less.
Non-tech users are being fed fake images when they browse idly. Tech users are restructuring their entire lives around these tools.
Also, this can’t be real. How many publications did they train this stuff on and why are there no acknowledgment even if to say - we partnered with xyz manga house to make our model smarter at manga? Like what’s wrong with this company?
If a work of art is good, then it's good. It doesn't matter if it came from a human, a neanderthal, AI, or monkeys randomly typing.
When I watch a Lynch film I feel some connection to the man David Lynch. When I see a AI artwork, there is nothing to connect with, no emotional experience is being communicated, it is just empty. It's highest aspiration is elevator music, just being something vaguely stimulating in the background.
I understand these are fundamental questions about aesthetics that people differ over. But that's how it works for me. However, ultimately, I think people will realize that I'm right around the time that AI does start generating good art.
Is that true? Don't think I'd get tired of images that are as good as human made ones just because I know/suspect there may have been AI involved
I‘m sure if they could they would have shown more all americans. Especially given how important the state connection is for them to keep up their spending..
That means they struggle to find american technical presenters
response: https://chatgpt.com/backend-api/estuary/content?id=file_0000...
result: FAIL
Maybe it's meant to convey pace & hype
But the broader concept of fake news and the manufactured nature of media and rhetoric is much more relevant - e.g. whether or not something's AI is almost immaterial to the fact that any filmed segment does not have to be real or attributed to the correct context.
Its an old internet classic just to grab an image and put a different caption on it, relying on the fact no one can discern context or has time to fact check.
AI generated voice over, likely AI generated script (You see, this model isn't just generating images, it's thinking!). From what it looks like only the editing has some human touch to it?
It does this Apple style announcement which everyone is doing, but through the use of AI, at least for me, it falls right into the uncanny valley.
I would imagine this will hit illustrators / graphics designers / similar people very hard, now that anyone can just generate professional looking graphical content for pennies on the dollar.
but in general though - will people believe in anything photographic ?
imagine dating apps, photographic evidence.
I'm guessing we're gonna reach a point where - you fuck up things purposely to leave a human mark.
Hopefully film makes a come back.
Storefronts like Steam require disclosing use of AI assets for art. In most indie dev spaces, devs are scolded for using AI art in their games. I wonder if this perspective will change in a few years.
It's just another step into hell.
Never before in history did humanity have the possibility of seeing a picture of a pack of wolves! The dearth of photographs has finally been addressed!
I told my AI girlfriend that I will save money to have access to this new technology. She suggested a circular scheme where OpenAI will pay me $10,000 per year to have access to this rare resource of 21th century daguerreotype.
The person you're replying to is making a joke about OpenAI shutting down Sora their video generation "social media" app recently.
https://www.gally.net/temp/20260422-chatgpt-images-2-example...
Later Google tried the same thing, Apple we will give you a $1 billion dollar a year refund, what’s changed in two and a half years?