I see this all the time from AI boosters. Flashy presentation, and it seems like it worked! But if you actually stare at the result for a moment, it’s mediocre at best.
Part of the issue is that people who are experts at creating ML models aren’t experts at all the downstream tasks those models are asked to do. So if you ask it to “write a poem about pizza” as long as it generally fits the description it goes into the demo.
We saw this with Gemini’s hallucination bug in one of their demos, telling you to remove film from a camera (this would ruin the photos on the film). They obviously didn’t know anything about the subject beforehand.