When in reality if you ask ChatGPT for 10 good movies from this year you will get this.
Anora - Directed by Sean Baker, a compelling drama about the life of a sex worker in Coney Island.
Challengers - A provocative tennis drama directed by Luca Guadagnino, starring Zendaya.
Dune: Part Two - Denis Villeneuve's continuation of the epic science fiction saga.
Furiosa: A Mad Max Saga - An action-packed prequel exploring the origins of Furiosa, directed by George Miller.
Inside Out 2 - Pixar's sequel that dives deeper into the complexities of human emotions.
Wicked - A musical fantasy adaptation directed by Jon M. Chu . The Zone of Interest - A thought-provoking film about Auschwitz, directed by Jonathan Glazer.
The Idea of You - A steamy romance starring Anne Hathaway.
Hit Man - A comedy thriller starring Glen Powell.
The Outrun - A powerful drama about a recovering alcoholic, starring Saoirse Ronan.
Let me know if you'd like more details about any of these!
Which is a great list.
I know that any mention of fallacies, valid or otherwise, causes instinctive eye rolls, but in this instance I agree with them that this amounts to moving the goalposts.
Originally the problem was supposedly that it would hallucinate complete and utter gibberish, but now here we are quibbling over one example and insisting that maybe it's not quite as good as alternative descriptions.
The gap between what was produced and what you're looking for is small enough that I think it could be covered with some slightly tweaked prompt instructions.
I'm not saying you're wrong but want to note how the goalposts keep seeming to shift whenever we talk about these capabilities.
That point is that the information provided above about these movies is worthless. It does not add any new value beyond what would already be available in the streaming interface. Several of the descriptions are nothing but the genre and one person involved in the making of the movie. And yet even with these descriptions being incredibly short and vague, they still manage to contain at least one misleading summary.
Despite your protestations to the contrary, these descriptions seem perfectly fine in that they're accurate and meaningful. And it if you want to start getting fast and loose with all kinds of new extra criteria and requirements for what it's supposed to do, they all seem squarely within the reach of the capabilities on offer, with some prompt tweaks.
The description of Wicked doesn't mention either The Wizard of Oz or the Broadway musical. So yes, the descriptions don't contain obscene mistakes like calling Wicked a courtroom drama. If that is enough for you to call these "accurate" while ignoring the vagueness or the 1 in 10 failure rate on the Anora description, fine by me. But you must have some weird definition of the word "meaningful" to apply that to descriptions like the one of Wicked. That simply isn't a helpful way to describe that movie.
> I would expect nothing but hallucinations and nonsense coming out of any LLM regarding recently-released movies (aka. the ones you often find on flights).
The comment that replied to it (the one that you're arguing against) provides evidence that proves it wrong. You are correcting someone who isn't incorrect, and I think the person you're responding to is very justified in saying you're moving the goalposts here.
> Those descriptions are less detailed than the information you will see on basically any streaming interface and yet it still manages to not being very good.
The points you made were not relevant to the discussion at hand. It's like if people were having a debate about where to find the best tacos in town and you stepped in to say "tacos aren't as good as hamburgers, you know" and then got upset that nobody wanted to debate that point with you. It's not everybody else's fault if you don't understand how conversations work!
It was literally the first movie in that list.
You tried making a counter example and the first part of it was already wrong.
That’s the point. Not that it _cant_ give good answers, but whether it does or not is a crap shoot.
Now to analyze how correct it was we need to verify each movie it gave… It’d be faster just to read the movie descriptions.
And I think that's quite obviously not the case, most, probably every other example on the list is just fine.
I now need to either trust a machine that I know gives incorrect information (as demonstrated by the first example) or I need to verify each example.
> probably every other example on the list is just fine.
Why don’t you check IMDb and let me know?
While you’re at it, don’t think about how much faster it would’ve been if you just looked up popular recent movies on IMDb or rotten tomatoes.
"Anora is a 2024 American comedy-drama film written, directed, and edited by Sean Baker. It follows the beleaguered marriage between Anora (Mikey Madison), a young sex worker, and Vanya Zakharov (Mark Eydelshteyn), the son of a Russian oligarch. The supporting cast includes Yura Borisov, Karren Karagulian, Vache Tovmasyan, and Aleksei Serebryakov."
I haven't seen the film, but it doesn't seem incompatible with ChatGPT's briefer description.
Instead of just checking with a first party source, you ask a statistical guessing machine for an answer.
There was a disagreement about the answer, so we needed to dig deeper.
You bring up Wikipedia, a 3rd party source of information. That description could also be wrong (it’s probably not, but stick with me)
Instead of just checking with a first party source (IMDb is very easy to search on), we went through several layers of obfuscation.
This was an issue for Wikipedia early on, but it has citations, at least. AI doesn’t and doesn’t have an army of people constantly fact checking every answer generated either.
There’s no benefit to asking AI for information like this. Especially since the in flight summary has accurate information that’s more than “drama, sex worker, cony island”
Maybe something like perplexity is better, since it has citations, but I haven’t tried it for very long yet.
In general you can't, but surely it's not that big a deal if ChatGPT offers an inaccurate summary of a movie you're about to use to kill time on a flight? I suppose it becomes important if, e.g., you're relying on it to tell you whether a movie is appropriate for children, but, if you're just asking it whether a movie is worth watching, that's a question that doesn't have an objective, factual answer anyway, so a hallucinated answer is probably about as useful as that of a not-previously-known reviewer.
Sure, but that's the filmmaker's interest. As someone sitting on a plane trying to decide whether to watch a movie, I care about my interest, not that of the person who made it. I'm not particularly arguing for the use of ChatGPT here (I wouldn't use it), just that the risks it usually poses are fairly minimal in this case.
You don't even seem to be disputing the actual results here, just gesturing towards a kind of philosophy class exercise of whether we can ever "really" verify its accuracy. I see Wittgenstein's name increasingly tossed around in these parts (a good thing!), so I'll just note that one of the reasons he's hailed as one of the great philosophers of the 20th century is because he felt these puzzles about "really" knowing were frivolous.
I don't think I agree that what's needed here is some new and extra process of verification. I think the same usual quality control criteria that are already being used are good enough in this case.
My wife wanted a pair of boots for Christmas that I couldn’t find in her size. Google was a wasteland of SEO, but ChatGPT found 5 sites and was able to tell me current stock levels.
' As of my knowledge cutoff in January 2022, the last movie I have information on is "Spider-Man: No Way Home", which was released in theaters in December 2021. It was one of the most highly anticipated films of that year, marking a major event in the Marvel Cinematic Universe (MCU) and the Spider-Man franchise. '
I pasted the same initial prompts in both, but Meta AI needed more clarification. When ChatGPT found multiple entries with similar titles, it gave information about all of them.
https://gist.github.com/appsforartists/004bafe11a9e23a418fd5...
The first thing I fact-checked, the Rotten Tomatoes scores are actually 66% and 51% respectively[1]. Probably not enough of a difference to sway any opinions, but an excellent example of the type of inaccuracy that the previous comment was referencing.
Hilariously it often believes that it can’t access the web and then hallucinates reasons for how it can know things beyond its knowledge cutoff date. But in any case, it works very well for this use case.