OpenAI’s Sora made me crazy AI videos then the CTO answered most of my questions
youtube.com
youtube.com
This thread is a good example of how damaging that can be because threads are so sensitive to initial conditions. The comments are basically all responding to the editorialized title, and most are just angry reflexive responses. It's not possible to salvage a thread once it's gotten going like this. Maybe there are interesting things in the video, maybe not, but if there are, it's too late for them to get discussed here.
Being sued in general is not that great of an indicator for anything, in a world filled with lawyers and angry people.
For these companies it’s especially important since even if copyright claims will come down the line they need to become big enough to be able to push back on them so they could force a favorable settlement.
[1] https://www.nnlm.gov/guides/data-glossary/data-provenance
- They say what they know
- They don't say what you don't know
- Let them know only what you want them to say
Because, you see it is absolute absolutely mind boggling to understand how a 16 y.o. Ermira girl from the third-largest city (very small actually) in a country where mafia and government was melted together at that time (and still pretty much is), so... how this girl won a scholarship of unspecified origin (https://en.wikipedia.org/wiki/Mira_Murati) which took her straight to USA, and later on landed her in a private and very expensive Ivy League university. You see, there are at least dozen people in my extended pool of contacts from that time, from this region, who have been at IoI, or IoM at the time, and won prestigious first places, and won scholarships of some sorts, and none was THAT lucky.
Of course, this may sound like a girls dream come true, but if you have even a limited insight how the Balkans operate, and particularly how Albania operated 20 years ago... And together with the fact that the present Albanian prime minister suddenly is very close to nowadays Mira, so much as to embraced OpenAI for legislation-something (https://www.euractiv.com/section/politics/news/albania-to-sp...).
Sorry, perhaps my imagination, but this really raises a brow.
From Wikipedia (https://en.wikipedia.org/wiki/Mira_Murati) - Throughout her school years, she participated in many Olympiads and math competitions. That was likely how she got a foreign scholarship.
But once again - the timeframe is very very important.
I understand that we humans have natural instincts to uncover plots because as a social animal we have been primed to develop such a skill.
We have also been primed to recognize faces but that can lead us to see them even when there are no faces (e.g. the sphinx on Mars).
We're very bad at intuitively grasping low probability events and large numbers.
I personally know lots of people who have played the lottery but I never met somebody who was THAT lucky to win a jackpot.
Yet those people exist. We understand how the lottery works. It happens regularly enough and transparently enough so it no longer tickles our "corruption/plot/conspiracy" instincts. But if lotteries were never invented and we had one run today and somebody won, I'm pretty sure the default assumption for most people would be to be suspicious about who that person was, why they won, was it a setup etc etc
Unless you have a specific allegation please refrain from insinuating wrong doing just because she came from a corrupted country and was successful.
Edit: Mind you, she did say that they used publicly available and licensed data. So if you're saying she has no motivation to say this, she already said it.
1. She doesn't know and appears incompetent.
2. She does know and is lying (as you see, this is the main belief).
3. She does know in part and that part could include an honest use of this claim. (She doesn't know all, but what she does know is legal)
4. She does know in part and knows that some was illegally obtained
Truth be told, unless it is in writing somewhere that she knows of illegal data being used, she could post hoc claim 3. But if she does know illegal, then option 1 still gets her in trouble for the same reasons 3 would. The only difference is she appears incompetent by saying I don't know. Remember, post hoc she can say "At the time I was not aware of any illegally obtained works being used to train SORA". Claiming she doesn't know now would be WORSE if that is the situation and it came to light because it is a more explicit form of deciept.Edit: Mind you, she basically took 1 and 3. She did say it used publicly available and licensed data.
5. "you won't understand" or "you're not yet ready to comprehend" or "insufficient data for meaningful answer".
When journalist spends half of the time speaking about safety and identification, the interviewee certainly realizes that the journalist and his audience do not understand the basic principles of algorithms. - Researcher: We have made an ultra-efficient algorithm for sorting data!
- Journalist: And what protections have you put in place against processing copyrighted results? What about gender neutrality? And how should we distinguish between data sorted by a conventional algorithm and yours?
- Researcher: ...You could see her mind click into a new gear when that question was asked.
Just feels like more shadiness from this company.
They 100% use full-length, copyrighted movies at OpenAI. Try using Whisper and see how often you get an obvious hint like "Subtitles by <RandomInternetUser1234>"
Similarly no answer about videos from Facebook or Instagram.
It's a kind of significant implementation detail, wouldn't you say? If the discussion were about minor sources of videos it would be a lot more understandable.
And pretty much everyone knows that this question is going to be asked during any interview. If there's anything to prepare for, it is this question.
The specific questions were
What data was used to train SORA?
So, videos on YouTube?
Videos from Facebook? Instagram?
What about Shutterstock? I know you guys have a deal with them.