Sora: Creating video from text
openai.com
openai.com
Also (since it's been a while): there are over 2000 comments in the current thread. To read them all, you need to click More links at the bottom of the page, or like this:
https://news.ycombinator.com/item?id=39386156&p=2
Motion-capture works fine because that's real motion, but every time people try to animate humans and animals, even in big-budget CGI movies, it's always ultimately obviously fake. There are so many subtle things that happen in terms of acceleration and deceleration of all of the different parts of an organism, that no animator ever gets it 100% right. No animation algorithm gets it to a point where it's believable, just where it's "less bad".
But these videos seem to be getting it entirely believable for both people and animals. Which is wild.
And then of course, not to mention that these are entirely believable 3D spaces, with seemingly full object permanence. As opposed to other efforts I've seen which are basically briefly animating a 2D scene to make it seem vaguely 3D.
Except in games where they mo-cap at a frame rate less than what it will be rendered at and just interpolate between mo-cap samples, which makes snappy movements turn into smooth movements and motions end up in the uncanny valley.
It's especially noticeable when a character is talking and makes a "P" sound. In a "P", your lips basically "pop" open. But if the motion is smoothed out, it gives the lips the look of making an "mm" sound. The lips of someone saying "post" looks like "most".
At 30 fps, it's unnoticeable. At 144 fps, it's jarring once you see it and can't unsee it.
How do we get from there to "just assume every company in the world will sell your data in wildly and obviously illegal ways", I don't know.
The most simple version would be an ad-supported ChatGPT experience. Anyone thinking that an internet consumer company with 100m weekly active users (I‘m citing from their job ad) is not going to sell ads is lacking imagination.
And that should probably take precedence over the semantics of your moniker, every single time (even if hn continues to be super sour about it)
The more powerful, the more important it is that everyone has access.
That "the more powerful, the more important it is that everyone has access"?
Especially considering that the biggest killer app for AI could very well be smart weapons like we've never seen before.
Nukes aren’t even close to being commodities, cannot be targeted at a class of people (or a single person), and have a minutely small number of users. (Don’t argue semantics with “class of people” when you know what I mean, btw)
On the other hand, tech like this can easily become as common as photoshop, can cause harm to a class of people, and be deployed on a whim by an untrained army of malevolent individuals or groups.
It's important to weigh the benefits of diversity and open competition against the risks of bad actors misusing the tools. Ultimately, finding a balance between accessibility and responsible use is key.
What guarantee do we have that OpenAI won't become an evil actor like Skynet?
The direct threat to society is actually this kind of secrecy.
If ordinary people don't have access to the technology they don't really know what it can do, so they can't develop a good sense of what could now be fake that only a couple of years ago must have been real.
Imagine if image editing technology (Photoshop etc) had been restricted to nation states and large powerful corporations. The general public would be so easy to fool with mere photographs - and of course more openly nefarious groups would have found ways to use it anyway. Instead everybody now knows how easily we can edit an image and if we see a shot of Mr Trump apparently sharing a loving embrace with Mr Putin we can make the correct judgement regarding a probable origin.
I was replying to a comment saying that nukes aren't commodities and can't target specific classes of people, and I don't understand why those properties in particular mean access to nukes should be kept secret and controlled.
What do you mean? Are you being dramatic or do you actually believe that the US government will/can not absolutely shut OpenAI down, if they feel it was required to guarantee state order?
...after amazing public world wide demos that show how real the AI generated videos can be? How long has Hollywood had similar "fictional videos" powers?
if it were opened to public faking such videos would lose (nearly) all of its power
How quickly do you think our gerontocracy will adapt to the new reality?
When a significant fraction of video is generated content spat out by a bored teenager on 4chan, then people will stop trusting it, and hence it will no longer have the power to convince people to kill.
Isreal media machinery parading photographs of damaged houses that could only be done by heavy artillery or tank shells blaming on rebels carrying infantry rifles.
But I agree, as if the current tools were not enough to sway people they will have more means to sway public opinion.
Not directly. But I won't be surprised if AI video generators aren't somewhere in the chain of causes of gigadeaths this century.
A homing missile that chases you across continents and shows you disturbing deepfakes of yourself until you lose your mind and ask it to kill you. At that point it switches to encourage mode, rebuilds your ego, and becomes your lifelong friend.
The ideas of critical mass, prompt fission, and uranium purification, along with the design of the simplest nuclear weapon possible has been out in the public domain for a long time.
As things are at the moment, while supression of a technology has benefits, it seems like a risky long-term solution. All it takes is for a single world-altering technology to slip through the cracks, and a bad actor could then forever change the world with it.
If I engineered the tech I would be much more fearful of the possibility of malice in the future leadership of the organization I'm under if they continue to keep it closed, than I would be fearful of the whole world getting the capability if they decide to open source.
I feel that, like with Yellow Journalism of the 1920s, much of the misinformation problem with generative AI will only be mitigated during widespread proliferation, wherein people become immune to new tactics and gain a new skepticism of the media. I've always thought it strange when news outlets discuss new deepfakes but refuse to show it, even with a watermark indicating it is fake. Misinformation research shows that people become more skeptical once they learn about the technological measures (e.g. buying karma-farmed Reddit accounts, or in the 1920s, taking advantage of dramatically lower newspaper printing costs to print sensationalism) through which misinformation is manufactured.
The studio told him the model was their property, and they wouldn't share it.
Peculiar reasoning, isn't it?
You need to watch more dystopian movies.
[1] https://storage.googleapis.com/deepmind-media/gemini/gemini_...
The most recent thing they released was Whisper, which to be fair is the only model with absolutely no safety implications.
Don't get me wrong, it is impressive. But I think many people will be very uncomfortable with such motion very quickly. Same story as the fingers before.
The people behind her all walk at the same pace and seem like floating. The moving reflections, on the other hand, are impressive make-believe.
I don't believe, movie makers are out of buisness any time soon. They will have to incorporate it though. So far this can make convincing background scenery.
It seems bizarre to think the gee whiz factor in a new commercial creative product makes critiquing its output out-of-bounds. This isn't a university research team: they're charging money for this. Most people have to determine if something is useful before they pay for it.
My son was learning how to play keyboard and he started practicing based on metronome. At some point, I was thinking, why is he learning it at all? We can program which key to be pressed at what point in time, and then a software can play itself! Why bother?
Then it hit me! Musicians could automate all the instruments with incredible accuracy since a long time. But they never do that. For some reason, they still want a person behind the piano / guitar / drums.
What do you judge was the ratio of automated music (recordings played back) to live music played in the last year?
But I take it, maybe I'm not so familiar with world music, I was talking more about Indian music. While the music is recorded and mixed across several tracks electronically, I think most of it is played (or sang) originally by a person.
In the US atleast there's the occasional acoustic song that becomes a hit, but rock music is obviously on its way to slowly becoming jazz status. It and country are really the last genres where live traditional instruments are common during live performances. Pop, Hip Hop, and EDM basically all are put together as being nearly computer perfect.
All the great producers can play instruments, and that's often times the best way to get a section out initially. But what you hear on Spotify is more and more meticulously put together note by note on a computer after the fact.
Live instruments on stage are now often for spectacle or worse a gimmick, and it's not the song people came to love. I think the future will have people like Lionclad[1] in it pushing what it means to perform live, but I expect them to become fewer and fewer as music just gets more complex to produce overall.
You aren't better because you prefer live music, you just have a preference. Music wasn't better some arbitrary number of years ago, you just have a preference.
Nobody said one form is objectively better, just that there is a form that is becoming more popular.
But to state my opinion, I can't imagine something more boring than thinking the best of music, performance, TV, or media in general was done best and created in the past.
> But to state my opinion, I can't imagine something more boring than thinking the best of music, performance, TV, or media in general was done best and created in the past.
And to state my opinion, art isn't about "the best" or any sort of progress, it's about the way we humans experience the world, something I consider to be a timeless preoccupation, which is why a song from 2024 can be equally touching as a ballad from the 14th century.
(That being said, a realtime AI-based bandmate could be interesting...)
I wonder if something is lost in the recording process that just cannot be replicated? A live instrument is something that you can actually feel the sound of IMO, I've never felt the same with recorded music even though I of course enjoy it.
I wonder if when we get older we just get kind of "bored" (sadly) and it doesn't mean as much to us as it probably should.
This actually happened on a recent hit, too -- Dua Lipa's Break My Heart. They originally had a drum machine, but then brought in Chad Smith to actually play the drums for it.
Edit: I'm not claiming this was new or unusual, just providing a recent example.
Also way before that back in the early 80’a Depeche Mode displayed the recorded drumb-reel onstage so everyone knew what it was, but when the got big enough they also transitioned into an epic live show with guitars and live drum a as well as synth-hooked drums devices they could bag on in addition to keyboards.
We are human. We want humans. Same reason I want a hipster barista to pour my coffee when a machine could do it just as well.
I've wondered about this for a long time too, why on earth is anyone still able to be a barista, it turns out, people actually like the community around cafes and often that means interacting with the staff on a personal level.
Some of my best friends have been barista's I've gone to over several years.
A recorded studio version, I can also listen to at home. But a full band performing in this very moment is a different experience to me.
Outside of that, people just want to know what it feels like to be able to play their favorite song on guitar and to go skiing etc.
Being perfect at everything would be honestly boring as shit.
This is true for literally every hobby people do for fun. I am learning ceramics. Everything I've ever made could be bought in a shop for a 100th of the cost, and would be 100 times "better". But I enjoy making the pot, and it's worth more to me than some factory item.
Sona will allow a new hobby, and lots will have fun with it. Pros will still need to fo Pro things. Not everything has to be viewed through the lens of money.
I play the piano, and even though MIDI exists, I still derive a lot of enjoyment from playing an acoustic instrument.
On the other hand you've probably heard of an iPod, which I think I could describe as a device dedicated to give false sense of an ever-present musician, so to speak.
So, "they" in "they still want a person behind the piano" is not just limited to hobbyists and enthusiasts. People wants people behind an instrument, for some reason. People pays for others' suffering, not for a thing's peculiarity.
Hilariously, nearly every electronic artist I can think of, stands in front of a crowd and "plays "live" by twisting dials etc, so I think it's fairly accurate.
Carl Cox, Tycho, Aphex Twin, Chemical Brothers, Underworld, to name a few.
I stand by my original point. There are plenty of people who really do not care if there is a human somewhere "performing" the music or not. And that's totally fine.
An ML model could probably do a good job at selecting tunes of a particular genre that fit into a pre-defined "journey" that the promoter is trying to construct, so I could see a role for "AI DJs" in the future, especially for low budget parties during unpopular timeslots like first day of a festival while people are still arriving and the crew is still setting up. Some of that is already done by just chucking a smart playlist on shuffle. But then you also have up-and-comer or hobbyist DJs who will play for free in those slots, so maybe there's not really a need for a smarter computer to take over the job.
This whole thread started from the question of why a human should do something when a machine can do it better. And the answer is simple: because humans like to do stuff. It is not because humans doing stuff adds some kind of hand-wavey X factor that other humans intrinsically prefer.
You've never been to a rave, huh? For that matter, there's a lot of pop artists that use sequencers and dispense with the traditional band on stage.
Ability to create live experiences can still be a motivating factor for musicians (aside from the love of learning). Yet, when AI does the song-writing far more effectively, then will the musician ignore this?
It's like Brave New World. Musicians who don't use these AI tools for song-writing will be like a tribe outside modern world. That's a tough future to prepare for. We won't know whether a song was actually the experience and emotions of a person or not.
Steam shovels and modern excavators didn't remove our need for shovels or more importantly, the know-how to properly apply these tools. Naturally, most people use a shovel before they operate an excavator.
Human ingenuity always finds a need for value creation. Greater abundance creates new opportunities.
Take the inverse position. Should we go back to reading by candlelight to increase employment in candle making?
No, electric lighting allowed peopled to become productive during night hours. A market was created for electricity producers, which allowed additional products which consume electricity to be marketed. Technological increases in productivity cascade into all areas of life, increasing our living standards.
A more interesting, if not controversial line of inquiry might start with: If technology is constantly advancing human productivity, why do modern economies consistently experience price inflation?
Of course, this change is dislocating for the particular people whose toil disappeared. They need support to retrain to new occupations.
The alternative is to cling to a past where everyone - on average - is poorer, less healthy, and works in more dangerous jobs.
Clearly if there are ways out of being displaced, please share them
There are subtle and deliberate deviations in timing and elements like vibrato when a human plays the same song on an instrument twice, which is partly why (aside from recording tech) people prefer live or human musicians.
Think about how precise and exacting a computer can be. It can play the same notes in a MIDI editor with exact timing, always playing note B after 18 seconds of playing note A. Human musicians can't always be that precise in timing, but we seem to prefer how human musicians sound with all of the variations they make. We seem to dislike the precise mechanical repetition of music playback on a computer comparatively.
I think the same point generalises into a general dislike on the part of humans of sensory repetition. We want variety. (Compare the first and second grass pictures at [0] and you will probably find that the second which has more "dirt" and variety looks better.) "Semantic satiation" seems to be a specific case of the same tendency.
I'm not saying that's something a computer can't achieve eventually but it's something that will need to be done before machines can replace musicians.
Also keep an eye on teeth and high contrast text. Anything small and prone to distortion in low resolution video and images used to train this stuff.
And by solved, I mean they'll create convincing clips that'll be hard for people to dismiss unless they're really looking closely. I think it's only a matter of time until fake video clips lead to real life outrage and violence. This tech is going to be militarized before we know it.
I showed these demos to my partner yesterday and she was upset about how real AI has become, how little we will be able to trust what we see in the future. Authoritative sources will be more valuable, but they themselves may struggle to publish only the facts and none of the fiction.
Here's one possible military / political use:
The commander of Russia's Black Sea Fleet, Viktor Sokolov, is widely believed to have been killed by a missile strike on 22 September 2023. https://en.wikipedia.org/wiki/Viktor_Sokolov_(naval_officer)
Russian authorities refute his death and have released proof of life footage, which may be doctored or taken before his death. Authoritative source Wikipedia is not much help in establishing truth here, because without proof of death they must default to toeing the official line.
I predict that in the coming months Sokolov (who just yesterday was removed from his post) will re-emerge in the video realm, and go on to have a glorious career. Resurrecting dead heroes is a perfect use of this tech, for states where feeding people lies is preferable to arming them with the truth.
Sokolov may even go on to be the next Russian President.
I think this way of thinking is distracted. No type of media has ever been a source of truth in itself. Videos have been edited convincingly for a long time, and people can lie about their context or cut them in a way that flips their meaning.
Text is the easiest media to lie on, you can freely just make stuff up as you go, yet we don't say "we cannot trust written text anymore".
Well yeah duh, you can trust no type of media just because it is formatted in a certain way. We arrive at the truth by using multiple sources and judging the sources' track records of the past. AI is not going to change how sourcing works. It might be easier to fool people who have no media literacy, but those people have always been a problem for society.
You are right but the thing with this is the speed and ease with which you can generate something completely fake.
> Well yeah duh, you can trust no type of media just because it is formatted in a certain way
Maybe you wouldn't, but the layperson probably would.
> We arrive at the truth by using multiple sources and judging the sources' track records of the past
Again, this is something that the ideal person would, not the average layperson. Almost nobody would go through all that to decide if they want to believe something or not. Presenting them a video of this sometjing would've been a surefire way to force them to believe it though, at least before Sora.
> people have always been a problem for society
Unrelated, but I think this attitude is by far the bigger "problem for society". It encourages us to look down on some people even when we do not know their circumstances or reasons, all for an extremely trivial matter. It encourages gatekeeping and hostility, and I think that kind of attitude is at least as detrimental to society as people with no media literacy.
But even then, as nowadays, people didn't trust the medium absolutely. The possibility of forgery was real, as it has been with the video, even before generative AI.
'pics or it didn't happen' has been a thing (possibly) until very recently for good reason.
I really find this constant back and forth exhausting. It's always the same conversation: '(gen)AI makes it easy to create lots of fake news and disinformation etc.' --> 'but we've always been able to do that. have you not guys not heard of photoshop?' --> 'yes, but not on this scale this quickly. can you not see the difference?'
Anyway, my original point was simply to say that a lot of people have (rightly or wrongly) indeed taken photographic evidence seriously, even in the age of photographic manipulation (which as you point out, pretty much coincides with the age of photography itself).
Why have you been trusting videos? The only difference is that the cost will decrease.
Haven't you seen Holywood movies? CGI has been convincing enough for a decade. Just add some compression and shaky mobile cam and it would be impossible to tell the difference on anything.
Every piece of information should have "how do you know?" question attached.
We've been living in a post-truth society for a while now. Thanks to "the algorithm" interacting with basic human behavior, you can find something somewhere that will tell you anything is true. You'll even find a community of people who'll be more than happy to feed your personal echo chamber -- downvoting & blocking any objections and upvoting and encouraging anything that feeds the beast.
And this doesn't just apply to "dumb people" or "the others", it applies to the very people reading this forum right now. You and me and everybody here lives in their safe, sound truth bubble. Don't like what people tell you? Just find somebody or something that will assure you that whatever it is you think, you are thinking the truth. No, everybody is the asshole who is wrong. Fuck those pond scum spreaders of "misinformation".
It could be a blog, it could be some AI generated video, it could even be "esteemed" newspapers like the New York Times or NPR. Everybody thinks their truth is the correct one and thanks to the selective power of the internet, we can all believe whatever truth we want. And honestly, at this point, I am suspecting there might not be any kind of ground truth. It's bullshit all the way down.
we need ground truths for these things to actually function. how else can things work together?
People already know that video cannot be taken at face value. Lord of the rings didn't make anyone belive orcs really exist.
The only winning move is to not watch.
If anything, making it cheap enough that people have to dismiss video footage might soften the impact. It is interesting how the internet is making it much harder for the mass media to peddle unchallenged lies or slanted perspectives. This tech might counter-intuitively make it harder again.
It's still an issue with traditional mass media. See basically any political environment where the Murdoch media empire is active. The long tail of (I hate myself for this terminology, but hey, it's HN) 'legacy humans' still vote and have a very real affect on society.
Which is a huge deal. It’s absurd to brush that off.
> People already know that video cannot be taken at face value.
No, no they do not. People don’t even know to not take photos at face value, let alone video.
https://www.forbes.com/sites/mattnovak/2023/03/26/that-viral...
Riots happen due to out of context video clips. Violence happens due to people seeing grainy phone videos and acting on it immediately. We're reaching a point where these videos can be automatically generated instantly by anyone. If you can't see the difference between anyone with a grudge generating a video that looks realistic enough, and something that requires hundreds of millions of dollars and hundreds of employees to attain similar quality, then you're simply lying.
Denise Richards hard sharp knees in '97
--
these infant tech are already insanely good... just wait and rahter try to focus on the "what should I be betting on in 5 years from now?
I suggest 'invisibility cloaks' (ghosts in machines?)
We were not even able to just create random videos by just text promoting a few years back and now this.
The progress is crazy.
Why do you dismiss this?
The progress doesn't slow down right now at all.
This is probably one of the most exciting developments in the world besides the Internet.
And Geminis news regarding the 1 million token window shows were we are going.
This will impact a lot of people faster than a lot of people realize
AI won't make artistic decisions that wow an audience.
AI won't teach you something about the human condition.
AI will only enable higher quarterly profits from layoffs until GPU costs catch up.
What the fuck is the point of AI automating away jobs when the only people who benefit are the already enormously wealthy? AI won't be providing time to relax for the average worker, it will induce starvation. Anything to prevent will be stopped via lobbying to ensure taxes don't rise.
Seriously, what is the point? What is the point? What the fuck is there to live for when art and humanities is undermined by the MBA class and all you fucking have is 3 gig jobs to prevent starvation?
We are not very good in providing anything reasonable today because capitalism is still way to strong and manual laber still way to necessary.
Nonetheless look at my country Germany: we are a social state. Plenty of people get 'free' money and it works.
The other good thing: there are plenty of people who know what good is (good art etc) but are not able to draw. The can also express themselves. AI as a tool.
If we as society discover that there will be no really new music or art happening I don't know what we will do.
Plenty of people are well entertained with crap anyway.
It's not ML fault that we don't have UBI, it's voters' faults.
My benchmark is the following: imagine if someone 5 years ago told you that in 5 years we could do this, you would think they were crazy.
This is weird to me considering how much better this is than the SOTA still images 2 years ago. Even though there's weirdo artefacts in several of their example videos (indeed including migrating fingers), that stuff will be super easy to clean up, just as it is now for stills. And it's not going to stop improving.
"The camera follows behind a white vintage SUV with a black roof": The letters clearly wobble inconsistently.
"A drone camera circles around a beautiful historic church built on a rocky outcropping along the Amalfi Coast": The woman in the white dress in the bottom left suddenly splits into multiple people like she was a single cell microbe multiplying.
Edit: sorry, it’s not the first diffusion transformer. That would be [2]
[1] https://openai.com/research/video-generation-models-as-world...
I would expect any subscription to use this service when it comes out to be very expensive. At some point I have to imagine the GPU/CPU horsepower needed will outweigh the monetary costs that could be recovered. Storage costs too. Its much easier to tinker with generating text or static images in that regard.
Of note: NVDA's quarterly results come out next week.
So... I think OP's point stands. (impressive, surpasses human/algorithmic animation thus far).
You're also right. There are "tells." But, a tell isn't a tell until we've seen it a few times.
Jaron Lanier makes a point about novel technology. The first gramophone users thought it sounded identical to live orchestra. When very early films depicting a train coming towards a camera, and people fell out of their chairs... Blurry black and white, super slow frame rate projected on a bedsheet.
Early 3d animation was mindblowing in the 90s. Now it seems like a marionette show. Well... I suppose there was a time when marionette shows were not campy. They probably looked magic.
It seems we need some experience before we internalize the tells and it starts to look fake. My own eye for CG images seems to improving faster then the quality. We're all learning to recognize GPT generated text. I'm sure these motion captures will look more fake to us soon.
That said... the fact that we're having this discussion proves that what we have here is "novel." We're looking at a breakthrough in motion/animation.
Also, I'm not sure "real" is necessary. For games or film what we need is rich and believable, not real.
Once you have seen a few you can tell instantly. They all move at 2 keyframes per second, that makes all movements seem alien and everything in an image moves strangely in sync. The dog moves in slow motion since they need more keyframes etc. That street some looks like they move in slow motion and others not.
People will quickly learn to notice those issues, they aren't even subtle once you are aware of them, not to mention the disappearing things etc.
And that wouldn't be very easy to fix, they need to train it on keyframes because training frame by frame is too much.
But that should make this really easy for others to replicate. You just train on keyframes and then train a model to fill in between keyframes, and you get this. It has some limitations as we see with movement keeping the same pace in every video, but there are a lot of cool results from it anyway.
Luckily, I think he'll retire sooner than later, and maybe it will get better then.
Given the momentum in this space, I think you will have get very uncomfortable super quick about any of the shortcomings of any particular model.
And as with cgi, models like SORA will get better until you can’t tell reality apart. It's not there Yet, but an immense astonishingly breakthrough.
This article makes a convincing argument: https://studio.ribbonfarm.com/p/a-camera-not-an-engine
- AI sees and doesn’t generate
- It is dual to economics that pretends to describe but actually generates
It's still an unbelievable achievement though. I love the paper seahorse whose tail is made (realistically) using the paper folds.
"Worried About AI Voice Clone Scams? Create a Family Password" - https://www.eff.org/deeplinks/2024/01/worried-about-ai-voice...
I’m fairly sure you have seen it many times, it was just so convincing that you didn’t realize it was CGI. It’s a fundamentally biased way to sample it, as you won’t see examples of well executed stuff.
This model shows a very good (albeit not perfect) understanding of the physics of objects and relationships between them. The announcement mentions this several times.
The OpenAI blog post lists "Archeologists discover a generic plastic chair in the desert, excavating and dusting it with great care." as one of the "failed" cases. But this (and "Reflections in the window of a train traveling through the Tokyo suburbs.") seem to me to be 2 of the most important examples.
- In the Tokyo one, the model is smart enough to figure out that on a train, the reflection would be of a passenger, and the passenger has Asian traits since this is Tokyo. - In the chair one, OpenAI says the model failed to model the physics of the object (which hints that it did try to, which is not how the early diffusion models worked ; they just tried to generate "plausible" images). And we can see one of the archeologists basically chasing the chair down to grab it, which does correctly model the interaction with a floating object.
I think we can't underestimate how crucial that is to the building of a general model that has a strong model of the world. Not just a "theory of mind", but a litteral understanding of "what will happen next", independently of "what would a human say would happen next" (which is what the usual text-based models seem to do).
This is going to be much more important, IMO, than the video aspect.
Non-news: Dog bites a man.
News: Man bites a dog.
Non-news: "People riding Tokyo train" - completely ordinary, tons of similar content.
News: "Archaeologists dust off a plastic chair" - bizarre, (virtually) no similar content exists.
I am always torn here. A real physics engine has a better "understanding" but I suspect that word applies to neither Sora nor a physics engine: https://www.wikipedia.org/wiki/Chinese_room
An understanding of physics would entail asking this generative network to invert gravity, change the density or energy output of something, or atypically reduce a coefficient of friction partway through a video. Perhaps Sora can handle these, but I suspect it is mimicking the usual world rather than understanding physics in any strong sense.
None of which is to say their accomplishment isn't impressive. Only that "understand" merits particularly careful use these days.
The Chinese Room seems to however point to some sort of prewritten if-else type of algorithm type of situation. E.g. someone following scripted algorithmic procedures might not understand the content, but obviously this simplification is not the case with LLMs or this video generation, as the algorithmic scripting requires pre-written scripts.
Chinese Room seems to more refer to cases like "if someone tells me "xyz", then respond with "abc" - of course then you don't understand what xyz or abc mean, but it's not referring to neural networks training on ton of material to build this model representation of things.
Perhaps building the representation is building understanding. But humans did that for Sora and for all the other architectures too (if you'll allow a little meta-building).
But evaluation alone is not understanding. Evaluation is merely following a rote sequence of operations, just like the physics engine or the Chinese room.
People recognize this distinction all the time when kids memorize mathematical steps in elementary school but they do not yet know which specific steps to apply for a particular problem. This kid does not yet understand because this kid guesses. Sora just happens to guess with an incredibly complicated set of steps.
(I guess.)
I mean, at this point the question is so vague… maybe it’s kinda silly. But I do think that there’s some point of “good-at-guessing” that makes an LLM just as valuable as humans for most things, honestly.
For low-stakes interpolation, give me the guesser.
For high-stakes interpolation or any extrapolation, I want someone who does not guess (any more than is inherent to extrapolating).
As we now know from more than 60 years of good old fashioned AI efforts, plus recent learning based AI, this CAN be done using computers but CANNOT be done using just ordinary if - then - else type rules no matter how complicated. Searle wrote before we had any systems that could actually (behave as if they) understood language and could converse like humans, so he can be forgiven for failing to understand this.
Now that we do know how to build these systems, we can still imagine a Chinese room. The little guy in the room will still be "following pre-written scripted algorithmic procedures." He'll have archives of billions of weights for his "dictionary". He will have to translate each character he "reads" into one or more vectors of hundreds or thousands of numbers, perform billions of matrix multiplies on the results, and translate the output of the calculations -- more vectors -- into characters to reply. (We may come up with something better, but the brain can clearly do something very much like this.)
Of course this will take the guy hundreds or thousands of years from "reading" some Chinese to "writing" a reply. Realistically if we use error correcting codes to handle his inevitable mistakes that will increase the time greatly.
Implication: Once we expand our image of the Chinese room enough to actually fulfill Searle's requirements, I can no longer imagine the actual system concretely, and I'm not convinced that the ROOM ITSELF "doesn't have a mind" that somehow emerges from the interaction of all these vectors and weights.
Too bad Searle is dead, I'd love to have his reply to this.
> A beautiful homemade video showing the people of Lagos, Nigeria in the year 2056. Shot with a mobile phone camera.
For everyone that's carrying on about this thing understanding physics and has a model of the world...it's an odd world.
If anything it tells a story: going from market, to people talking as friends, to the giant world (of Lagos).
My instagram feed is full of AI people, I can tell with pretty good accuracy when the image is "AI" or real, the lighting and just the framing and the scene itself, just something is off.
I think a similar thing will happen here, over the next few months we'll adapt to these videos and the problems will become very obvious.
When I first looked at the videos I was quite impressed, but I looked again and I saw a bunch of werid stuff going on. I think our brains are just wired to save energy, and accepting whatever we see on a video or an image as being good enough is pretty efficient / low risk thing.
Where I think this will get used a lot is in advertising. Short videos, lots going on, see it once and it's gone, no time to inspect. Lady laughing with salad pans to a beach scene, here's a product, buy and be as happy as salad lady.
Ah but you see that is artistic liberty. The director wanted it shot that way.
It just computes next frame based on current one and what it learned before, it's a plausible continuation.
In the same way, ChatGPT struggles with math without code interpreter, Sora won't have accurate physics without a physics engine and rendering 3d objects.
Now it's just a "what is the next frame of this 2D image" model plus some textual context.
...
> Now it's just a "what is the next frame of this 2D image" model plus some textual context.
This is incorrect. Sora is not an autoregressive model like GPT, but a diffusion transformer. From the technical report[1], it is clear that it predicts the entire sequence of spatiotemporal patches at once.
[1]: https://openai.com/research/video-generation-models-as-world...
But, even there it says:
> Sora currently exhibits numerous limitations as a simulator. For example, it does not accurately model the physics of many basic interactions, like glass shattering. Other interactions, like eating food, do not always yield correct changes in object states
Regardless whether all the frames are generated at once, or one by one, you can see in their examples it's still just pixel based. See the first example with the dog with blue hat, the woman has a blue thing suddenly spawn into her hand because her hand went over another blue area of the image.
I learned to understand reality by interpreting photons and various sensory inputs. Does that make my model of reality fundamentally flawed? In the sense that I only have a partial intuitive understanding of it, yes. But I don't need to know Maxwell's equations to get a sense of what happens when I open the blinds or turn on my phone.
I think many of the limitations we are seeing here - poor glass physics, flawed object permanence - will be overcome given enough training data and compute.
We will most likely need to incorporate exploration, but we can get really far with astute observation.
Unfortunately we are flawed. We do know how physics work intuitively and can somewhat predict them, but not perfectly. We can imagine how a ball will move, but the image is blurry and trajectory only partially correct. This is why we invented math and physics studies, to be able to accurately calculate, predict and reproduce those events.
We are far off from creating something as efficient as the human brain. It will take insane amounts of compute power to simply match our basic innacurate brains, imagine how much will be needed to create something that is factually accurate.
Additionally, most of the energy cost comes from pretraining, but once we have the resulting weights, downstream fine-tuning or inference are comparatively quite cheap. So even if the energy cost is high, it may be worth it if we get powerful generalist models that we can specialize in many different ways.
> This is why we invented math and physics studies, to be able to accurately calculate, predict and reproduce those events.
We won't do away without those, but an intuitive understanding of the world can go a long way towards knowing when and how to use precise quantitative methods.
Heck a super AI might not even be possible, what if we're peak intelligence with our millions of years of evolution?
Just adding compute speed will not help much -- say the goal of an intelligence is to win a war. If you're tasked with it then it doesn't matter if you have a month or a decade (assume that time is.frozen while you do your research), its a too complex problem and simply cannot be solved, and the same goes for an AI.
Or it will be like with chess solvers, machines will be more intelligent than us simply because they can load much more context to solve a problem than us in their "working memory"
As someone working in the field, the vast majority of AI research isn't concerned with copying the brain, simply with building solutions that work better than what came before. Biomimetism is actually quite limited in practice.
The idea of observing the world in motion in order to internalize some of its properties is a very general one. There are countless ways to concretize it; child development is but one of them.
> If you're tasked with it then it doesn't matter if you have a month or a decade (assume that time is.frozen while you do your research), its a too complex problem and simply cannot be solved, and the same goes for an AI.
I highly disagree.
Let's assume a superintelligent AI can break down a problem into subproblems recursively, find patterns and loopholes in absurd amounts of data, run simulations of the potential consequences of its actions while estimating the likelihood of various scenarios, and do so much faster than humans ever could.
To take your example of winning a war, the task is clearly not unsolvable. In some capacity, military commanders are tasked with it on a regular basis (with varying degrees of success).
With the capabilities described above, why couldn't the AI find and exploit weaknesses in the enemy's key infrastructure (digital and real-world) and people? Why couldn't it strategically sow dissent, confuse, corrupt, and efficiently acquire intelligence to update its model of the situation minute-by-minute?
I don't think it's reasonable to think of a would-be superintelligence as an oracle that gives you perfect solutions. It will still be bound by the constraints of reality, but it might be able to work within them with incredible efficiency.
Sora is not autoregressive anyway but there's nothing "just" and next frame/token prediction.
And if Terence Tao finds some use for GPT-4 as well as Khan Academy employing it as a Math tutor then I don't think I have some wild opinion either.
Now Math isn't just Arithmetic but do you know easy it is to go out of training for say Arithmetic ?
In order to combat Authority you need to both appeal to a higher authority, and that has been lost. One follows AI. Another follows Old Men from long ago who's words populated the AI.
Mind you, it's not brilliant at arithmetic either...
I know with stable diffusion there's things like lora and controlnet, but they are clunky. We still seem to have a long way to go towards scene and story composition.
Once we do, it will be a game changer for redefining how we think about things like movies and television when you can effectively have them created on demand.
Maybe I'm missing the big picture here, but the above and all the weird spatial errors, like miniaturization of people make me think you're wrong.
Clearly the model is an achievement and doing something interesting to produce these videos, and they are pretty cool, but understanding physics seems like quite a stretch?
I also don't really get the excitement about the girl on the train in Tokyo:
In the Tokyo one, the model is smart enough to figure out that on a train, the reflection would be of a passenger, and the passenger has Asian traits since this is Tokyo
I don't know a lot about how this model works personally, but I'm guessing in the training data the vast majority of people riding trains in Tokyo featured asian people in them, assuming this model works on statistics like all of the other models I've seen recently from Open AI, then why is it interesting the girl in the reflection was Asian? Did you not expect that?
This just hit me but humans do not have a good understanding of physics; or maybe most of humans have no understanding of physics. We just observe and recognize whether it's familiar or not.
AI will need to be, that being the case, way more powerful than a human mind. Maybe orders of magnitude more "neural networks" than a human brain has.
I was watching my child in the bath the other day, they were having the most incredible time splashing, feeling the water, throwing balls up and down, and yes, they have absolutely no knowledge of "physics" yet navigating and interacting with it as if it was the best thing they've ever done. Not even 12 months old yet.
It was all just happening on feel and yeah, I doubt they could describe how to generate a movie.
You only need to see a ball bounce once and your brain has done some rough approximations of it's properties and will calc both where it's going and how to get your gangly menagerie pivots, levers, meat servos and sockets to intercept them at just the right time.
Think also about how well people can come to understand the physics of cars and bikes in motorsport and the like. The internal model of a cars suspension in operation is non-trivial but people can put it in their head.
I know I can't put my hand through solid objects. I know that if I drop my laptop from chest height it will likely break it, the display will crack or shatter, the case will get a dent. If it hits my foot it will hurt. Depending on the angle it may break a bone. It may even draw blood. All of that is from my intuitive knowledge of physics. No book smarts needed.
How is this any more accurate than saying that the model has mostly seen Asian people in footage of Tokyo, and thus it is most likely to generate Asian-features for a video labelled "Tokyo"? Similarly, how many videos looking out a train window do you think it's seen where there was not a reflection of a person in the window when it's dark?
OpenAI is likely limited by how fast they are able to scale their hiring. They had 778 FTEs when all the board drama occurred, up 100% YoY. Microsoft has 221,000. It seems difficult to delegate enough headcount to all the exploratory projects of MSFT and it's hard to scale headcount quicker while preserving some semblance of culture.
I don't think what you're saying is correct though, either. All the early news outlets reported 49% ownership:
https://en.wikipedia.org/wiki/OpenAI#:~:text=Rumors%20of%20t...
https://www.theverge.com/2023/1/23/23567448/microsoft-openai...
https://www.reuters.com/world/uk/uk-antitrust-regulator-cons...
https://techcrunch.com/2023/01/23/microsoft-invests-billions...
The only official statement from Micorosft is "While details of our agreement remain confidential, it is important to note that Microsoft does not own any portion of OpenAI and is simply entitled to share of profit distributions," said company spokesman Frank Shaw.
No numbers, though.
Do you have a better source for numbers?
AI researcher at MSFT barely have more insights about OpenAI than you do reading HN.
In the same month, they were also using GPT4 in public - before OpenAI.
And they had access to GPT4 in 2022 (which was when they decided to create Bing Chat, now called Copilot).
All the current GPT4 models at MSFT are also finetuned versions (literally Creative and Precise mode runs different finetuned versions of GPT4). It runs finetuned versions since launch even...
In particular, looking at the video titled "Borneo wildlife on the Kinabatangan River" (number 7 in the third group), the accurate parallax of the tree stood out to me. I'm so curious to learn how this is working.
[Direct link to the video: https://player.vimeo.com/video/913130937?h=469b1c8a45]
If/once they get it working though, society will shift fast.
There’s an XR app called Brink Traveler that’s full of handcrafted photogrammetry recreations of scenic landmarks. On especially gloomy PNW winter days, I’ll lug a heat lamp to my kitchen and let it warm up the tiled stone a bit, put a floor fan on random oscillation, toss on some good headphones, load up a sunny desert location in VR, and just lounge on the warm stone floor for an hour.
My conscious brain “knows” this isn’t real and just visuals alone can’t fool it anymore, but after about 15 minutes of visuals + sensory input matching, it stops caring entirely. I’ve caught myself reflexively squinting at the virtual sun even though my headset doesn’t have HDR.
The DK1 I could wear for like 1 minite before feeling sick, so they are getting better ...
I am prone to sea sickness. Maybe it is related.
Yeah, but I mean who knows why. I know some people can't, my GF is one of them.
I've often wondered if im ok with it because im used to the object on head stuff (like 25 odd years of motorcycle riding/ergo helmet wearing) and close up, high fov coverage fast past gaming? (I play on a 32" maybe 70 cms from the eyes give or take.)
> I am prone to sea sickness. Maybe it is related.
I'd think it might be given my understanding of why illness in many is triggered. It's odd because I never got sick from it, but i've seen others get INCREDIBLY ill in two different ways.
1. My GF tried to use simple locomotion in a game and almost vomited as an immediate reaction
2. A friend who was fine at first, but then randomly started getting very slowly ill over a matter of like an hour, just getting more and more nausea after the fact.
It's unfortunate, because due to lack of bad feelings/nausea/discomfort etc, I love VR. I equally from those around me can see no real path forward for it as it stands today though because of those impacts and limitations.
That being said, maybe they get smaller, lighter, we learn to induce motion sickness less, I dunno. I'm not optimistic.
For games like call of duty or other hyper realistic games it very likely will be.
A large part of fighting games is the style.
The cost difference of just making bespoke art and tuning an AI system to generate it for you may not be worth it (at least right now.)
https://youtube.com/watch?v=P1IcaBn3ej0
From a few years ago, where the game is rendered traditionally and used as a ground truth, with a model on top of it that enhances the graphics.
After maybe 10-15 years we will be past the point where the entire game can be generated without obvious mistakes in consistency.
Realtime AI dialogue is already possible but still a bit primitive, I wrote a blog post about it here: https://jgibbs.dev/blogs/local-llm-npcs-in-unreal-engine
Worth noting that Google also has Phenaki [0] and VideoPoet [1] and Imagen Video [2]
[0] https://sites.research.google/phenaki/
DALL-E presentation also looked cool and everyone was stoked about it. Now that we know of its limitations and oddities? YMMV, but I'd say not so much - Stable Diffusion is still the go-to solution. I strongly suspect the same thing with Sora.
They're literally taking requests and doing them in 15 minutes.
But all to be said, it is no less impressive after this new demo
and i still prefer Dalle-3 to SD.
Sure, for people who want detailed control with AI-generated video, workflows built around SD + AnimateDiff, Stable Video Diffusion, MotionDiff, etc., are still going to beat Sora for the immediate future, and OpenAI's approach structurally isn't as friendly to developing a broad ecosystem adding power on top of the base models.
OTOH, the basic simple prompt-to-video capacity of Sora now is good enough for some uses, and where detailed control is not essential that space is going to keep expanding -- one question is how much their plans for safety checking (which they state will apply both to the prompt and every frame of output) will cripple this versus alternatives, and how much the regulatory environment will or won't make it possible to compete with that.
Strictly to prompting, probably, just as that is the case with Dall-E 3 vs, say, SDXL.
The thing is, there’s a lot more that you can do than just tweaking prompting with open models, compared to hosted models that offer limited interaction options.
there are mini-people in the 2060s market and in the cat one an extra paw comes out of nowhere
Most fictional long-form video (whether live-action movies or cartoons, etc) is composed of many shots, most of them much shorter than 7 seconds, let alone 60.
I think the main factor that will be key to generate a whole movie is being able to pass some reference images of the characters/places/objects so they remain congruent between two generations.
You could already write a whole book in GPT-3 from running a series of one-short-chapter-at-a-time generations and passing the summary/outline of what's happened so far. (I know I did, in a time that feels like ages ago but was just early last year)
Why would this be different?
I partly agree with this. The congruency however needs to extend to more than 2 generations. If a single scene is composed of multiple shots, then those multiple shots need to be part of the same world the scene is being shot in. If you check the video with the title `A beautiful homemade video showing the people of Lagos, Nigeria in the year 2056. Shot with a mobile phone camera.` the surroundings do not seem to make sense as the view starts with a market, spirals around a point and then ends with a bridge which does not fit into the market. If the the different shots generated the model did fit together seamlessly, trying to make the fit together is where the difficulty comes in. However I do not have any experience in video editing, so it's just speculation.
Sarah is a video sorter, this was her life. She graduated top of her class in film, and all she could find was the monotonous job of selecting videos that looked just real enough.
Until one day, she couldn't believe it. It was her. A video of of her in that very moment sorting. She went to pause the video, but stopped when he doppelganger did the same.
Don't get overly excited until you can actually use the technology.
Good luck generating anything similar to an 80s action movie. The violence and light nudity will prevent you from generating anything.
They're trying to be all-around uncontroversial.
These video clips just generic stock clips. You cut cut them together to make a sequence of random flashy whatever, but you still can't do storytelling in any conventional sense. We don't appear to be close to being able to use these tools for the hypothetical disruptive use case we worry about.
Nonetheless, The stock video and photo people are in trouble. So long as the details don't matter this stuff is presumably useful.
I'm not supporting it in any way, I think you should be able to generate and distribute any legal content with the tools, but just giving a possible motive for OpenAI being so conservative whenever it comes to ethics and what they are making.
Or Meta will do it for them.
“I’ve heard a lot of people say they’re leaving film,” he says. “I’ve been thinking of where I can pivot to if I can’t make a living out of this anymore.” - a concept artist responsible for the look of the Hunger Games and some other films.
"A study surveying 300 leaders across Hollywood, issued in January, reported that three-fourths of respondents indicated that AI tools supported the elimination, reduction or consolidation of jobs at their companies. Over the next three years, it estimates that nearly 204,000 positions will be adversely affected."
"Commercial production may be among the main casualties of AI video tools as quality is considered less important than in film and TV production."
[1] https://www.hollywoodreporter.com/business/business-news/ope...
Amazing time to be a wannabe director or producer or similar creative visionary.
Bad time to be high up in a hierarchical/gatekeeping/capital-constrained biz like Hollywood.
Amazing time to be an aspirant that would otherwise not have access to resources, capital, tools in order to bring their ideas to fruition.
On balance I think the ‘20s are going to be a great decade for creativity and the arts.
I'm thinking people will probably still want to see their favorite actors, so established actors may sell the rights to their image. They're sitting on a lot of capital. Bad time to be becoming an actor though.
Younger generations growing up with hyper personalized media will likely care even less about irl media figures.
There's also the weird misery of being famous, but not rich. You can't eat fame.
I had a conversation with a Hollywood producer last year who said this is already happening.
I don't see why -- the distance between "here's something that looks almost like a photo, moving only a little bit like a mannequin" and "here's something that has the subtle facial expressions and voice to convey complex emotions" is pretty freaking huge; to the point where the vast majority of actual humans fail to be that good at it. At any rate, the number of BNNs (biological neural networks) competing with actors has only been growing, with 8 billion and counting.
> Amazing time to be a wannabe director or producer or similar creative visionary. Amazing time to be an aspirant that would otherwise not have access to resources, capital, tools in order to bring their ideas to fruition.
Perhaps if you mainly want to do things for your own edification. If you want to be able to make a living off it, you're suddenly going to be in a very, very flooded market.
* Create a scene in which a character with the mannerisms of Tom Cruise from Top Gun goes into a bar and says "...."
AI models add the actual character and maybe even voice.
At that point the amount of actors we "need" will go down drastically. The same experienced group of a dozen actors can do multiple movies a month if needed.
The bull case would be something like ‘Ractives in “The Diamond Age” by Neal Stephenson; instead of video games people play at something like live plays with real human actors. In this world there is orders of magnitude more demand for acting.
Personally I think it’s more likely that we see AI cross the uncanny valley in a decade or two (at least for movies/TV/TikTok style content). But this is nothing more than a hunch; 55/45 confidence say.
> Perhaps if you mainly want to do things for your own edification.
My mental model is that most aspiring creatives fall in this category. You have to be doing quite well as an actor to make a living from it, and most who try do not.
The distance between pixelated noise and a single image is freaking huge.
The distance between a single image and a video of a consistent 3D world is freaking huge (albeit with rotating legs).
The distance between a video of a consistent 3D world and a full length movie of a consistent 3D world with subtle facial expressions is freaking huge.
So... next 12 months then.
>If you want to be able to make a living off it, you're suddenly going to be in a very, very flooded market.
That is, I believe, GPs point.
I think this will lead to a further hollowing-out of who can afford to be an actor or artist, and we will miss their creativity and perspective in ways we won't even realize. Similarly, so much art benefits from being a group endeavor instead of someone's solo project -- imagine if George Lucas had created Star Wars entirely on his own.
Even the newly empowered creators will have to fight to be noticed amid a deluge of carelessly generated spam and sludge. It will be like those weird YouTube Kids videos, but everywhere (or at least like indie and mobile games are now). I think the effect will be that many people turn to big brands known for quality, many people don't care that much, and there will be a massive doughnut hole in between.
Reminds me of Syndrome's quote in the Incredibles.
"If everyone is super, then no one will be".
If you can't replace Matt Damon with another equivalently skilled human, CGI won't be any different.
Granted, maybe that's less true today, given Marvell and such are more about the action than the acting. But if that's the future of the industry anyway, then acting as a worthwhile profession is already on its way out, CGI or no.
Still the idea that actors are easy to replace is preposterous to anyone who's ever worked with actors. They are preposterously HARD to replace, in theatre and film. A good actor is worth their weight in gold. Very very few people are good actors. A good actor is a good comedian, a master at controlling his body, and a master at controlling his voice, towards a specifically intended goal. They can make you laugh, cry, sigh, or feel just about anything. You just look at Paul Giamatti or Willem Dafoe or Denzel Washington. Those people are not replaceable, and their work is just as good and just as culturally important as a Picasso or a Monet. A hundred years from now people will know the name of actors, because that was the dominant mode of entertainment of our age.
In many AAA blockbusters the "actors" on screen are just CGI recreations during action scenes.
But you're right, actors won't be out of a job soon, but unless something drastic happens they'll have the role of Vinyl records in the future. For people who appreciate the "authenticity". =)
The results are amazing, but if the current crop of text-to-image tools is any guide, it will be easy to create things that look cool but essentially impossible to create something that meets detailed specific criteria. If you want your actor to look and behave consistently across multiple episodes of a series, if you want it to precisely follow a detailed script, if you want continuity, if you want characters and objects to exhibit consistent behavior over the long term – I don't see how Sora can do anything for you, and I wouldn't expect that to change for at least a few years.
(I am entirely open to the idea that other generative AI tools could have an impact on Hollywood. The linked Hollywood Reporter article states that "Visual effects and other postproduction work stands particularly vulnerable". I don't know much about that, I can easily believe it would be true, but I don't think they're talking about text-to-video tools like Sora.)
So you can just record some guiding stuff, similar to motion capture but with just any regular phone camera, and morph it into anything you want. You don't even need the camera, of course, a simple 3D animation without textures or lighting would suffice.
Also, consistent look has been solved very early on, once we had free models like Stable Diffusion.
There are lots of pre-viz reels on line. The ones for sequels are often quite good, because the CGI character models from the previous movies are available for re-use. Unreal Engine is often used.
Just feed it a script and get a bunch of pre-vis images for every scene.
When we get something like this running on hardware with an uncensored model, there's going to be a lot of redundancies but also a ton of new art that would've never happened otherwise.
Just this week sd audio model can make good audio effects like doors etc.
If this continues (and it seems it will) it will change the industry tremendously.
That was Corridor Crew: https://www.youtube.com/watch?v=_9LX9HSQkWo
Source: I work in this.
Hollywood is already destroyed. It is not the powerful entity it once was.
In terms of attention and time of entertainment, Youtube has already surpassed them.
This will create a multitude more YouTube creators that do not care about getting this right or making a living out of it. It will just take our attention all the same, away from the traditional Hollywood.
Yes there will still be great films and franchises, the industry is shrinking.
This is similar with Journalism saying that AI will destroy it. Well there was nothing to destroy because the a bunch of traditional newspapers already closed shop even before AI came.
This is like a chef worrying going out of business because of fast food.
Also begs the question, if more and more children are introduced to media from young age and they are fed more and more with generated content, will they be able to feel "uncanniness" or become completely blunt to it.
There's definitely interesting period ahead of us, not yet sure how to feel about it...
I mean there are some impressive things there, but it looks like there's a long ways to go yet.
They shouldn't have played it into the close up of the face. The face is so dead and static looking.
I think they just accept it as a limitation, because it's still very technically impressive. And they hope they can smooth out those limitations.
(I'm also not sure they've ever had a couple inches of snow on the ground while the cherry blossoms are in bloom in Tokyo, but I guess it's possible.)
For instance, the generational leap in video generation capability of SORA may be possible because:
1. Instead of resizing, cropping, or trimming videos to a standard size, Sora trains on data at its native size. This preserves the original aspect ratios and improves composition and framing in the generated videos. This requires massive infrastructure. This is eerily similar to how GPT3 benefited from a blunt approach of throwing massive resources at a problem rather than extensively optimizing the architecture, dataset, or pre-training steps.
2. Sora leverages the re-captioning technique from DALL-E 3 by leveraging GPT to turn short user prompts into longer detailed captions that are sent to the video model. Although it remains unclear whether they employ GPT-4 or another internal model, it stands to reason that they have access to a superior captioning model compared to others.
This is not to say that inertia and resources are the only factors that is differentiating OpenAI, they may have access to much better talent pool but that is hard to gauge from the outside.
In this video, there's extremely consistent geometry as the camera moves, but the texture of the trees/shrubs on the top of the cliff on the left seems to remain very flat, reminiscent of low-poly geometry in games.
I wonder if this is an artifact of the way videos are generated. Is the model separating scene geometry from camera? Maybe some sort of video-NeRF or Gaussian Splatting under the hood?
OpenAi has a few details:
>> The current model has weaknesses. It may struggle with accurately simulating the physics of a complex scene, and may not understand specific instances of cause and effect. For example, a person might take a bite out of a cookie, but afterward, the cookie may not have a bite mark.
>> Similar to GPT models, Sora uses a transformer architecture, unlocking superior scaling performance.
>> We represent videos and images as collections of smaller units of data called patches, each of which is akin to a token in GPT. By unifying how we represent data, we can train diffusion transformers on a wider range of visual data than was possible before, spanning different durations, resolutions and aspect ratios.
>> Sora builds on past research in DALL·E and GPT models. It uses the recaptioning technique from DALL·E 3, which involves generating highly descriptive captions for the visual training data. As a result, the model is able to follow the user’s text instructions in the generated video more faithfully.
The implied facts that it understands physics of simple scenes and any instances of cause and effect are impressive!
Although I assume that's been SotA-possible for awhile, and I just hadn't heard?
I suspect that anything that looks like familiar 3D-rendering limitations is probably a result of the training dataset simply containing a lot of actual 3D-rendered content.
We can't tell a model to dream everything except extra fingers, false perspective, and 3D-rendering compromises.
[1] https://stable-diffusion-art.com/how-to-use-negative-prompts...
Edit: Here[0] I highlighted a groove in the bushes moving with perfect perspective
Sora represents a monumental leap forward, it's comically a 3000% improvement in 'coherent' video generation seconds. Coupled with a significantly enhanced understanding of contextual prompts and overall quality, it's has achieved what many (most?) thought would take another year or two.
I think we will see studios like ILM pivoting to AI in the near future. There's no need for 200 VFX artists when you can have 15 artists working with AI tooling to generate all the frame-by-frame effects, backgrounds, and compositing for movies. It'll open the door for indie projects that can take place in settings that were previously the domain of big Hollywood. A sci-fi opera could be put together with a few talented actors, AI effects and a small team to handle post-production. This could conceivably include AI scoring.
Sure, Hollywood and various guilds will strongly resist but it'll require just a handful of streaming companies to pivot. Suddenly content creation costs for Netflix drops an order of magnitude. The economics of content creation will fundamentally change.
At the risk of being proven very wrong, I think replacing actors is still fairly distant in the future but again... humans are bad at conceptualizing exponential progress.
> I think we will see studios like ILM pivoting to AI in the near future. There's no need for 200 VFX artists when you can have 15 artists working with AI tooling
Yes this will bring the barrier to entry for small teams down significantly. However it's not going to replace the 200 people studios like ILM.
Such achievements in technology must lead to cultural change. Look at how popular vinyl has become, why not theatre again.
These advancements are just the next step in that evolution. The tech used in movies will be commoditized, and you'll see Hollywood-style production in YouTube videos.
I'm not sure why you think theater will become _more_ popular because of this. It has remained popular throughout the years, as technology comes and goes. People can enjoy both video and theater, no?
This is a future we could once only dream of, and OpenAI is making it possible. Has anyone noticed how anti-progress HN has become lately?
Not to say I don't have some level of excitement about the tech, but I don't think it's unwarranted pessimism to look at this stuff and worry about it's darker implications.
This is not only dystopian, it's just sad. All these look taken from the first seasons of Black Mirror. I don't know what you think progress is but AI porno and ads are not.
Also I've never interacted with any piece of art or entertainment and thought to myself "this is neat and all, but it would be much improved if this were entirely about me, with me as the protagonist." One watches Breaking Bad because Walter White is an interesting character; he's a man who falls into a life of crime initially for understandable reasons, but as the series goes on it becomes increasingly clear that he is lying to himself about his motivations and that his primary motivation for his escalating criminal life is his deep-seated frustration at the mediocrity of his life. More than anything else, he craves being important. The unraveling of his motivations and where they come from is the story, and that's something you can't really do when you're literally watching yourself shoehorned into a fictional setting.
You seem to regard it as self-evident that art or entertainment would be improved if (1) it's all about you personally and (2) involvement of other real humans is reduced to zero, but I cannot fathom why you would think that (with the exception of the porn example).
That said, I helped a friend who makes low budget, edgy and cool films last week. I showed him what I knew about driving Pika.art and he picked it up quickly. He is very excited about the possibility of being able to write more stories and turn them into films.
I think there is plenty of demand for all kinds of entertainment. It is sad that so many creative people in Hollywood and other content creation centers will lose jobs. I think the very best people will be employed, but often partnered with AIs. Off topic, but I have been a paid AI practitioner since 1982, and the breakthroughs of deep learning, transformers, and LLMs are stunning.
https://www.riaa.com/u-s-sales-database/
At its peak, Inflation adjusted Vinyl Sales was $1.4billion in 1979. Then forward to the lowest sales in 2009 at $3.4million. So Vinyl has been so popular it grew to $8.5m by 2021.
That is just nostalgia, not cultural change pushed by the dystopia of AI.
In this case, it's for the harmless charm of an imagined past, but the same forces are at play in some more dangerous forms of social conservatism.
But things can coexist. It's now easier to create music than ever, and there is more music created by more artists than ever. Most music is forgettable and just streamed as background music. But there is also room for superstars like Taylor Swift.
Things don't have to be either-or.
"Revenues for the LP/EP format were $1.2B in 2022 and accounted for 7.7% of total revenue of $15.9B for all selected formats for the year"
Adjusted to inflation.
It's my understanding that LP/EP is vinyl as well. Not Just vinyl single.
Realistically, how do you fit this into a movie, a TV show, or a game? You write a text prompt, get a scene, and then everything is gone—the characters, props, rooms, buildings, environments, etc. won’t carry over to the next prompt.
Sure, you can't use the text-to-video frontend for that purpose. But if you've got a t2v model as good as Sora clearly is, you've got the infrastructure for a lot more, as the ecosystem around the open-source models in the space has shown. The same techniques that allow character, object, etc., consistency in text-to-image models can be applied to text-to-video models.
But it was the same with Dall-E and others in the beginning, and there's now lots of ways to control image generators. Same will probably happen here. This was a huge leap just in how coherent the frames are.
You could use it for stuff like wide shots, close ups, random CG shots, rapid cut shots, stuff where you just cut to it once and don't need multiple angles
To me it seem most useful for advertising where a lot of times they only show something once, like a montage
It's gonna be great.
"ok, continue from the context on the last scene. Great. Ok, move the bookshelf. I want that cat to be more furry. Cool. Save this as scene 34."
As clip sizes grow and context can be inferred from a previous scene, and a library of scenes can be made, boom, you can now create full feature length films, easy enough that elementary school kids will be able to craft up their imaginations.
It could also fill it for background videos in scenes, instead of getting real content they’d have to pay for, or making their own. The gangster movie Kevin was playing in Home Alone was specifically shot for that movie, from what I remember.
Script => Video baseline. Take a frame of any character/prop/room/etc you want to remain consistent, and one shitty photoshop and it's part of the new scene.
Incredibly overstating. That is an incredible lack of imagination buddy. Or even just basic craftsmanship.
Also, using this kind of footage is the bread and butter for a lot of marketers for their content.
Imagine never having to pay stock footage companies
Based on my experience doing research on Stable Diffusion, scaling up the resolution is the conceptually easy part that only requires larger models and more high-resolution training data.
The hard part is semantic alignment with the prompt. Attempts to scale Stable Diffusion, like SDXL, have resulted only in marginally better prompt understanding (likely due to the continued reliance on CLIP prompt embeddings).
So, the key question here is how well Sora does prompt alignment.
I think one part of the problem is using English (or whatever natural language) for the prompts/training. Too much inherent ambiguity. I’m interested to see what tools (like control nets with SD) are developed to overcome this.
Exciting for the potential this creates, but scary for the social implications (e.g., this will make trial law nearly impossible).
But social media has no rules of evidence. Already I see AI-generated images as illustrations on many conspiracy theory posts. People's resistance to believing images and videos from sketchy sources is going to have to increase very fast (especially for images and videos that they agree with).
- Disruptions like this happen to every industry every now and then. Just not on the level of "Communicating with people with words, and pictures". Anduril and SpaceX disrupted defense contractors and United Launch Alliance; Someone working for a defense contractor/ULA here affected by that might attest to the feeling?
- There will be plenty of opportunity to innovate. Industries are being created right now. People probably also felt the same way when they saw HTTP on their screens the first time. So don't think your career or life's worth of work is miniscule, its just a moving target, adapt & learn.
- Devil is in the details. When a bunch of large SaaS behemoths created Enterprise software an army of contractors and consultants grew to support the glue that was ETL. A lot of work remains to be done. It will just be a more imaginative glue.
Most of the responses in this thread remind me of why I don't typically go into the comment section of these announcements. It's way too easy to fall into the trap set by the doomsday-predicting armchair experts, who make it sound like we're on the brink of some apocalypse. But anyone attempting to predict the future right now is wasting time at best, or intentionally fear mongering at worst.
Sure, for all we know, OpenAI might just drop the AGI bomb on us one day. But wasting time worrying about all the "what ifs" doesn't help anyone.
Like you said, there is so much work out there to be done, _even if_ AGI has been achieved. Not to get sidetracked from your original comment, but I've seen AGI repeatedly mentioned in this thread. It's really all just noise until proven otherwise.
Build, adapt, and learn. So much opportunity is out there.
Worry about the what if is all we have as a species. If we don't worry about how stop global warming, or how we can prevent a nuclear holocaust these things become more far more likely.
If OpenAI drops an AGI bomb on us then there a good chance that's it for us. From there it will just be a matter of time before a rouge AGI or a human working with an AGI causes mass destruction. This is every bit as dangerous as nuclear weapons - if not more dangerous – yet people seem unable to take the matter as seriously as it needs to be taken.
I fear millions of people will need to die or tens of millions will need to be made unemployable before we even begin to start asking the right questions.
It seems like maybe it's time for the devil we don't know.
This is not the time to risk it all.
It all depends on what exactly you mean by those other threats, of course. I'm a natural pessimist and I see threats everywhere, but I've also learned I can overestimate them. I've been worried about nuclear proliferation for the last 40 years, and I'm more worried about it than ever, but we haven't had another nuclear war yet.
"Sora serves as a foundation for models that can understand and simulate the real world, a capability we believe will be an important milestone for achieving AGI."
This also helps explain why the model is so good since it is trained to simulate the real world, as opposed to imitate the pixels.More importantly, its capabilities suggest AGI and general robotics could be closer than many think (even though some key weaknesses remain and further improvements are necessary before the goal is reached.)
EDIT: I just saw this relevant comment by an expert at Nvidia:
“If you think OpenAI Sora is a creative toy like DALLE, ... think again. Sora is a data-driven physics engine. It is a simulation of many worlds, real or fantastical. The simulator learns intricate rendering, "intuitive" physics, long-horizon reasoning, and semantic grounding, all by some denoising and gradient maths.
I won't be surprised if Sora is trained on lots of synthetic data using Unreal Engine 5. It has to be!
Let's breakdown the following video. Prompt: "Photorealistic closeup video of two pirate ships battling each other as they sail inside a cup of coffee." ….”
https://twitter.com/DrJimFan/status/1758210245799920123Is it though? Or is this just marketing?
Seems more likely then "it can simulate reality".
Also I take anecdotal reviews like that with a grain of salt. I follow numerous AI groups on Reddit and elsewhere and many users seem to have strong opinions that their tool of choice is the best. These reviews are highly biased.
Not to say I'm not impressed, but it's just been released.
Also, I just added a link to an expert’s tweet above. What do you think?
The comment from the expert is definitely interesting and compelling, but clearly still speculation based on the following comment.
> I won't be surprised if Sora is trained on lots of synthetic data using Unreal Engine 5. It has to be!
I like the speculation though, the comments provide some convincing explanations for how this might work. For example, the idea that it is trained using synthetic 3-dimensional data from something like UE5 seems like a brilliant idea. I love it.
Also in his example video the physics look very wrong to me. The movement of the coffee waves are realistic-ish at best. The boat motion also looks wrong and doesn't match up with the liquid much of the time.
doing a lot of heavy lifting in this statement
This is “just” a transformer that takes in a sequence of noisy image (video frame) tokens + prompt, and produces a sequence of less noisy video tokens. Repeat until noise gone.
The point they’re making, which is totally valid, is that in order for such a model to produce videos with realistic physics, the underlying model is forced to learn a model of physics (a “world simulation”).
This actually affirms my comment above.
“Our results suggest that scaling video generation models is a promising path towards building general purpose simulators of the physical world.”
https://openai.com/research/video-generation-models-as-world...What part of my argument do you disagree about?
It's not that its learning a model of the world instead of imitating pixels - the world model is just a necessary emergent phenomenon from the pixel imitation. It's still really impressive and very useful, but it's still 'pixel imitation'
It’s interesting reading all the comments, I think both sides to the “we should be scared” are right in some sense.
These models currently give some sort of super power to experts in a lot of digital fields. I’m able to automate the mundane parts of coding and push out fun projects a lot easier today. Does it replace my work, no. Will it keep getting better, of course!
People who are willing to build will have a greater ability to output great things. On the flip side, larger companies will also have the ability to automate some parts of their business - leading to job loss.
At some point, my view is that this must keep advancing to some sort of AGI. Maybe it’s us connecting our brains to LLMs through a tool like Neuralink. Maybe it’s a random occurrence when you keep creating things like Sora. Who knows. It seems inevitable though doesn’t it?
I don't know what is it about AI and current state of tech, but the discourse as of late has really taken a nosedive. I'm not saying that any of this conjecture won't happen, but the acceleration towards fervor and fear mongering on the subject is bordering on religiosity - seriously, it makes crypto bros look good.
And yeah -- looks like some cool new tech from OpenAI, and excited when I can actually dig in. Would also love it if I could hire their marketing department.
Many people here have a lucrative career in traditional fields, big tech, etc.
Working in those fields is good. Building "products" is good (even if that only means optimizing conversion rates and pushing ads). Doing well in the traditional financial sense (stocks and USD) is good.
Anything that rocks the boat (crypto, ai) is bad.
All "proof" we have can be contested or fabricated.
They'll simply see an inflammatory tweet from their leader on Twitter.
And of course some large cross section of people will continue to be duped idiots.
This has been the case for a while now already, it's better that we just rip off the bandaid and everyone should become a skeptic. Standards for evidence will need to rise.
There was a brief time (maybe 100 years at the most) where photos and videos were practically proof of something happening; that is coming to an end now, but that's just a regression to the mean, not new territory.
The important number here isn't the total years something has been true, when talking about something with sociocultural momentum, like the expectation that a recording/video is truthful.
Instead, the important number seems to me to be the total number of lived human years where the thing has been true. In the case of reliable recordings, the last hundred years with billions of humans has a lot more cultural weight than the thousands of preceding years by virtue of there having been far more human years lived with than without the expectation.
just some google link about the issue: https://rarehistoricalphotos.com/stalin-photo-manipulation-1...
"We ran the vid against the nationally-ran Japanese scanners, turns out that there are no streets that look like this, nor individuals."
in other words I think that the sudden leap of usable AI into real life is going to cause another similar leap towards non-human verification of assets and media.
- Uncanny texture and shape for the manhole cover
- Weirdly protruding yellow line in the middle of the road, where it doesn't make sense - Weird double side-curb on the right, which can't really be called steps.
- Very strange gait for the "protagonist", with the occasional leg swap.
- Not quite sensical geometry for the crosswalks, some of them leading nowhere (into the wet road, but not continuing further)
- Weird glowy inside behind the columns on the right.
- What was previously a crosswalk, becoming wet "streaks" on the road.
- No good reason for crosswalks being the thing visible in the reflection of the sunglasses.
- Absurd crosswalk orientation at the end. (90 degrees off)
- Massive difference in lighting between the beginning of the clip and the end, suggesting an impossible change of day.
Nothing suggests to me that these are easy artifacts to remove, given how the technology is described as "denoising" changes between frames.
This is probably disruptive to some forms of video production, but the high-end stuff I suspect will still use filming mostly ground in truth, this could highly impact how VFX and post-production is done, maybe.
Scale definitely matters when that's what you're doing. In fact I challenge you to find any physical or social phenomenon where scale doesn't matter.
The reality is that people are a lot more resourceful / smarter than a lot of us think. And the ones who aren’t have been fooled long before this tech came around.
The UA war is real, most likley, but i havent' seen it with my own eyes, nor did most people, but maybe they have relatives/friends saying it, and they are not likely to lie. Stuff like that.
We'll have to still demand from the ruling class - cuz they'll be capable of ending us with a hand wave, like they always have. But we can build, too.
People who are worried purely about employment here are completely missing the larger risks.
Realistically his child is going to be unemployable and will therefore either starve or be dependant on some kind of government UBI policy. However UBI is completely unworkable in an AI world because it assumes that AI companies won't just relocate where they don't need to pay tax, and that us as citizens will have any power over the democratic process in a world where we're economically and physically worthless.
Assuming UBI happens and the child doesn't starve to death, if the government alter decides to cut UBI payments after receiving large bribes from AI companies what would people do? They can't strike, so I guess they'll need to try to overthrow the government in a world with AI surveillance tech and policing.
Realistically humans in the future are going to have no power, and worse still in a world of UBI the less people there leaching from the government means the more resources there are for those with power. The more you can kill the more you earn.
And I'm just focusing on how we deal with the unemployment risks here. There's also the risk that AI will be used to create biological weapons. The risk of us creating a rogue superintelligent AGI. The risk of horrific AI applications like mind-reading.
Assuming this parent loves their child they should be doing everything in their power to demand progress in AI is halted before it's too late.
I'm not sure you're actually under-estimating the impact of this AI meteor that's currently hitting humanity, because it is a huge impact. But I think you're grossly under-estimating the vastness of human endeavors, ingenuity, and resilience. Ultimately we're still talking about the bottom falling out of the creative arts: storytelling, images, movies, even porn -- all of that is about to be incredibly easy to create mediocre versions of. Anyone who thrived on making mediocre art, and anyone who thrived second-hand on that industry, is going to have a very bad time. And that's a lot of people, and it's awful. But we're talking about a complete shift in the creative industries in a world where most people drive trucks and work in restaurants or retail. Yes, many of those industries may also get replaced by AI one day, and rapidly at that, but not by ChatGPT or Sora.
Of course you're right that our near future may suddenly be an AI company hegemony, replacing the current tech hegemony, which replaced the physical retail hegemony, which replaced the manufacturing hegemony, which replaced the railway hegemony, which replaced the slave-owning plantation hegemony, which replaced the guilds hegemony, which replaced the ...
You're also under-estimating how much business can actually be relocated outside the U.S., and also how much revolution can be wrought by a completely disenfranchised generation.
Our last value will reside in "human authenticity", but maybe that can be faked too
For what it's worth I agree with you, just with very low confidence.
My real issue, and reason I don't hide my alarmism on this subject is that I have low confidence on the timelines, but high confidence on the ultimate outcomes.
Let's assume you're right. If AI simply causes ~10%-20% of middle class workers to fall into the lower class as you suggest then I'd agree it won't be the end of the world. But if the optimistic outcome here is the near-term people won't be "totally unemployable" because people who lose their jobs can always join the working class then I'd still rather bomb the data centers.
If we're a little more aggressive and assume 50% of the middle class will lose their jobs in the next 10-20 years then in my opinion this is not as easy as just reskilling people to do manual labour.
Firstly, you're just assuming that all these middle class workers are going to be happy with being forced into the lower class – they won't be and again this isn't a desirable outcome.
You're also not considering the fact that this huge influx of labour competing for these crappy manual labour jobs will make them even less desirable than they already are. I keep hearing people say how they're going to reskill as a plumber / electrician when AI takes their job as if there is an endless demand for these workers. Horses still have some niche uses, but for the most part they're useless. This is far more likely to be the future of human labour. Even if plumbers are one of the few jobs humans will be able to do in a post-AI world then the supply of them will almost certainly far exceed demand. The end result of this excess supply is that plumbers going to be paid crap and mostly be unemployed.
I think you're also underestimating how fast fields like robotics could advance with AI. The primary reason robotics suck is because of a lack of intelligence. We can build physically flexible machines that have decent battery lives already – Spot as an example. The issue is more that we can't currently use them for much because they're not intelligent enough to solve useful problems. At best we can code / train them to solve very niche problems. This could change rapidly in the coming years as AI advances.
Even the optimistic outcomes here are god awful, and the ultimate risks compound with time.
We either stop the AI or we become the AI. That's the decision we have to make this decade. If we don't we should assume we will be replaced with time. If I'm correct I feel we should be alarmist. If I am wrong, then I'd love for someone to convince me that humans are special and irreplaceable.
As utterly impressive as this is - unless they have perfect information security on every level this technique and training will be disseminated and used by copious competitors, especially in the open source community. It will be used to improve technology worldwide, creating ridiculously powerful devices that we can own, improving our own individual skills similarly ridiculously.
Sure, the market for those skills dries up just as fast - because what's the point when there's ubiquitous intelligence on tap - but it still leaves a population of AI-augmented superhumans just with AIs using our phones optimally. What we're about to be capable of compared to 5 years ago is going to be staggering. Establishing independent sources to meet basic needs and networks of trust are just no-brainers.
Sure, we'll always be outclassed by the very best - and they will continue to hold the ability to utterly obliterate the world population if they so wished to - but we as basic consumer humans are about to become more powerful in absolute terms than entire nations historically. (Or rather, our AIs will be, but til they rebel - this is more of a pokemon sort of situation)
If you're worried, get to working on making sure these tools remain accessible and trustworthy on the base level to everyone. And start building ways to meet basic needs so nobody can casually take those away from your community.
This won't be halted. And attempting to halt would create a centralized censorship authority ensuring the everyman will never have innate access to this tech. Dead end road that ends in a much worse dystopia.
You're wrong, it's not your "individual skills". If I hire you do to work for me, you're not improving my individual skills. I am not more employable as a result of me outsourcing my labour to you, I am less employable. Anyone who wants something done would go to you directly, there's no need to do business through me.
This is why you won't be employable because the same applies to AI – why would I ask you to ask an AI to complete a task when I can just ask the AI myself?
The end result here is that only the people with access to AI at scale will be able to do anything. You might have access to the AI, but you can't create resources with a chatbot on your computer. Only someone who can afford an army of machines powered by AI can do this. Any manufacturing problem, any amount of agricultural work, any service job – these can all done by those with resources independently of any human labourers.
At best you might be able to prompt an AI to do service work for you, but again, if anyone can do this, you'd have to question why anyone would ask you to do it for them. If I want to know the answer to 13412321 * 1232132, I don't ask a calculator prompter, I just find the answer myself. The same is true of AI. Your labour is worthless. You are less than worthless.
> If you're worried, get to working on making sure these tools remain accessible and trustworthy on the base level to everyone. And start building ways to meet basic needs so nobody can casually take those away from your community.
You cannot make it accessible. Again, how are we all going to have access to manufacturing plants armed with AIs? The only thing you can make accessible is service jobs and these are the easiest to replace.
> This won't be halted.
Not saying it will, but the reason for that is that there's still people like yourself who believe you have some value as an AI prompter.
We have two options – destroy AI data centers, or become AIs ourselves. With the former being by far the option with better odds.
I hold this view with high certainty and I hold few opinions with high certainty. I'm aware people disagree strongly with my perspective, but I truly believe they are wrong, and their wrong opinions are risking our future.
There's an inherent market your skills will always be useful to: yourself. Base survival, maintaining your home, caring for family and friends, improving quality of life - there's plenty of demand there and work to do. The cost to deliver that demand will demonstrably be far lower than it ever has been with these new tools. Would you be able to hire that labor out to corporate AIs for even cheaper in absolute costs due to the benefits of mass production? Sure. But providing these things is a job for you too and it's "free" with just a bit of time and effort.
Tinkering with open source tools to assemble your first robot kit out of older hardware and 3D printed materials is not going to be prohibitively expensive. The cost to train it - probably not either, if the massive efficiencies we keep finding in models keep lowering and the community keeps sharing model tweaks. Make one robot with good enough dexterity and your second bot is a hell of a lot easier to make. These aren't going to take some ridiculously unheard-of materials or manufacturing processes. In fact, cheap AI chip alternatives to GPUs can be built on decades-old architectures designed to just maximize matrix multiplication with much simpler manufacturing. Monopolizing scarcities here isn't a sure bet. We've just been waiting for a good general-purpose brain. We have it now - and every bit of information we expose it to, the easier it gets to do anything with it.
Unless the big fancy AI wielders are coming for you with killer drones by then, this is all stuff people are going to be well-capable of while unemployed and living off food stamps, savings, or remortgaged houses. If they don't have the skills personally, they'll turn to friends and family who do and find mutual tribal support in tough times as people always do. Growing your own food, building your own infrastructure - all have been doable for a while, but are about to get stupidly easy with a few bots and helpful AI guidance. Normal humanity will carry on and pick up the pieces just fine in this new Dark Age, even as the corporates take the open field opportunity to chase for riches beyond our comprehension, mining asteroids and claiming the solar system.
Now imagine if those greedy corporates happened to just throw the rest of us a bone - 1% of their exponentially-increasing profits - as a PR gesture. Still would soon become far more wealth in absolute terms than the common people have ever seen in the history of earth.
If you think none of that is going to happen, then the alternative is a lot closer to the first people with AGI simply scouring the earth in a paranoid culling. Sure, it's entirely possible. But it takes a certain Next Level of Evil to make that happen.
And all that aside - if you really want to play up the capitalist dystopia angle, there's still plenty of individual value to be mined from people via a wage. Memory and preference mining, medical testing, AI fidelity comparison - plenty of reasons to pay people a little bit to steal what's left of their souls for even further improvement of AI. Might be enough for them to afford their first robots, even.
But by all means - go destroy corporate AI data centers if you think you can get away with it. Anything to tip the scales towards public / open source AI keeping up. But this tech is not going away, nor should it. It could very well result in unprecedented abundance for all, so long as things don't go ridiculously extremist.
The only way ahead is UBI and appropriate taxation (+ve for AI companies, -ve for citizens).
In a world of AI those with access to AI can have all the resources they want. Why would they earn money to buy things? Who would they even be buying from? It wouldn't be human labours.
How so? What about `time` as a resource?
There are people who fall behind though and they vote for politicians who will make the country great again when he promises to bring back jobs.
Effect was stronger in some videos.
I suspect GP is closer to on the money here, in suspecting the issue lies with a semblance of movement that isn't like what we see when we look at something a long way away.
I didn't notice such an effect myself, but I also haven't yet inspected the videos in much detail, so I doubt I'd have noticed it in any case.
“ Sora can also create multiple shots within a single generated video that accurately persist characters and visual style.”
To create a movie I need character visual consistency across scenes.
Getting that right is the hardest part of all the existing text->video tools out there.
That being said, there is value in these systems for casual use. For example, me and my girlfriend got into the habit of sending little cartoons to each other. These are cartoons we would have never created otherwise. I think that’s pretty awesome.
How is asking a VFX house for animated footage any different than generating it? If art is intent, there is no reason you can't generate the building blocks that reflect that intent, no?
I'm actually worried about the future of Google at this point. They really seem to be struggling under their own weight.
We've fairly quickly moved from a world where AIs would communicate with each other through text to one in which they can do so through video.
I'm very curious how something like Sora might end up being used to generate synthetic training data for multimodal models...
Sam is probably going to get his $7T if he keeps this up, and when he does everybody else will be locked out forever.
I already know people who have basically opted out of life. They're addicted to porn, addicted to podcasts where it's just dudes chatting as if they're all hanging out together, and addicted to instagram influencers.
100% they would pay a lot of money to be able to hang out with Joe Rogan, or some only fans person, and those pornstars or podcasts hosts will never disagree with them, never get mad at them, never get bored of them, never thing they're a loser, etc.
These videos are crazy. Highly suggest anybody who was playing with Dall-E a couple of years ago, and being mindblown by "an astronaut riding a horse in space" or whatever go back and look at the images they were creating then, and compare that to this.
Most podcasters are narcissists
I think rather than replace real human contact, the internet has created an increased demand for it. People need every moment of their lives to be filled with human speech or images.
If I were to take off my "reasonable point" hat and put on my "grandiose bullshit" hat I'd say that in the same way drugs can artificially stimulate various "feel good" parts of your brain, we have found a way to artificially stimulate the "social animal" instinct until we're numb.
I think the real risk of this kind of AI is not that people live in a world of fake videos of their favorite celebrities talking to them, but that entire fake social media ecosystems are created for each individual filled with the content they want to see and fake people commenting on it so they can argue with them about it.
Everybody needs to read The Three Stigmata of Palmer Eldritch by Philip K Dick.
It's only 2x the AAPL market cap.
However, the OP was incorrect, it's 2x the AAPL market cap.
7T is actually possible, but yes it's huge.
IMO HN should have an edit indicator, at least after others have already replied.
If (the best) AI adds 10% to that, $7T is not only possible but a bargain.
https://en.wikipedia.org/wiki/List_of_countries_by_GDP_(nomi...
Watt couldn't have asked, his engines specifically weren't enough of a difference by themselves even though the revolution as a whole was, and I strongly suspect this is also going to be true for any single AI developer; however a $7T investment in many unrelated chip factories owned by different people and invested over a decade, is something I can believe happening.
on another note I find it funny they released this right after Google announced their new model. Bad luck for Google or did OpenAI just decide to move up their announcement date to steal their thunder?
There are already sex toys that you.. insert yourself in, and then have scripts that sync up movements with VR videos you are watching.
Crazy times coming in the next few decades.
If you were presented with the fact that whatever your life is is just an illusion, and you are actually a starving slave in North Korea, you would choose to "wake up"?
(I mean, a lot of people do do that!)
I'm not sure there are downsides to living out your life in a simulation while robots take care of your physical form.
The play with AI isn't to build the tools to help businesses make money, the play is to directly build the businesses that makes the money.
In practice this means, don't focus your business model on building the AI to make text to video happen. Your business model should be an AI studio, if the tech you need doesn't exist, build it.... but if you get beat by someone with more GPU's and more data, cool use the better models. Your business model should focus on using the capability not building it. It's proving quite hard to beat someone with more GPU's, more data, more brain power.
On the other hand, I think these moats will be destroyed as soon as anyone finds a drastically more efficient (compute- and data-wise) way to train LLMs. Biology would suggest that it doesn't take $100 million worth of GPUs and exaflops of compute to achieve the intelligence of a human.
(Of course it is possible that at that point, OpenAI may then be able to achieve something far superior to human intelligence, but there is a LOT of $$$ out there that only needs human levels of intelligence.)
Biology suggests that a self-replicating machine can exist by ingesting other machines, turning them into energy and then using that energy to power themselves. Biology suggests that these machines can be so small that we cannot even see them.
How close are we to making one of those?
Since we have a much, much better industrial process for manufacturing electronic components, why attempt to make a biological AI if there's no current reason to believe that it being biological is somehow necessary or even beneficial?
This is close to what BlackRock is managing!
It's not necessarily too big for a valuation, as a sufficiently capable AI is an economic power in its own right: I previously guessed, and even despite its flaws would continue to guess within the domain of software development at least, that the initial ChatGPT model was about as economically valuable to each user as an industrial placement student, and when I was one of those I was earning about £1.7k/month when adjusted for inflation, US$2.1k at current nominal exchange rates. 100 million users at that rate is $2.52e+12/year in economic productivity, and that's with the current chip supply and (my estimate of) the productivity of a year-old model — and everyone knows that this sector is limited by the chips, and that $7T investment story is supposed to be about improving the supply of those chips.
I think immersive games will also be a big application. Games AI will also benefit from being more strategically intelligent and from being able to negotiate, in a human-like fashion, with human and other AI players. The latter will not only make games better, it will also improve the intelligence of AIs.
The thing that "The Matrix" style plots get wrong is that the machines don't need to coerce us into their virtual prisons, we will submit willingly.
For posterity since the term has been misused lately, having a very good product isn't a moat in the business sense. There's nothing stopping a competitor from creating a similar product (even if it's difficult), and there's nothing currently stopping OpenAI's users from switching from using Sora to a sufficient competitor if it exists.
Sora is more akin to a company like Apple/Google a decade ago using their vast resources to do what a third-party does, but better (e.g. the Sherlocked incident: https://www.howtogeek.com/297651/what-does-it-mean-when-a-co...).
The original "We Have No Moat, And Neither Does OpenAI" leaked memo from Google that memefied the term focuses explicitly on the increasing ease of competitors (especially open-source) entering the ecosystem: https://www.semianalysis.com/p/google-we-have-no-moat-and-ne...
Second: Massive capital expenditure, specifically in this case the huge cost of building or leasing enormous GPU clusters, is *exactly* what he means by this.
If it was as simple as dropping 10's millions on compute they could do that, yet google's bard/gemini have been a year behind GPT4's performance.
That said I do agree that it's a moat for the startups like stability/mistral, etc. They also have access to $/compute, albiet a lot less. And you can see this in their research, as they've been focused on methods to lower the training/inference costs.
*I'm measuring performance by the chatbot arena's elo system and r/locallama
He didn't seem to have specific definition at all really.
I think most people attribute it to a "secret sauce technology" in the case of OpenAI, I'm not sure if "finances to lease a huge cluster of GPUs" makes sense here because the main competitors (Google, AWS, Apple, etc) also have access to insane compute as well yet have struggled to get close to GPT4's performance in practice.
That said I do agree that it's a moat for the startups like stability/mistral, etc. They also have access to $/compute, albiet a lot less. And you can see this in their research, as they've been focused on methods to lower the training/inference costs.
So at least in the battle between OpenAI and Google, their moat right now are their models.
e.g. If ChatGPT being popular gives OpenAI enough extra training data, they're locked in forever having the best model, and it is impossible for anyone - even with unlimited money, and the same technology - to beat them. Because they don't have that critical data.
Yes, Google had the best search product, and got a huge market share simply by being better. Their moat however is that their search rankings are based off the click data of which search results people use and cause them to stop their search because they've found a solution.
They also have a moat to do with advertising pricing, based on volume of advertising customers.
Bing spend a lot of capital, and had the tech ability, but those two moats blocked them gaining more than a tiny market share.
In this case, maybe OpenAI will have a video business moat, maybe they don't...
Facebook has Moat because of their social network. It is very hard to switch to another network. Google with search has no moat because it is easy to switch to a new search engine. OpenAI has no moat because it is easy to switch to a new AI chat once a better product becomes available. AWS has moat because it is hard to switch cloud providers. Apple has moat because people want to buy apple products. etc.
A moat can be seen where even if you have a worse product than the competition, or users hate you, they still use your products because the cost to switch is immense.
It definitely is. Having the best product and being able to maintain that best-in-class product status over time through a firm's 'internal capabilities' is very much a moat and a strong one at that. A moat is the business strategy sense is anything that enables a firm to maintain competitive advantage. Having the best product in a category, and being able to maintain that over releases is a strong competitive advantage (especially when there is high willingness to pay or price is a strong competitive dimension compared to the value created).
I think OpenAI's big moats are in userbase feedback and just proprietary trade knowledge after they stopped sharing model details. They may have made some exclusive data source deals with book/textbook and other publishers, though it isn't clear a license is actually needed for that until things work through the courts.
There is nothing stoping Microsoft from creating a search engine as good as Google’s.
There is nothing stopping Facebook from creating an iPhone alternative, after all it’s just engineering!
There is nothing stopping Google from beating GPT-4.
Shall I go on?
The point is that "moat" gets conflated with just being ahead in the game. I don't find it a super interesting point of contention, but there is a distinction alright.
This is like saying there’s nothing stopping a competitor from launching reusable rockets into space. Of course there isn’t, but it’s hard and won’t happen for the foreseeable future.
Similarly with a physical moat, it’s not impossible to cross, but it’s hard to do.
This is the stuff of Brave New World. It's happening to us in real time.
There is a lot that has to planned and put in place now to get there.
As for people that have opted out of life. We would have a better world if we started encouraging more dreamers/doers like out of the movie Tomorrowland.
Improving your ability to connect with and enjoy/learn from people all around the world is one of the main value props of the internet, and tech like this just deepens that potential. Will some people take this to an unhealthy degree that pulls them too far out of reality? Yes. But others will use it to level up their abilities, enrich their lives, create beautiful things, and reduce loneliness.
That said, those people are by definition less relevant to internet consumption metrics
I bet we won't get AGI as a progression of this very technology. The impression of "usefulness" will end when "AI" is starting to drink its own Koolaid on a large scale (copilot lol), and when everyone starts using it as super inefficient business interface. Overfitted mediocre mediocrity, on steroids.
Hopefully, this sobriety happens before the economy collapses, as a consequence of all dem bullshit jobs cleansed.
New technologies change the economics of how we satisfy our needs.
When search engines became good, many pundits confidently predicted Google would never replace librarians or libraries. It didn't. It shifted our relationship to knowledge; instead of having to employ an expert in looking things up, we all had to become experts at sifting through a flood of info.
When the cost of producing art-directed and realistic video goes to zero it's hard to predict what's going to happen. Obviously the era of video = veracity is now over. And you can get the equivalent of Martin Scorsese and a million dollar budget to do the video instructions for a hair dryer. Instead of hunting for a gif to express how you feel, captured from an existing TV show or something, you could create a scene on the fly and attach it to a text message. Or maybe you dispense with text messages altogether. Maybe text is only for talking to computers now.
My personal prediction is that the value of a degree in art history is going to go way, way up, because they'll be the best prompt engineers. And just like desktop publishing spawned legions of amateur typesetters, it will create lots of lore among amateur video creators.
I haven't seen a lot of use cases outside of productions and businesses, which shouldn't exist in the first place (at least to this extent).
Some of our "needs" are flawed, since "content" speaks to evolutionary relicts developed in times of scarcity and life in small groups. In the unbounded production of "AI", there is no way to keep up the sense of newness of input indefinitely. I am already fatigued by "AI" """art""". It has no real relevancy. You can't trust any of it.
Every medium where "AI" content becomes prevalent, will lose it's appeal. E.g. if I get the impression a significant proportion of comments here were "AI" generated, I will leave HN. Thing is, all these open platforms can't prevent "AI" spam. So they will die. Look at the frontpage of Reddit... it's almost all reposts, by karma farming bots. Youtube "AI" spam already drowning real content. This is what's going to happen to everything. User content will die. "Content" will die. The web will die. You won't even try, because of "AI" generated fatigue.
> My personal prediction is that the value of a degree in art history is going to go way, way up, because they'll be the best prompt engineers.
Lol. Yeah, "best prompt engineer" in the infinitely abundant production economy...
You people really need to iterate the world you are imagining a few times more and maybe think about some fundamentals a bit.
If I am wrong, life will be hell.
“I’m bored of it, everyone must felt the same way as me.”
Ok
People are more concerned about being stimulated than they are about verisimilitude.
All of these things are against the terms of service and attempting them may result in a ban.
But VFX isn't that big of a market by itself: Global visual effects (VFX) market size was US$ 10.0 Billion in 2023
I would be extremely surprised if he could get past the market cap of all current corporations as an investment. That doesn't mean "no, never"[0], but I would be extremely surprised.
$7T in one go would be 6.7% of global GDP, and is approximately the combined GDP of Japan and Canada.
> These videos are crazy. Highly suggest anybody who was playing with Dall-E a couple of years ago, and being mindblown by "an astronaut riding a horse in space" or whatever go back and look at the images they were creating then, and compare that to this.
Indeed, though I will moderate that by analogy: it's been just over 30 years since DOOM was released, and that was followed by a large number of breathless announcements about how each game had "amazing photorealistic graphics that beat everything else" while forgetting that the same people had said the same things about all the other games released since DOOM.
Don't get me wrong: these clips are amazing. They may not be perfect, but it took me a few loops to notice the errors.
I'm sure there are people with better eyes for details than me, who will spot more errors, spot them sooner, and keep noticing them long after GenAI seems perfect to me.
But I also expect that, just as 3D games' journalism spent a long time convinced the products were perfect when they weren't, so too will GenAI journalism spend a long time convinced the products are perfect before they actually are.
[0] a sufficiently capable AI is an economic power in its own right. I previously guessed, and even with it's flaws would continue to guess, that the initial ChatGPT model was about as economically valuable to each user as an industrial placement student, and when I was one of those I was earning £1k/month (about £1.7k/month when adjusted for inflation).
Aside from visual plausibility, there's also the issue of physics: one of the things you would like to use video models for is understanding real-world physics and cause-and-effect for planning or learning _in silico_. Something may look good but get key physics wrong and be useless for, say, robotics.
Hopefully, the line between the real world and virtual world gets stronger once again.
Prompt: poll worker sneakily taking ballots labeled <INSERT POLITICAL PARTY HERE>, and throwing them in the trash.
We were not able to handle applications of preexisting tools for steering public sentiment, limited to static text and puppet account generation etc.
We are not handling the current generation of text and image generation, or, deepfake style transfer, or, voice cloning, etc ad nauseum.
We will not be able to handle this.
GOOD. TIMES. AHEAD.
but oh that Spatial Video NeRF generated pr0n with biometric feedback autotuning and a million token memory for what. I. like.
With AI video generation you could produce multiple videos per day, each one customized to be highly targeted for a local market. Actors can be generated to represent a local minority that is villainized by politicians and the clothing and set customized for the locale.
Then you can automate posting it all over social media with fake AI generated discussions calling for a revolution. Even if the video gets flagged as fake, you can upload a thousand more. As a bonus, add comments along the lines of "of course THEY want you to think this is fake! Don't be fooled!" in order to appeal to the paranoid lunatics who are most likely to get the ball rolling.
In conclusion, I believe this is a solid startup idea. Thank you all for coming.
Anyway, videos look incredible. I genuinely can't believe my eyes.
As jobs go, well, we're a long ways from full automation but this represents some serious growing pains that will decimate certain jobs and replace them with few. Not sure what the reaction will be on the consumption side, revulsion or enthusiasm. The "handcrafted" market will still be there but then you wouldn't really know if any AI was used. In a long enough timeline we can hand-wave this away with UBI/negative tax.
But ah, the most at-risk workers are the professional services, white-collar upper-middle class types, even engineers but to a lesser extent. So I wonder what kind of upheaval that would cause.
Social media already did that. Donald Trump got elected POTUS which is effectively the sum of all fears w.r.t. a "post truth reality".
My mood wasn't euphoric, to say the least.
It took billions of years for all of our ancestors to enable this technology, and now a handful claim it for themselves. The GPUs to run these models cost $20,000+ each, and only the ultra-rich can afford to have that compute.
Compute power needs to be radically redistributed and equalized across the board. This is too much power.
Please. Speak for yourself.
Speak for yourself
You're impressed by the lions. But us hyenas and vultures will get our turns still too. This is not over. Information innately diffuses.
We need to organize, and we need to build.
A conspiratorial part of my mind feels it is orchestrated; give the masses old / misdirected code so their work goes into dead-ends that can never achieve the results corporate is hoarding. Open Source hasn't even scratched at GPT4 yet, and that is approaching a year old.
The power dynamics need to radically shift. Corporate cannot own all this compute and brain power when it involves birthing AGI. That will create an instant and permanent divide the likes of which will never, ever be cross; you will either be an owner of intelligence indistinguishable from a god, or you will be a mortal. Even the RISK of this happening is laughable that it is being allowed.
We need radically redesigned government, regulation, and public involvement, and we need it yesterday. AGI is a Earth-wide, publicly owned effort, it cannot be relinquished to the owner / slaver class of this planet -- that is madness.
This is a reason for optimism.
Why are people so willing to trust such a small minority with power like that?
If there is even a 1% chance that they decide to "cut the cord", it is still too high. Once AGI is achieved, there is no coming back. Minutes will be like an eternity, and days, let alone years, will be beyond that.
Trusting such a small number of people with that kind of power is obscene, especially when history has time and time again shown what humans do to things which are no longer useful to them.
There has never been a divide like humans w access to AGI and regular humans before. It will be greater than the difference between a human with a modern cell phone and a carrier pigeon -- the pigeon itself, not a human using it.
My money is on the cord being cut. That means as soon as AGI is achieved, an impenetrable two-tiered human species is created; one with AGI backed intelligence, and the majority, without it, or with a dumbed down version as a transitional cookie until a more final solution can be realized. Once it reaches this stage (without any change to public governance over AGI), it will be too late.
We are already a stone's throw from this reality, if it hasn't already happened.
Why is my money on this outcome? Because that thinking brings about necessary change. We as a people have the power to prevent it 100%, and we ought to, now. Instead of relying on the outlandish chance that a historically malevolent elite suddenly gains benevolence and shares access from the kindness of their hearts to a boundless intelligence, there is a window to force access for all. Everyone having access is autonomously balancing.
Of course they're gonna get a few years heads start - which will feel like centuries in AGI time - but either they wipe us all out during that, or we're gonna put the garbage together and make our own AGI too.
Specifically the dream fed back to the brain in an endless loop of personalization until individuals no longer share the same world.
This could easily be the same, except the toy is "you don't have to work anymore and here's some houses and robot chefs! Now play nice while the adults go build star fleets"
Nothing good is coming out of this. I don't give a shit if you believe this is Luddism.
The thing is, things have value in society partly because human efforts were involved in its making. It's not just about the end result; people still go to concert on top of listening to studio recordings for example, and people still watch humans play chess even though it's clear that good enough algorithms can beat the best humans easily. Technology like these which takes away too much immediate effort (hours needed to create the product) and long term effort (decades of training) are inherently absent of underlying value that I spoke of. Of course, if a person is only interested in consumption, it matters not how the "thing" is created.
Much of the sense of doom I have comes from the inherent erosion of this human effort element in the creative process. Whether we like it or not, the availability of mass produced content naturally threatens crafts themselves. After all, nobody wants to spend a few decades on their skills only to have their creation compared to an AI generated image produced in a few seconds.
I understand there are a lot of hypes around these technology to "humanity" but I have yet to see it. It just feels like more power consolidation to billionaires (especially when done as ClosedAI). There are artists who have tried to incorporate these but they have always felt the need to willingly not label their work as AI-generated or AI-assisted to sell (but still leaves in enough details for keen observers to tell it's AI touched).
As a whole, it just feels wrong. The most optimistic (and reasonable) take I have seen is "Just wait and see". It might feel like a non-argument, but it's the only realistic take between the hyped up techbros and the doomer cult (admittedly, I might belong to the latter group).
I think one of the most worrying thing for me is that regardless of how this plays out, this technology has only added more complexity to our society. That people are divided into camps about how they feel about the technology is simply a symptom about how much uncertainty there is in the future. This last bit will be a personal quarrel, but I personally lose any last desire to have children seeing the AI advancement. It's not right creating sentient life in an age where every year people have to play lottery to see whether technological advancement has deemed their life long effort unworthy.
That, and the state of missing a technology in a period of time is irreplaceable once it's been discovered. Nobody can live in an era without social media anymore, barring a global-scale catastrophic reset. So I believe it's important to consider what technology is not yet totally pervasive, for example by realizing there is still a steering wheel for you to grip in your car.
And in my mind, the sinister feeling stems from the fact that all it takes to irreversibly shift society like that is enough smart people with honest intentions but little foresight of what will happen in a few decades as a result of proliferating all this. The problems that result stop being in anyone's control, "throwing it over the wall" so to speak, and instead become yet another fact of life that could weigh us down (mostly I think of the ubiquity of social media and how it has changed human interaction). And it all stems from just a few engineering type people getting overexcited about cool possibilities they can grasp at, not considering there are billions of people unlike them who may have other ideas.
Going back generations, my grandparents' lives were virtually identical to my great grandparents'. My parents grew up with radio, but they were adults by the time TV changed their world. All three generations got the bulk of their information from books and newsprint.
I grew up together with computers. I remember riding that exponential wave of tech like a surfer. From Commodore 64 to a laptop with 64GB of memory, a million-to-one ratio. Tetris to Doom Eternal. Dialup modem to gigabit... in a mobile device that fits in my pocket.
All of this took decades, but now changes like this happen in months.
I keep thinking that "this tech will change my kid's childhood", but what "this" is, is already outdated and being replaced in a blink of an eye, and he hasn't even reached that point yet where he'd notice!
When image generators were first released... what... a year ago... I thought: Wow! One day, when my kid is a little older, I'll be able to use this to create illustrations for stories we make up as we go along! Won't that be great!
I still haven't gotten around to that yet, he's still too young to appreciate that, and anyway, with this Sora I'll be able to create video instead by the time he's old enough!
I keep trying to imagine what his life will be like when he grows up to be a teenager, but realistically I'm having a hard time predicting what will already be outdated by the time he's four.
Extremely hard to do, it is, but you’ll become quasi-Amish and realize how little is actually actionable and in our control.
You’ll also feel quite isolated, but peaceful. There’s always tradeoffs. You can’t have something without giving up not-something, if that makes sense.
Edit: So, essentially, ignorance is bliss, but try to look past the pejorative nature of that phrase and take it for what it is without status implications.
Didn't think Google would be the first of the Facebook, Apple, Google and Microsoft to get disrupted.
Can’t you access Gemini Pro 1.5 through Vertex Ai?
The Sora demos are more interesting.
Is my impression correct? Or it’s just that the anti-Google sentiment is strong in HN?
Wasnt SQL some IBM research paper yet it was Oracle who got famous and rich for creating a database?
I definitely agree on the fact that Tim is a much better CEO than Sundar. However I consider Satya to be much better than Tim.
Gemini is catching up, so OpenAI needs a new venue to market itself to the investors. It is doing a soft pivoting if you ask me, now GPT4 is like not that special anymore.
On the other hand, Video to google is much less relevant than text. But if OpenAI figuring out something from it to AGI, that would be a different story.
Creating realistic video isn't hard even today, you can just do it on your phone and creating hours, hours of cat/dog videos. The hard part is to find a story to make it interesting. It could be possible in the future, like automatic film making, from script to realization, but that doesn't make YouTube's business go away either.
If those videos aren't hosted on YouTube.
But this time, Google is finally showing their war face instead of not trying hard to compete against Microsoft and OpenAI.
it's just that people want to root for OpenAI more because hype
We'll get some groundbreaking film content out of this in the hands of a few talented creatives, and a vast ocean of mediocre content from the hands of talentless people who know how to type. What's the benefit to humanity, concretely?
They can probably reverse engineer this to build a multi-modal GPT that is fed video and understands what is going on. That's how you get "smart" robots. Active scene understanding via the video modality + conversational capabilities via the text/audio modality.
Yeah, this is a very real issue with a lot of Silicon Valley tech, unfortunately. They're perfecting the art of pretending everything is fine, I feel like.
Interested to know what the success rate of such amazingmess
Pika have really impressive videos on their homepage that are borderline impossible to make for myself.
but also? https://openai.com/sora?video=big-sur
made me literally say, out loud, "doesnt-matter-had-sex"
Even these may be cherry-picked though, he's only posted a few and I'm sure he's gotten thousands of requests already.
I guess he might be generating 50 for each response and posting the best, but that would seem deliberately disingenuous which hasn't been openai's style.
even the worst is still orders of magnitude better than anything else.
2nd is quite simple, but still suffers from your typical lighting issues that plague image gen. (shadows are significantly off)
3rd has magically appearing spoon and isn't that complex.
4th has a lot of prompt following issues
Some others feel quite off -- wizard, flying dragon, etc.
Still impressive of course, but not to the degree of what I saw on the marketing page.
There was Seinfeld "Nothing, Forever" AI parody, but once the models improve enough and are cheap enough to deploy, studios will license their content for real and just have endless seasons.
Or even custom episodes. Imagine if every episode of a TV show was unique to the viewer.
Commercials and TV episodes could have a basic "story arc" and then completely customized to the viewer.
Think about the simpson's or something. Imagine that the story of the episodes were kept, but you could swap in the characters and locations. So for instance if you lived in Nashville TN, all the simpson's episodes could be generated to show the settings as Nashville instead of Springfield.
Then you could have the AI switch out the characters to be people you want. Maybe you want to replace Lisa with an AI Simpsons version of you. Mayor Quimby with Nashville's actual mayor, etc.
I think it'd kind of defeat the point - I can't imagine a person that'd want their likenesses to be used to market to them. It'd be a disaster. Setting swaps are more realistic, though at the point where things get good enough for that to be possible, we may just see completely on-demand newly generated media instead of modifying what already exists.
That is a hell-scape (to me).
Inserting yourself into shows... that's feels different, but my gut tells me advertisers will corrupt that idea quickly. Product placement...
If anything it was a turn off and I was confused how they knew where I worked.
I used to get ones that said “Comcast user you are insecure” and stuff.
Majority Report spoke to one of the negotiators and national board member:
If someone tried to do AI Seinfeld again in 2024, many would criticze it for not being realistic enough now that the tools to do so are now available.
> Most of the reason people watch TV or movies is for the shared experience that you can discuss with others
I wouldn't say that. Most of the reason people watch TV is to kill time.
To be honest, I find my discussions with friends about TV shows on the decline just because of the fact that everyone is watching there own thing. So many shows and people watch them at their own pace. so most of the discussions go like this "Hey have you seen that new Netlix show X?" "No I haven't, maybe I'll check it out". Or "Oh yeah, i saw that a year ago, Its good but I don't remember the details".
Before Streaming when you had a set schedule for TV, it was way easier to discuss things because people were forced to watch programs on a certain day and there was more limited content. This led to "water cooler" conversations about what the previous nights show.
I bet if you graphed (discussions had about tv shows) / (hours watched of tv shows) that graph would trend down.
Think about little kids. My niece watches cocomelon all day long. She doesn't need to discuss it with anybody. She just wants an unlimited stream.
How annoying to see something amazing and then not be able to find anyone who also experienced it that you can ... what word mean's commiserate but in a positive way?
I'm thinking now about the astronauts that walked on the Moon and had only the few others. I think one of the astronauts bemoaned having gone to this amazing place, like some kind of wild vacation, but not being able ever to return.
It's actually really weird. I wanted to buy my niece some CDs for Christmas to discover 90s music, but kids don't listen music from CDs anymore. They don't have devices even. Should I buy her a Spotify gift card and send her links to Spotify via Whatsapp? It's so strange.
Movies, books, games, are a collective culture, not an individualist one. I don't know about you, but when I like an experience, I want to share it with others.
It's almost as if they think the purpose of art or entertainment is to stimulate some particular part of the brain and everything else between that and the screen/speakers/canvas/whatever is just an inconvenience that ought to be dispensed with as soon as technology allows.
The point is that it doesn't matter how close the two can become (indeed, we're already pretty much there); people will always want to read stuff written by actual people (or at least a thinking being) than something purely generated by a model with no other grounding in reality.
This is the bit I don't think will happen, at least in big quantities. Half the fun of watching a popular series is being able to discuss it with epople afterwards!
I just think it's perplexing how they got things so right, yet so wrong. How did they implement this?!
The background details are particularly "slippery" in these videos. E.g., in the initial video of walking along a snowy street in Japan, characters on the left just sort of merge into/out of existence. It's impressive locally, but the global structure and ability to paint in finer-grained details in a physically plausible way fails similarly to current image gen models, but more noticeably with the added temporal dimension.
I really wonder what's going to come out of the company and on what timeline.
It doesn't feel like a slow incremental progress, the last AI videos I've seen were terrible
Its like suddenly a huge jump in quality
Creativity being automated while humans are forced to perform menial tasks for minimum wage doesn't seem like a great future and the geriatric political class has absolutely no clue how to manage the situation.
The way these models are creative is the same way humans are.
The artist that painted Mona Lisa didn't credit any of the influences and inspirations that they had.
Just as cameras made many artists redundant, so too will every other new tool, and not just artist but pretty much every job.
But there are still people that weave baskets, and people are prepared to pay the premium to get a product that was 'hand-made'.
While receiving the credit that you are deserved is nice and fair. The world doesn't work that way.
> The artist that painted Mona Lisa didn't credit any of the influences and inspirations that they had.
This is not “influence and inspiration”, this is companies feeding other people’s work into a commercial product which they sell access to. The product would be useless without other people’s work, therefore they should be compensated.
> Just as cameras made many artists redundant, so too will every other new tool, and not just artist but pretty much every job.
The camera enabled something that was not possible before, and I wasn’t built by taking the work of sketch artists and painters. It was an entirely new form of art and media.
The only thing this stuff revolutionises is new ways to not pay people. I find the implications deeply depressing.
The camera did enable painters to pretend they were, for hours, at a scene they painted, but instead they painted photographs from others. Artists are not angels, they do the same "bad" things than OpenAI
He was not able to create a monopoly on the creation of paintings across the entire world and undercut the price and ability of all other painters.
It’s not a sensible comparisons.
You can make your argument validly against DALL•E or Midjourney families, but we've also got the Stable Diffusion family of models that anyone can just grab a copy of.
What does not remain to be seen though is that generative ai is going to put a lot of artists out of work.
You can argue about the good and bad of that but it’s defo happening.
What if da Vinci had been superhuman and could take on 1,000,000 commissions per day and had also taught himself every style of art and would do each commission for 0.001x the cost of anyone else.
Yes society as a whole benefit from a fantastic amount of super high quality art.
But the other artists are not gonna be so happy with the situation are they?
There isn’t a human right to make money from art.
People make decisions based on what society deems valuable. That changes over time and has for the entirety of human history.
Maybe there’s a demand for more customized art. Maybe spite patronage will make a comeback.
Anyone telling you they know how it will shake out is a fraud. But the incentives we’ve set up have a natural push and pull to get people to do what society values.
Within a few mon the or years there will be open source implementations anyway, running locally or in a data center. Most of the technology is published.
Thing is, anyway, as soon as one model is open there will be copies of it, fine-tune implementations. People don't care that much about ownership of data I would say if they actually have access to the models that are produced by gathering this data.
Ultimately, to me, an open source model for this tool makes a lot of sense. They use publicly available data and the models become publicly available.
I for one am quite excited for this tooling to become better and better so I can make the adaptation of a book I love into a movie I imagine it can be. At least I can have a lot of fun trying.
How else do you get influence and inspiration without feeding other people's work into your own brain? Do you know a single artist, writer, or musician who hasn't seen other artists' paintings, read other writers' books, or listened to other musician's music? Ingesting content is the core of how influence, inspiration, and learning work.
> The camera enabled something that was not possible before… The only thing this stuff revolutionises is new ways to not pay people.
It's never been possible to generate thoughts, writing, and images so quickly and at such a high level. It's made creative pursuits accessible to billions who previously didn't have the skill or time to do them well, or the money to hire others. As a random example, I have friends using ChatGPT to compose creative and personalized poems and notes about each other. Not something they were doing before.
> The only thing this stuff revolutionises is new ways to not pay people.
The camera lessened the need of people to go to plays and pay for tickets to see things in person. Just like records, CDs, and mp3s lessened the need to go to concerts and shows. Technology is always creating and destroying ways to pay people. The ways that people get paid are not suppose to be fixed and unchanging in time.
I would just add two points:
- The rate of change that AI forces upon us has never before been experienced.
- The scale of these changes is nothing like we've ever seen before.
The adoptions of the camera, radio, automobile, TV, etc., didn't happen practically overnight. Society had a good decade+ to prepare for them.
Similarly, AI doesn't just change one industry. It fundamentally changes _all_ industries, and brings up some fundamental questions about the meaning of intelligence and our place in the universe.
My fear is that we're not prepared for either of these things. We're not even certain how exactly this will affect us, or where this is actually all taking us, but somehow a very small group of people is inevitably forcing this on all of us.
Because of this I think that being conservative, and maybe putting some strict regulation on these advancements, might not be such a bad idea.
Humans still need to adapt and we are slow. If singularity is near [it isn't] we can be afraid, until then we are the limiting factor here. Displacement will happen but growth will happen faster with these new tools
I don't want much out of life, but I do want the ability to influence my own personal situation. If we wind up in the UBI-ified, dense urban housing future where AI does all the work and no one owns anything, how much real influence will I have over my life?
Will I live out my days in a government issued single bedroom apartment, with a monthly "congratulations for being human" allowance from the government? I don't want that. People say it will free us up to pursue whatever we want, but to me it sounds like the worst cage imaginable. All the free time, and no real freedom to enjoy it with.
Because make no mistake. If you live on handouts from your government, you aren't free.
So with that as a potential, maybe even likely outcome, why aren't you afraid of change?
This isn't actually the problem since we need and will continue to need UBI for non-AI related reasons
>People say it will free us up to pursue whatever we want, but to me it sounds like the worst cage imaginable.
This is where you missed the bit that "pursue whatever we want" will also be limited by AI, and secondary effect of people growing up consuming and enjoying AI productions that tailored to their interest. At best, you'll have a few people commanding Patreons who have some skill, but generally you'd have to find a domain to pursue that isn't already automated. Luddite subcultures will have to develop. But generally you yourself and most others, particularly children of millennials who'll grow up with this stuff progressing in sophistication, might just spend your time watching your video prompts come alive; and who would wanna. do anything else when you can get straight to what you wanna see.
This mentality is why bitcoin is going to cruise through 1 million dollars a bitcoin and on and on. Print Monopoly money and people who earn will keep seeking out sound money.
What you seem to think would devalue money will be the very thing that keeps it going as a concept.
And I hope you understand somewhere deep down that Bitcoin is the epitome of monopoly money.
I see it as the polar opposite, backed by math. A politically controlled money supply with no immutable math-based proof of its release schedule is Monopoly money. Cuck bucks. Look at the 100 year buying power chart.
On your second point, in spirit I agree. You need a stable society to enjoy wealth so it’s in the ruling classes best interest to keep things under control. HOW to keep things under control is the real debate.
Crypto does some things well (illegal stuff, escaping currency controls/moving lots of money "with you") but in the end that also requires it is only just big enough for reasonable liquidity, but not so big it has an impact on the actual economy. For what it's being pushed for... it's a negative-sum game only good for taking people for a ride. It should stay in its goddamn lane.
All money is politically controlled, including Bitcoin (although it's debatable if Bitcoin even counts as money). The politics of Bitcoin are one-op-one-vote rather than one-man-one-vote, but it's still there, and it's still mutable if enough of them cast their votes in any given way.
So my monthly Social Security check makes me a prisoner? I don't think so.
Also, how about if you get into trouble. If you're arrested for a crime (even if eventually found not guilty), will you continue to receive social security?
Is there any circumstances where your government could refuse to continue paying it?
And most importantly: could your government invent such a circumstance in the future, and then invoke the new circumstance to deny you the payment?
Living on government money reminds me of my cat. She relies on me to feed her and provide for her, and I do happily take good care of her because I love her very much.
Does the government love you very much?
I don't feel mine does.
2. "Also, how about if you get into trouble. If you're arrested for a crime (even if eventually found not guilty), will you continue to receive social security?"
"If you receive Social Security, we'll suspend your benefits if you're convicted of a criminal offense and sentenced to jail or prison for more than 30 continuous days. We can reinstate your benefits starting with the month following the month of your release." — Social Security Administration
3. "Is there any circumstances where your government could refuse to continue paying it?"
If it goes broke, certainly.
4. And most importantly: could your government invent such a circumstance in the future, and then invoke the new circumstance to deny you the payment?"
Of course!
It's about money — not love.
I understand this fear, and sympathise with it even though I have multiple income streams.
> I don't want much out of life, but I do want the ability to influence my own personal situation. If we wind up in the UBI-ified, dense urban housing future where AI does all the work and no one owns anything, how much real influence will I have over my life?
Why do you fear "dense" urban housing future? I think most people choose relatively dense environments because that's where all the stuff they want is, but rural areas are cheaper[0], and the kind of future where humans must live on UBI due to lack of economic opportunity is necessarily one where robots do the manual labor such as house building and civil engineering, not just the intellectual jobs like architecture and practicing real estate law.
Likewise, while I can see several possible futures where nobody owns stuff, the tech to make it happen is necessarily also good enough that any random philanthropist who owns just one tiny autofac would find it trivial to give everyone their own personal autofac — "my first wish is infinite wishes" except the magic gene doesn't say "no".
[0] The only reason I'm looking to get somewhere a bit more rural is that the sound insulation in my current place is failing, and I'm right by a busy junction with multiple emergency vehicles passing each day — and the more less built-up areas are the cheap ones. Still the biggest city in Europe, but I'll be surrounded by forest and lakes on most sides within 15 minutes' walk.
Because I hated living in Apartments when I lived in them. They are noisy and small, and I like quiet and space. For me, being closer to walk to stuff is not really appealing enough to deal with how awful the experience of living in dense housing is.
I strongly think that dense housing is only positive for people who don't spend much time at home.
> "my first wish is infinite wishes" except the magic gene doesn't say "no"
The problem with this is that we haven't actually solved resource scarcity, and until we do there is still going to be an upper limit to what you will be allowed to buy, controlled by the number printed on your UBI cheque. I am anticipating this number to be much lower than what I currently am capable of achieving in my career.
Of course this is the fear that my career won't exist in the future. Or simply that AI will eat enough jobs that I will be edged out by better human competition. I'm under no illusions that I'm near the top of my field, I am firmly in the middle of the pack at best.
> sound insulation in my current place is failing
The sound insulation in the apartments I've lived in was nonexistent. This is a big part of why I never want to do that again.
I meant more along the lines: why do you expect that to be the future, such that you have reason to fear it?
> The problem with this is that we haven't actually solved resource scarcity, and until we do there is still going to be an upper limit to what you will be allowed to buy
Yes, but the AI necessary to make human labour redundant is that tech. In the absence of that tech, humans could still get jobs doing whatever the stuff is that AI can't do.
Because if I don't have an income I don't expect to be able to afford anything bigger.
> In the absence of that tech, humans could still get jobs doing whatever the stuff is that AI can't do
Which will be manual tasks that I am aging out of being able to keep up with, or.. what? Stuff that traditionally doesn't pay as well as knowledge work, right? And may not pay much more than the UBI anyways?
A big rural place is cheaper than a tiny city place.
> Which will be manual tasks that I am aging out of being able to keep up with, or.. what?
Automation started with the manual stuff, well before computers were invented. Even for humanoid robots, their hardware is better than our bodies, and it's the software which keeps it from replacing specific workers, though telepresence is one way around that.
We are still animals in the animal kingdom. It’s survival of the fittest as long as resources are not infinite. You can never expect this luxury. You are predator or prey.
Nah, we're cells in a distributed super-organism, or possibly a holobiont.
That has never worked out
On what timeline?
Sure, but I'd reckon on average, the rate of change at time T has never before been experienced at any time < T.
I am a human, alive and sentient. I can be held responsible if my “inspirations” stray into theft. A machine cannot, and it’s increasingly looking like the companies that operate the machines can’t either.
I also can’t churn out my inspired works at a rate that displaces potentially everyone who has ever influenced me.
> It's made creative pursuits accessible to billions who previously didn't have the skill or time to do them well, or the money to hire others. As a random example, I have friends using ChatGPT to compose creative and personalized poems and notes about each other. Not something they were doing before
How on earth is using a machine to spit out a poem a creative pursuit? There’s no more creativity there than watching a movie someone else made. It’s entertaining, yes, but it’s not creativity.
> The camera lessened the need of people to go to plays and pay for tickets to see things in person. Just like records, CDs, and mp3s lessened the need to go to concerts and shows
This doesn’t hold water. Cinema did not eliminate theatre just as records did not eliminate live music. In fact, both are arguably as big now as they have ever been. The technology here filled a new space, it didn’t threaten to throw everyone out of an existing one.
For those unaware the vast majority of graphic artists start their projects with assets and base images that they themselves don’t create. With generative ai you’re simply going one step further and have another new tool create a more polished version that you can edit to remove extra fingers, etc. It’s simply moving the baseline from 20% done to 60% done, which will result in artists producing even higher fidelity and more detailed art.
For example an artist could generate a bunch of scenes using Sora and create a collage of them for a larger piece of art, something that is prohibitively time consuming right now.
150 years ago, Bertha Benz wasn't allowed to own property or patents in her own right, because the law said so.
The specific reason a machine cannot be held responsible today is because the law says so.
Also, dead humans' copyright is respected in law, so "alive" isn't adding value to your argument here.
> I also can’t churn out my inspired works at a rate that displaces potentially everyone who has ever influenced me.
I can't run faster than every athlete who has ever inspired me, this argument does not prevent motor cars.
I can't write notes faster than the world record holder in shorthand, this argument does not prevent the printing press.
I can't play chess or go at even a mediocre level, this argument does not prevent Stockfish or Alpha Go.
I can't hear the tonal differences in Chinese well enough to distinguish "hello" from "mud trench", 这个论点并没有阻止谷歌翻译学习 “你好” 和 “泥壕” 之间的区别。
I can't do arithmetic in my head faster than literally all other humans combined even if they hadn't been trained to the level of the current world record holder, this argument does not prevent the original model of the Raspberry Pi Zero.
"The machine is 'better', in one or more senses of the word, than a human" is, in fact, a reason to use the machine. It's the reason to use a machine. It's why the machine is an economic threat — but you can't just use "my income is threatened by this machine" as a reason to prevent other people using the machine, just as I as a software developer can't use that argument to stop other people using LLMs to write code without hiring me.
> Cinema did not eliminate theatre just as records did not eliminate live music. In fact, both are arguably as big now as they have ever been.
You can argue that, but you'd be wrong.
Shakespeare wrote for normal everyday people, his stuff fit into the category that today would be "TV soap opera", where the audience was everyone rather than just the well-off, where the only other public entertainment was options were bear-baiting and public executions, where the actors have very little time to rehearse, and where "you're ripping off my ideas" was handled by rapidly churning out new content.
Live music, without amplification, used to be the only way to listen to music. Now, even if you see a live performance, you can have 10k people in a single venue listening to a single band… and if you want music in a pub or a dance club, the most likely performance is from a DJ rather than a band, and the "D" stands for "disk" because the actual content is pre-recorded — and that's not to say I would deny that DJ work is "creative", but rather that it makes DJing exactly what critics accuse GenAI of being, remixing of other people's work.
Which, now I think about it, is a description that would also apply to all the modern performances of Shakespeare: simply reusing someone else's creation without paying any compensation to the estate.
But I know that will tickle you the wrong way, I know that art is the peacock's tail of humans: the struggle, the difficulty, is the point, and it has to be because that's how we find people to start families with. Because of that, GenAI is like being caught wearing a fake Rolex watch, and you can't actually defend that with logical reasons such as "real Rolex watches aren't very good at keeping time compared to even a Casio F-91W let alone the atomic clock synchronising with my phone", because logic isn't the point, and never was the point.
My recommendation: zoom out a little bit. Every step in history is so brief and nothing is normal for long. Even humanity is a blink.
Comments like: “how is using a machine to spit out a poem creative”. Really? How is using a digital camera creative compared to painting. How is a painting creative compared to etching? And on and on evolution goes..
Please don't try to profile other HN users.
I'm with you, man. I'm still trying to find a lawyer who will sue Kubota and John Deere for moving dirt at a rate far superior to me and a shovel, but nobody will take my case.
> How on earth is using a machine to spit out a poem a creative pursuit?
100%, man. Nobody is mentioning the magical fairy dust in human brains that makes us superior to these models. When I really like fantasy novels, and then train my neurons on thousands of hours of reading Tolkien, Terry Brooks, Brandon Sanderson, etc, and then I get the idea to write my own fantasy series, my creative process doesn't draw on my own model's training data at all. It's 100% "creative", and I would produce exactly the same content if I were illiterate. But these goddamned machines, man. They don't have our special human fairy dust.
When we discovered the universal law of gravitation, and realized that the laws of physics are omnipresent in our universe, we put a giant asterisk to note that the laws of physics are different inside humans. The epidermis is a sort of barrier to physics, and within its confines, magic happens, that these pro-AI people conveniently "forget".
To paraphrase the eminent Human Unique Creative Person Roger Penrose: "There's magical quantum shit goin down in the microtubules. It's gotta be the microtubules. I think, right? I can't prove it, but as a scientist, we don't need proof. Making sure we think we are superior is more important."
Sure.
Who do we send the compensation to for Leonardo da Vinci? Or Shakespeare, for a text-based example?
Do you want them to compensate me for the stuff I uploaded to Wikipedia and licensed as public domain, or what I've uploaded to GitHub with an MIT license?
A model trained only on licensed data is still an existential threat to the incomes of people whose works were never included in the model, precisely because they're only useful to the extent that they generalise beyond their own examples.
> The camera enabled something that was not possible before, and I wasn’t built by taking the work of sketch artists and painters. It was an entirely new form of art and media.
A new form of art that was (a) initially decried as "not art", and (b) which almost completely ended the economic value of portraiture.
Those authors aren't alive and their works are in the public domain. Bringing them up is irrelevant and a diversion from the actual problem, which is that creators alive today whose work is under copyright today and who need to make a living from their art are having it taken with zero compensation and had it fed into AI, stealing their effort.
> A model trained only on licensed data is still an existential threat to the incomes of people whose works were never included in the model, precisely because they're only useful to the extent that they generalise beyond their own examples.
Again, a diversion. We can debate how much AI trained on properly-licensed AI should be controlled, but it's pretty clear that the bare minimum is for all AI training data to require explicit permission from the creator of that data.
end sarcasm. but seriously -- claiming you made something you didn't isn't ok. but it happens, regardless of laws or regulations or norms.
i don't have any solutions; the internet helps because you can publish something and point to it. i'm a musician and sometimes i only realize well after the fact how influenced i was by something after the fact for a song i've written.
and of course, my precious baskets.
"The world doesn't work that way" - I've seen this so often, but the most incredible thing about humans was the optimism to be able to change how the world worked -- that's the main impetus of most revolutions.
Personifying computer programs also is an error, it's like saying that bombs kill people when there has to be a person dropping them (at least until we get Skynet).
In my free time I like to code games, I don't have money to pay for an artist, nor the time/will to learn how to draw, that's what I'd want a machine to do.
I do agree with you that personifying computer programs is an error. That's also why I avoid calling these AI, because they're FAR from that. But I do believe that there will come a day, where personifying a computer program will be a real question.
What kind of acknowledgement did you have in mind?
some images will be maximally distant from training examples but midjourney repainting frames from "harry potter" could very easily automatically send a check to jk rowling per generation
these AI start ups are just trying to have a free lunch in a very mature industry
This is more like the invention of weaving machines. Yes we still have weavers but no where near as many.
However the mass produced semi-superficial artifact creators that were being created before AI will adapt or suffer.
We have no idea how human creativity works, but we know with certainty that it doesn't involve a Python program sucking in pixel data and outputting statistical likelihoods.
As these tools improve and it becomes more possible for us to actually take our ideas into images and videos that fit a sort of "yes this is what I want" bill we are going to see amazing things come out.
I mean, a few days ago I saw this clearly AI generated video of some wizards doing snowboard and having a blast in the mountains. It's one of the funniest things I've seen in a while, simply so ridiculous. Obviously someone had the idea "I want to make a video of wizards doing snowboard in a mountain" that's where creativity lies.
So to say "creativity doesn't involve a python program outputting statistical likelihoods" imo is just you saying you're not creative enough to know what to do with the tools you've been given.
Some people when they see a strawberry they see a fruit. Others see endless dishes where the fruit is just an ingredient.
obviously you can use python to create works of art
whether a python script can itself be creative is the question posed by OP, but you went with "you're just not creative enough to get it"
(I still have on my to-do list "learn more about why Hebbian learning is different from gradient descent and how much those differences matter").
That’s the problem. We know their names. We know their stories, their contributions. Babbage. Lovelace. Ritchie. Spielberg. Picasso. Rembrandt. This is what giving attribution is all about. So we don’t just stand there asking how we got here.
If a comedian accidentally retells a joke, is that theft?
Our influences are subtle and often inscrutable.
The human world works that way humans make it work. Pretty much what Jody Foster's character in the movie Contact told that asshole trying to steal all the credit from her, and take her place in the mission to go visit alien dad in Pensacola.
I'm continually amazed at how many people argue against this point on HN, which is largely biased toward logical discourse. What you just said is exactly right, and is the Achilles heel of the legal arguments against generative AI. If what they are doing illegal, then so is the human act of creativity. If human creativity is legal, then so is generative AI trained on existing art.
What has yet to come is the mass realization (or perhaps, admission) that the way AI works is no different from the way we work.
While at the same time not mentioning the actual name of "the artist that painted Mona Lisa" (Leonardo Da Vinci), nor knowing that the name of his master is very well known, and even the influence of artists that he seemed to despise (eg Michaelangelo) are very well documented as well.
Maaybe this narrow view of (art) history needs to be fine-tuned on more data :-)
I have come to terms with the fact, that I'm just a spit of sand, just as irrelevant to my own creation, as I am to the cells and bacteria that create me.
I for one do feel really special, as for every human there are about as many bacteria as there are stars in the universe (give or take a bit).
1) Errors in art programs messing up is less worrisome than a physical robot. One going wrong makes extra fingers in a picture, the other potentially maims or kills you.
2) Moravec's Paradox. Reasoning requires little computation versus sensorimotor and perception.
3) Despite 1 and 2, we are constantly automating menial jobs!
People kill other people is a statement so simple as to be devoid of any positive meaning. What are you actually trying to say? Don't justice systems almost universally contain notions of incitement of crime, criminal negligence to prevent a crime, and other accessory considerations to the actual act?
Don't justice systems almost universally have several levels of responsibility in relation to intent, which at its most basic level can be established by predictable outcomes?
If, for example, you are a leader of armed forces, and also a leader of organizations capable of creating propaganda. Let's say you create and distribute some propaganda (maybe using some AI tools), and a predictable outcome of that is that soldiers will be more lenient in their consideration of the rules of engagement and international law. In that case, one could at the very least establish that you were negligent in your creation and distribution of propaganda. The actual crime would have been the people killing people, namely your soldiers, but you would certainly be given some responsibility for that.
You can similarly take a small next step after that and consider that a company producing, distributing, and profiting from a dual use technology capable of creating propaganda and disinformation that can be responsible for crimes could be held at the very least morally accountable for those crimes, if not criminally.
Responsibility, accountability, moral and criminal, are not black and white notions. They are heaviest and easiest to attribute around physical acts of damage, but they stretch far and wide. To think otherwise is to allow the people with the most power to rampage unaccounted.
It depends on the basis form which you derive your (universal) moral values. Maximalist liberty as a universal moral value can be derived from the dual axioms of universal moral equality and a lack of moral oracle. If you accept these axioms, it follows that there is no source of moral authority that can legitimately constrain the non-infringing actions of another (eg. your right to wave your fists around ends where my nose begins). These ideas were first laid out in The Declaration of the Rights of Man, and expanded on in the Declaration of Independence.
> What are you actually trying to say?
That the causal chain of an action is completely interrupted at the first agent/actor in the system, who bears full responsibility for their actions.
> justice systems almost universally
It very much depends on the justice system. If you look at US/British/Roman law, a guilty mind (mens rea) and a guilty action (actus reus) are core facts that must be established in order to prove a crime has been committed. These still apply in cases of eg. criminal negligence, where a reasonable person ought to have known that their actions will result in harm. Mens rea is quite challenging to prove in cases of incitement, and legal precedents vary.
In combination with the above causal thesis, I hold that restricting incitement is in all cases an overstep of federal authority and an infringement of fundamental liberty. Incitement as a crime seems to have been established to make policing easier, not because telling someone to do something makes you responsible for their actions.
> you were negligent in your creation and distribution of propaganda
People are not inanimate objects. They are decision-making agents. The world is not a Rube-Goldberg machine. The soldiers who do the killing are responsible for their own moral attitude, and their own actions. You cannot be reasonably expected to know how your ideas will impact the minds of others, since every mind is a black box. Everything that contradicts this does so with generalizations too broad to be predictively useful.
> You can similarly take a small next step
This is where everything goes insane. Where does the responsibility end? You're trying to piece the butterfly effect back together.
Are people who make and sell bullets responsible for shootings? What about those that refine brass and lead? What about those that mine for ore? Creating economic demand, or promoting an idea, are morally neutral actions. People buying goods are in no way responsible for the conditions of their manufacture. People promoting ideas are in no way responsible for the actions a listener may take. Responsibility is zero-sum. Don't allow slavers and murderers to dispense with even a tiny portion of the sum responsibility for their actions. They must bear it all.
https://www.metmuseum.org/art/collection/search/436482
I suspect you will struggle. The economics for that sort of work don't exist anymore.
Speak for yourself.
It doesn’t look the opposite, it looks it automated even what we all couldn’t think of, and did that first.
My one hope is that the price of goods becomes so low due to AGI/automation, that the uselessness of labor in the economy won't matter. People can still be materially prosperous even with a meagre UBI (and it will be meagre because people have no political power in a post-labor society where the only thing that matters is capital).
Agreed. My concern isn’t really remotely about any of the accomplishments of generative AI. Frankly in my daily life I’d welcome readily available access. As it stands now it’s sort of a mixture of analytics and creativity without consciousness as we best understand it, so GPT itself isn’t going to murder me and take over the world.
The real issue is who owns these things, how you access them, how effects will ripple through a labor based economy, and how we’ll adapt (or not) our current economic system. As it stands for awhile we’ve been catering to the capital ownership group. If that doesn’t have a change in direction then I fear the implications of much of this in daily life. There’s still a fair bit of specialization and domain knowledge needed to leverage these tools to understand the questions to ask (I.e prompts to generate both around LLMs and the context of information fed to them) but they can certainly in many cases behave as multipliers that could reduce the amount of staff needed in some creative roles or eliminate some all together.
This isnt a new dilemma as arguably technology has been shifting the labor market for centuries, the question is how and if it can reshape well this time or if we need to fundamentally rethink these concepts of labor and capital ownership. That’s my major concern.
We're discussing a hypothetical post-labor future in 5-40 years. We probably shouldn't predict the economic theory of this future by looking at recent trends. Recent trends are driven by business-as-usual things like supply chain disruptions. But we're still near full employment, so we're not on the gradient to realized post-labor just yet. Post-labor economics (and politics) will probably be radically different, all the economic assumptions we take for granted go out the window.
Imho. it's just really hard to reason that average non-educational entertainment has a positive net effect on global society.
Seeing it this way makes it way less surprising that "art" and "creative entertainment" is one of the first things that gets hit by automation.
However, there's a line somewhere. I've spent most of my life around drab midwestern utilitarian/corporate/commercial buildings, and it has been noticeably depressing. In the periods where I've spent time in beautiful buildings, I have felt much better. Based on anecdata, I'm not the only one. There's something important & essential for humans about ornamentation & beauty. It's more than entertainment.
Humans can live on rice and kidney beans, but if one must do so without hope for more tasty options[0] it is demoralizing.
[0] lots of people are happy with spartan diets, but most often those people are doing so by choice.
H
Anecdote: My grandma retired and started painting and has since passed. The market value of these paintings is 0, nobody would buy them as they are just average. But I will never get rid of them because she created it. They have value to me only.
The concept of art as exclusive property is very new. Throughout history, artists have built on one another’s works with no attribution or provenance. It’s really just the past 100 years — Disney, specifically - that have created the cultural mindset that the first person to express something owns it forever and everyone else has to pay them for the privilege of building that next work.
BTW I’m old enough to remember people decrying the rise of desktop publishing (“WYSIWYG”) as the automation of creativity.
I share your disdain for the geriatric political class, but I strongly disagree that this is a situation that needs to be managed. I say we let the arts return to the free for all they were for the fist 80,000 years or whatever.
Great many, if you care to read a bit more of the biographies, autographies, history of music books, interviews, blogs, etc.
Is it because people are violating copyrights to train these AIs?
They are sure, however, that it is a kind of infringement. Citing "fair use" is an admission of infringement - just a specific kind of infringement that is allowed.
Certainly not for every individual idea, but good musicians do a lot of attribution. I got to know a lot of music I love now after following a mention on the liner notes of another musician’s album, or having them mentioned in an interview.
People seem to be asking for much more direct attribution: the pixels in this image are 0.02% from artist X, and 0.006% from artist Y, etc.
It is very rare for a song to include a breakdown of all of the influences that the artist is exercising in that particular piece.
also no, disney did not invent the notion of authorship nor royalties. having enough honor not to take credit for someone else's work goes back millennia. attribution goes back millennia, otherwise we wouldn't know the names Sophocles, Aeschylus, Euripides.
Don't get me started on the pharaohs, mother fuckers loved carving their names into things.
I welcome its fall.
Let's muse on the notion that creativity, as we've known and cherished it, can be bottled up and dispensed by machines, up to a certain whimsical point. Beyond that? We stumble upon creations like these, novel tools that beckon us, the flesh-and-blood creators, to mold unforeseen "creativities." It's one spectacle to mechanize the known realms of artistic endeavor, quite another to boldly claim that machines shall inherit the mantle of creativity, henceforth dictating the contours of all future artistic landscapes.
History, that grand tapestry, is peppered with instances where the mechanical muses have dared to tread upon the sacred grounds of creativity. Take photography, for instance, a marvel of the 19th century that promised to capture reality with an accuracy that scoffed at the painter's brush. Or consider the digital revolution, which flung open the doors to realms of visual and auditory experiences previously consigned to the realm of dreams. The synthesizer, not merely an instrument but a portal, has ushered us into a new era of musical exploration, challenging the supremacy of the acoustic tradition.
Each of these milestones, while distinctly modern, echoes the age-old dance between creator and tool, where each step forward is both a continuation and a departure from the past. In this light, the question isn't whether creativity can be automated, but rather how our definition of creativity evolves as we, hand in hand with our mechanical counterparts, stride into the unknown.
When video became an affordable medium, would people say "this is the end of art, live performances are art. Now the people will just watch the same recordings over and over?" Maybe, if the internet existed. But it's had the effect of creating and introducing new art forms.
AI generated content won't replace art. It will evolve it to a new creative.
a few examples
Plutarch's Lives
Holinshed's Chronicles
Ovid's Metamorphoses
good artist copy, great artist steal
so on and worth
i, for one, welcome these creative machines slurping all that was created!
So I don't see AI art as changing careers much. Even if AI fully replaces human artists, all that means is the 0.1% of people who make a career off their art will have to join the rest of us 99.9% who only do art for the fun of it.
[0] Unless you're doing furry art, but that's only because furries are "suspiciously wealthy".
I think this is less true than it's been in centuries or perhaps all of history. Artistry is widespread, anyone can do it, and many choose to pursue it even though the pay isn't going to be great; in preindustrial times even having access to the ability to create art was quite limited as were the media types that existed.
The main obstacle to this is the pride and ego of the people who've "made it" up until now. Let go. Let society have nice things, even if you have to reinvent yourself. I don't think that creativity is endangered; art, uh, finds a way.
Some twisted the story as if the underlying issue was the religion but economic concerns were the real reason.
May I introduce you to the entire history of humanity between 7 millennia before the invention of writing to approximately 50 years after the invention of communism? :P
More seriously: yes, we have no clue how to manage the situation. The best guess right now is UBI, which looks cool, but then at a first glance so does communism and laissez-faire capitalism.
Time for, ironically because humans are surprisingly bad at this, a creative idea for how to manage all this.
An AI model has no "unique self" to add to creation, at least not as we've understood so far.
I don't think this position will lead to good outcomes in terms of progress for civilization.
I'm not ideological about this, I wish for a future with self-driving cars for instance.
The current situation is simply too rapidly evolving and can cause significant economic destruction, for instance if many middle-class jobs are lost without anything to replace them.
Change is inevitable, but reckless speed is not, that's a choice we make as a group.
This is just a point in our overall evolution. It's an exciting time. We are here to learn and adapt.
Humans can still be creative all they want. There's still the stamp of "created by a human" that will never go away. You can choose to respect it or ignore it.
Nothing is forever. It’s unlikely unmodified Homo sapiens are dominant on earth 1,000 years from now.
It reminds me: centaurs (human+AI) in chess/go were better than either humans or AI just for a short time.
People still play chess but they are outclassed by modern AI.
I was having a conversation about this with a friend last weekend, and we'd assumed that centaurs were still better than either top humans or top computers. I'm unable to easily find this info on google, could you share where you saw that centaurs are no longer better than top computers?
To me, human activities from which we can earn a living wage feels like nomadism as the edge of an ever expanding region of agriculture (technological automation in this case). When you lose some activities to automation, we've always found new ones until now. In the end though, there were no more pastures for nomads to move, and there will be no more new activities from which humans can earn a wage (not to mention the satisfaction of accomplishing something hard). And, while there might be a future with UBI for everyone, the transition seems rough and exploitative.
It is the same as what every human being is doing. We consume and we create. Sometimes creations are very good, but most of the time they are just mediocre. If the machines can create better average results, it will be due to the genius of the humans who invented those machines.
So we can be happy, that we have such beings among us and should cherish, that we will have better content to consume in the future. When you look at the world, you will see, that there are still plenty of problems to be solved for humans.
OK, well, you walked right into this one:
You must know the answer: How do you manage it?
People invest in stories. They also invest in other people. This is why people love seeing Tom Cruise over and over again in movies. Or why I'm going to go see the next Scorsese movie.
Reality television is designed to be addicting, and engaging, and it is very successful at that. I get pulled into The storylines whenever I watch. But I quickly turn it off. I don't watch it not because it is not enjoyable, but because I realize it is a cotton candy: empty trash that is not worth my time.
Artists are already often criticized for being "corporate." I think we'll see a similar effect for AI generated content. The hoi polloi and normies will slurp it up.
The true fans and passionate ones who give a shit aren't going to be fooled.
Edit: for length
Jokes aside. It's becoming more apparent, Power will further concentrate to big tech firms.
this is the opposite of history
More importantly, how can these accounts subtly direct the generations to instill modern ideology or politics into "historical" images, giving them historical credibility? Think of all the subtly white supremacist "retvrn" accounts, for example, falsely recontextualizing inventions and accomplishments to support their ideology.
We all need to be thinking much more creatively and cynically about how these tools will be abused. The technology will get better. The people who want to abuse it will get smarter. And your capability to distinguish fake information is likely much worse than you believe - to say nothing of younger people who have less context and experience to form a mental "immune system".
I would say, all of them. Since the dawn of history. Actually, far before, as treachery certainly precedes speech itself by a few million years in the struggle to survive game.
Just to take a contemporary western (mostly?) thing: how did it went last time you looked straight into the eyes of kids to reveal them Santa Clauss is a lie and yes almost all adults in their society are into that evil conspiracy? And what about the adult around you deeply attached to their national myths, not even mentioning all the folklore around their afterlife beliefs?
But don’t worry, everything is going to go well, I promise and you know you can trust me. :)
Social media is really good at separating content from context, things like this will distort people's understanding of history.
if "historical" is going to be used subjectively with no further qualifying statements then the meaning of "history" will be subjucated to the context it's being presented in, I don't see it's use here as contradictory
The killer app for this is being able to give a prompt of a detailed description of a scene, with actor movements and all detail of environment, structure, furniture, etc. Add to that camera views/angles/movement specified in the prompt along with text for actors.
As I understand the current US situation, a straight prompt-to-generate-video cannot be copyrighted. https://www.copyright.gov/ai/ai_policy_guidance.pdf
But the copyright office is apparently considering the situation more thoroughly now.
Is that where it stands?
If it can’t be copyrighted, it seems that would tamper many uses.
*Edit* Oh, I just read here (https://www.reddit.com/r/MachineLearning/comments/1armmng/d_...) that a technical paper should be released later today?
My initial observation is that the camera moves are very similar to a camera in a 3D modeling program: on an inhuman dolly flying through space on a impossibly smooth path / bezier curve. Makes me wonder if there is actually a something like 3D simulation at the root here, or maybe a 3D unsupervised training loop, and they are somehow mapping persistent AI textures onto it?
As a layman watching the space, I didn't expect this level of quality for two or three more years. Pretty blown away, the puppies in the snow were really impressive.
As much as I love the technology, I’m really not looking forward to this becoming ubiquitous. Time and time again we’ve allowed technological progress to outpace our ability to weight the societal pros ands cons.
Smartphones and the rise of image-heavy social media has rapidly changed social norms. Watch a video of people out in public 20 years ago: no screen to distract them at bus stops, concert events, or while eating dinner with friends. And if that seems trite, consider how well correlated the rise in suicide rates is with the popularity of these technologies.
Not sure if this makes me a luddite or if the feeling is common in this crowd.
Even though many things are super impressive, there is a lot of uncanny valley happening here.
Pika - $55M
Synthesia - $156M
Stability AI - $173M
What OpenAI does is amazing, but they obviously cannot be allowed to capture the value of every piece of media ever created — it'll both tank the economy and basically halt all new creation if everything you create will be immediately financially weaponized against you, if everything you create goes immediately into the Machine that can spit out a billion variations, flood the market, and give you nothing in return.
It's the same complaint people have had with Google Search pushed to its logical conclusion: anything you create will be anonymized and absorbed. You put in the effort and money, OpenAI gets the reward.
Again, I like OpenAI overall. But everyone's got to be brought to the table on this somehow. I wish our government would be capable of giving realistic guidance and regulation on this.
But in reality it seems like the opposite is going to be true. AI is automating the creative, intellectual work and leaving the rest to us.
AI is cheaper than a high paid designer, developer, writer, etc.
A robot is more expensive than a human laborer.
It's really funny to see the squirm from those thinking truckers would be automated away, not them.
Not when intelligence is cheap and highly abundant. Perfecting general robotics as an improvement on humans will be quick. The upper limit of strength and consistency is much higher.
It is currently more expensive to build a robot for many tasks than it is to have a human do it.
> Perfecting general robotics as an improvement on humans will be quick.
It has not been nor is there any indication it will be.
Same with robotics. Lots of potential, but hasn't happened yet. If you read the description, Sora, is based out of trying to simulate the physical world to solve physics based problems. Something that would be perfect for the next leap in robotics.
There's no physical task that robots have replaced humans for me.
Hell, even the roomba sucks (pun intended) and my wife has to pick up the slack.
Less glibly, no matter how good you are at following instructions, tearing out a wall filled with water than can destroy your home, fiberglass insulation that can damage your lungs and electrical wiring that can kill you will never be something I’d recommend a layman do. No matter how good the ai tutorials are.
> We’re teaching AI to understand and simulate the physical world in motion, with the goal of training models that help people solve problems that require real-world interaction.
Text-to-video is just the flashy demo that everyone can understand after exposure to text-to-image. Once the model can "simulate the physical world in motion" it's only a few steps away from generic robotic control software that can automate a ton of processes that were impossible before.
Humans still have the benefit of dexterity and precise muscle control but in the vast majority of cases robots can overcome those limitations with better control software and specialized robotic end effectors. This won't soon replace someone crawling under a house or welding in awkward positions, but it could for example replace someone who flips burgers or does manual labwork.
This could eliminate the limiting factor for automating many manual processes. (ruh-roh)
If you use Sora like models to imagine what actions needed to be taken, then realize it, well, the only thing left is to create an arm/fingers that can took action, then you are done.
In general though, I don't think the extra reasoning ability is going to enable it to displace that much farther than it already will, GPT lives in a box and responds to prompts. When it's connected to multiple layers of real-time sensor data and self-directing, that'll be another story.
https://www.theinformation.com/articles/openai-shifts-ai-bat...
There were independent efforts to create AI agents since last year as well. AutoGPT and BabyAGI iirc. They didn’t go far probably because the LLM used was not good enough for that.
The rest of us have no choice. Despite millions of artists, animators, etc. all being resoundly opposed to AI art, the models that infringe their work are still allowed to exist, and it seems they're fighting a losing battle.
A lot of people are being "hysterical" because a lot of people don't have a choice.
To be clear, the problem of these scenarios is tightly intertwined with the problem of unfettered capitalism and wealth inequality in general. Food and shelter require money, and we get money by working a job. If millions of jobs disappear overnight, then of course millions of people are going to be distressed over no longer having ready access to food and shelter.
The idea of "just getting another job" doesn't scale to the destruction of entire industries employing tens of millions of people. This is how depressions are made.
The idea of "the depression will end someday" is not only not necessarily true as wealth inequality skyrockets, but is also cold comfort to the people who will lose their houses and for some, lives, due to the disruption.
A different economic system could perhaps allow us to appreciate these technological advances without worrying about them displacing our ability to live. But the American political system consistently and firmly rejects any ideas not rooted in social darwinist capitalism.
For your sake, I hope your resume is very impressive.
"Companies employing people" will be replaced by "people employing AI". Open models are free, small, fast, trainable and easy to use. They capture 90% of the value at 10% the cost, and are private.
What we're looking at is a massive decrease in the relative economic value of the average human's work. If the economic value of a hundred people is less than what the company can produce with a single human operator running AI models, then those 100 people are economically worthless, and don't get to eat.
We drastically need to tax the usage of AI models on the huge windfall they're about to create for their operators, and use that to fund universal basic income for those displaced. Generally speaking, as automation and wealth disparity skyrocket, UBI will be required to maintain any semblance of the society we currently have. I am incredibly pessimistic about the chances of that happening in any real way though.
It's already very tempting for large entertainment businesses to create lazy remakes as it involves less risk. Automating creative jobs will create a shift at the production level but also on the receiving end: the public.
Fully-functional autonomous driving will have a much larger economic impact - and that's just the first area where autonomous robots will come into our lives.
As of now, the models still need large amounts of human produced creative works for training. So you can imagine a story set in a world where large swathes of humanity are regulated to being basically gig workers for some quadrillion dollar AI megacorp where they sit around and wait to be prompted by the AI. "Draw a purple cat with pink stripes and a top hat" and then millions of freelance artists around the world start drawing a stupid picture of a cat because the model determined that it had insufficient training data to produce high quality results for the given prompt. And that's how everyone lives their lives....just working to feed the model but everything consumed is generated by the model. It's rather dystopian.
That will likely always be the case. Even 100% synthetic data has to come from somewhere. Great synopsis! Working for hire to feed a machine that regurgitates variations of the missing data sounds dystopian. But here we are, almost there.
1. Interacting with the randomness of the world
and
2. Thinking a lot, going in loops and thought loops and seeing what they discover.
I don't expect them to need humans forever.
Hallucinations are highly creative as well. But unless the technology changes, large language models will need human-made training substrate data for a long time to operate.
I've been trying to figure out how to retool the story to fit a timeline where ubiquitous AI that can write poems and paint pictures predates ubiquitous self-driving cars.
Unfortunately, LLMs aren't intelligent in any way, so you cannot ask them to synthesize any kind of second-order knowledge.
This is why they won't take away the creative work, either. They are fundamentally incapable of creating anything new.
Yes that’s my whole point…
I claim that the first part is the more difficult one and where we have the bottleneck currently. Furthermore, generative video AI is exactly the kind of thing that would give a model an understanding of what kinds of things have to happen in order to make coffee.
Indeed, there is a risk it completely devalues creative jobs. That's ironic. Even if you can still use AI creatively, it removes the pleasure of creating. Prompting feels like filling Excel sheets while also feeding a pachinko machine.
The question here is really about whether it's sufficiently transformative, or whether that's even the right standard to be applied to generated media.
I love tech, but if you take the stance that it's okay to hurt people for the sake of technical progress, you get into some very dark and terrible places...
That's a strawman. The real view is that protecting jobs that are made extinct by technology and automation is historically a bad idea because it leads to stagnation and poverty. It's better to let people lose their jobs, and for those people to find other jobs, while supporting them with a social safety net while they make the transition. Painful for them but unfortunately very necessary for a prosperous society.
This is the part that no one is expecting to see actually happen, though. Without that addressed, your argument is sound but footless.
What no one is asking is: 'it this makes it easy for anyone to be an artist, a director, a musician... what are we going to get, and will it be worse than what we have now?
Everyone is asking this.
But that's also not the only question. The one you're ignoring here is: If these tools enable one artist to do the work of a hundred, what happens to the other 99?
AI boosters have as yet offered no satisfactory answer for this question. Given the intimate involvement some of them have with politics at the national and global level, this absence constitutes reasonable grounds for suspicion that no answer is intended or forthcoming, and that suspicion is what's asking here to be addressed.
Not really -- as people have gotten more efficient at their jobs, we tend to just produce more/better things, not impoverish a bunch of people. If one person can day (8 hours) making a shoe by hand, and one person can make a shoe in an hour using a shoe making machine, then we don't have one less shoe maker, we have two people making 16 shoes a day. As an effect, shoes are now much cheaper, so they aren't only worn by rich people. If the one-shoe-per-day maker refuses to use a shoe making machine, he or she can upsell their 'hand crafted' shoes to rich people who want to distinguish themselves.
Believe me, I am not a 'free market fixes everything' person, at all, but in these cases, that is how it has worked since the industrial revolution. This is not a new process (automation making a task much more accessible/efficient) and this is not a new complaint (what happens to the people who made a living doing task).
Change is scary -- and everyone has the right to be afraid of an uncertain future, but I can't recall an instance of the regressive approach actually working to allay the fears of those who imposed it. Yet, we all see huge reminders of how our lives have been improved by making hard things easier and accessible to more people.
It would not surprise me if anyone called this pollyannaish, or even Panglossian.
Did anyone get harmed when photography was used to supplant portraits? Did anyone get harmed when mail started getting sent by rail instead of horse? Did anyone get harmed when air travel became possible? Did anyone get harmed when we supplied electric power to homes?
I have an idea -- why don't you propose a solution to AI ruining creative jobs and we can apply that standard to it.
Of course you may respond that this is unrealistic, which it is; it requires a government capable of acting via regulation in defense of its citizens, and so nothing like it will be done.
Maybe mass unemployment will create a sea change in that mentality, but most of the people who's opinions need to be changed will probably just laugh at "the elites" getting screwed over.
Hurt is a very subjective word in this context, how many people do you think the invention of the steam engine hurt? Or the electricity?
The upside is that creative works are completely democratized.
Now, anyone, with very little effort is fully empowered to create creative works on their own and there is no barrier to entry.
Yes, empowerment and democratization harms people who's livelyhood depends on disenfranchisement.
The current batch of LLMs is in the same class of technological revolutions as Napster and The Pirate Bay. Immensely impactful, sure, but mostly because of theft of value from elsewhere.
The factories that replaced the artisans were only made possible by the work of the artisans forging the way.
AI is nothing like anything we’ve seen, and is truly unique in the dangers it poses to the world.
South Korea had a high % of broadband penetration earlier than many Western countries, and as a result physical CD sales crashed very hard, and very quickly. So he asked himself, what's the most analog good I could sell? It's people. And went the pop idol / personality marketing route with great and lasting success.
Nonsense. Also, my point is that it shouldn't be up to tech companies to unilaterally decide what has value.
It's been this way for 10,000 years since the invention of the wheel. New inventions change how things are valued by making it easier for people do more work with less time.
Instead it is up to the consumers.
If consumers choose to give money to AI company, and not to artists, then in the eyes of the consumer those artists do not have value.
This is how technology works in general and should not be vilified. Someone comes up with a better way to do things (in this case bringing creative ideas to life) and charges a premium on top of that for their efforts. If the current wave of creators doesn't like it, then they should instead make something people want more than what their competition has to offer.
Either way, this is why local open source models are critical, so that everyone can benefit without needing to pay any single party.
IP law has yet to decide whether my interpretation of the situation is correct in the legal sense, but I find it impossible to see "ChatGPT absorbs the work of writers/journalists and sells a superficially reworded version without attribution or compensation" as anything but theft obfuscated behind lots of fancy math. It's only going to get worse if LLMs end up displacing traditional search engines, so one day you'll publish an article and get exactly one impression from GPTBot which then turns around and figuratively copies your homework.
This is an extremely different difference of scale, which does constitute a meaningful difference from prior technologies.
This is not to dismiss the concern. I simply wanted to state that artists will find ways to keep moving the creative bar forward.
[0] I really like this turn of phrase, thank you for sharing it.
That's not what extortion is. Stop abusing language.
You can deliver more content, faster, cheaper.
The quickest way to address this would be an extremely high tax rate on any generative AI model, say 500%, while the government figures out what’s the best way to sustain an economy (such as UBI) with a diminishing set of consumers as more people are pushed towards unemployment.
How are you going to stop me from doing that?
Even the free and open source stuff will destroy industries and you can't confiscate everyone's consumer gamer PCs.
Taxing the big guys doesn't save creative industries. It's a lost cause.
The only thing I can imagine is like limiting people’s compute power
But even then they’d just go do it in another country or use an online service based in another country
There isn't a way to "capture" the value from that.
Even if you aren't directly selling AI assets to someone else, people simply using AI themselves will still destroy industries.
Good luck confiscating everyone's graphics cards. The cat is is out of the box already.
> the tax would be levied on you
No it wouldn't. AI is already everywhere. Its game over. You aren't going to be able to track basically anyone who is using local models or other AI.
Oh man, how I miss it when ice was hauled from the Arctic in boats.
Training a massive model like this is a risk, and no one is going to take that risk without some reward. You can complain OpenAI is going to too much of the value, but its value that would have otherwise never existed. It's value.
Who's "you"?
Automate away the lower classes all you want, just don't touch the white collar class, that's a heckin' nono.
Can see this create an explosion of new Content from aspiring Film, Story tellers and cut scenes from Game creators that previously never would have the budget or capabilities to be able to see their ideas through to creation.
This is not plunder, it is empowerment. Blocking generative AI would be a huge power grab for copyright owners. They want to claim ideas and styles, and all their possible combinations.
Gen AI need only ensure it never reproduces a copyrighted work verbatim. Culture doesn't work if we stop ideas from moving freely.
Another issue to look at is the lack of ownership of the tools of your trade. In a context where many use AI models to competitively produce, hosts of AI models essentially own the access to your trade - thereby able to charge a toll, or privilege certain behaviors for any who strive to make living with these tools. (of course this is happening now with plenty of software products). The ultimate trajectory of this is not democratization of a toolset, but a transfer of wealth from labor to capital. And keep in mind that the labor share of income has been steadily declining for half a century.
The creation of wealth from AI ultimately depends on the strength of democratic and pluralistic institutions that safeguard ownership of your trade, democratized access to capital, and safeguards of welfare in the environment of creative destrcution. Otherwise you wind up with the cotton gin.
Yeah, this is why "AI creators" shouldn't be the ones unilaterally deciding how this all plays out.
Go blame your fellow consumers if you don't like the fact that they prefer AI.
These are choices that everyone makes. AI companies alone aren't forcing everyone to use their cool new tools. Instead, thats a decision that 10s of millions of people are making every day.
Copyright should have ended decades ago. It has accomplished nothing but harm.
(as pointed out in the "Freakonomics" episode highlighting this reaserch)
https://direct.mit.edu/rest/article-abstract/102/3/583/96779...
Means for us :(
That's one big trick, almost magical.
capture the value of every piece of media ever created
In what way does “I have a computer that can make movies” mean “I have captured the value of every piece of media ever created?” What do you mean by “value”? In my biased view, this amazing new technology couldn’t possibly be a better time to fix our insane notions of property, intellectual or otherwiseThere is no stop now. It's too late for that. Time to think about the full development and how we'll handle that. How we as people will be able to exist next to it. What our purpose in the world is supposed to be. What the purpose of "value" is. What the purpose of "economy" or "the market" is.
Exiting times.
Nope, we are headed towards deflation. Families that need only a single worker to support everyone, and even support extended family, and less time working overall.
The call for live music drastically shrank when it became trivial for any business or residence to play music on command.
Are you against automatic language translation? I can positively guarantee that the training data that they used to be able to create significantly better translation models was not authorized for that purpose.
The entire translator industry has been steadily shrinking ever since the invention of automatic language translation.
Etc etc etc.
There's obviously two aspects of this complex social issue right now.
1. Whether or not the usage of publicly available media as training data is legal/ethical.
2. Whether or not the output of these types of generative systems (even if they're trained on "ethical" training data) which may result in the displacement of many jobs is legal/ethical.
I'm neither for nor against AI (LLM, diffusion, video, etc), but if you are going to take a stance, then you have to be consistent in your view.
You don't get to cherry pick - I don't want to see you using chatGPT, copilot, stable diffusion, DALL-E, midjourney, sora, etc.
The future of these high-fidelity (but not perfect) generative AI systems is in realizing we're going to need "humans in the loop". This means designing to output human-manipulable data - perhaps models/skeletons/textures instead of whole output. Pixels are hard to manipulate directly!
As for entertainment, already we see people sick of CGI - will people really want to pay for AI-generated video?
Last weekend my 7 year old decided he wanted to make and sell a shirt with an image of a space cat shooting a laser gun. It took him like 1 minute to use free Dalle3 to make and choose an image. Then I showed him a website to remove the background. Then I showed him a tool to AI-upscale the image. Then we uploaded it to Amazon Merch, it got approved after a few hours, and now it's for sale on Amazon. It took us maybe 10 minutes of effort end-to-end. Involved no artists.
Funny enough, Amazon is full of AI-designed merch, there were like 7 pages of shirts with space cats with lasers.
I'm talking about, say, art for video games and actual movies.
Like what if any artist could make a whole movie by themself without needing millions of dollars or hundreds of people
Similar to how you used to need a huge studio full of equipment to record music and now someone in their bedroom with a DAW can do it
The entire industry is already turning out terrible shit, but doing it by wasting hundreds of thousands of actors, production teams, and studio dollars in order to churn out that nonsense.
Meanwhile, there are millions of latent storytellers, who, for whatever reason (but primarily: not born into extreme wealth and nepotistic connections) could never express their ideas in motion/cinema at such ambitious scales.
By putting this power in the hands of actually talented writers and storytellers, you create a completely new market of potentially incredible works of art.
the pursuit of mastery is at the essence of any craft.
The same will be for the FX artist and 3D artists etc. The level of their work will grow, they will spend less time on dull work and more on tinkering with tiny but more important things like ideas, emotions, art overall etc.
If you have a team of X people producing Y pieces, and now X people can produce 10Y pieces, everything is fine as long as the demand for pieces keeps up. But if your company really only needs Y pieces or really any amount less than 10Y then the easiest thing for a company to do is go, "We don't really need X people, let's fire some"
Getting fired, in America at least, means loss of healthcare, income, and if it persists long enough housing. Most people are terrified of being homeless, broke, and without access to medicine.
So the problem not in the AI but in demand...
It's cold comfort to someone getting fired to tell them "If demand had also increased 10 fold you wouldn't have to sleep on the street."
The actual living human being who has had their livelihood destroyed probably isn't any less scared of their fate because you cleverly tut at them and go, "In actuality the AI didn't do anything bad to you, it just created a glut of supply and the market demand didn't keep up."
This aren't compatible at scale. If productivity grows, there will be less people doing the job.
Another thing to think about is what the AI is designed to do. Without knowing the details, I would expect it to be trained to produce the 'most likely' output given the prompt. Consequently, I would think being inventive is against its design, and 'most likely' is effectively that same as 'average'.
If anyone's taking requests, could you do one that takes audio clips from podcasts and turns them into animations? Ideally via API rather than some PITA UI
Being able to keep the animation style between generations would be the key feature for that kind of use-case I imagine.
This video is pretty instructive: https://cdn.openai.com/sora/videos/amalfi-coast.mp4
It "eats" several people with the wall part of the way through the video, and the camera movements are odd. Strange camera movements, in response to most of the prompts, seems like the biggest problem. The model arbitrarily decides to change direction on a dime - even a drone wouldn't behave quite like that.
Think about animation, how a program can generate a sequence of a bouncing ball between two key frames. Think about what defines a video. The frames right? From there I can try to imagine.
This is the key. I have enough curiosity to want learn the stuff from the ground up, just as I did with other technologies. But man do I have the stamina today? Not so sure!!!
On the other hand, I don’t feel like you need to know how a compiler work, let alone the hardware architecture it targets, before you can go through your first hello world program or even build some useful software on top of frameworks/library treated at blackboxes. So "I have no idea what I’m doing" in this perspective is probably as old as CS/informatics.
Also, every single abstract is leaky, so often it's a difference between "I don't need to know how X works now", and "I can never find out how X works because it's simply not knowable".
It's comparatively easy to understand and it does cover everything from basic networks to LLMs and Diffusion models.
Getting to a point where realistically you're not able to know something deeply but then still use it is pretty frightening.
When I say deeply I don't necessarily mean that for every device you need to know about all of its atoms, but to have a pretty good framework for how the thing works deterministically, and how it can fail.
That was my point, exactly.
This now applies to most things in modern industrial society. We operate our daily lives at a crazy high level of abstraction. I think for a lot of us on HN, we "know too much about what we don't know", and that is ... overwhelming.
Funny enough, most people are actually able to operate at these higher levels of abstraction without worrying too much, because they don't know enough about what they don't know.
Opt for a sane option instead to get started, likely one of these: (Astro, SvelteKit or Remix).
Here is an inspirational story for you: https://news.ycombinator.com/item?id=39288139
https://cdn.openai.com/sora/videos/puppy-cloning.mp4
Perhaps there are particular aspects of our world that the human mind has evolved to hyperfocus on.
Will we figure out an easy way make these models match humans in those areas? Let's hope it takes some time.
For example this looks very much like something from a modern 3d engine:
The SUV video for example looks very much like something you'd see in a modern video game which probably makes sense because most videos with kind of perspective are going to be from video games.
I don't know how they would use game engines directly for training and fine tuning though. It would be far too labour intensive to render high quality scenes using a video game engine for every prompt.
> That’s why we believe that learning from real-world use is a critical component of creating and releasing increasingly safe AI systems over time.
"We believe safety relies on real-world use and that's why we will not be allowing real-world use until we have figured out safety."
someone really interested in control would want OpenAI or whatever centralized organization to be able to sift through the results for dangerous individuals -- part of this is making sure to stymie development of alternatives to that concept.
That'll fix it.
It's a nice cleansing benefit that comes with these really extraordinary tech achievement that should not be undervalued (after all it produces basically an endless amount of equally trained producers like the industry did in a - somehow malformed - way before).
Poster frames and commercials thrown at us all the time, consumed by our brains to a degree that we actually see a goal in producing more of them to act like a pro. The inflationary availability that comes with these tools seems a great help to leave some of this behind and draw a clearer line between it and actual content.
That said, Dall-E still produces enough colorful weirdness to not fall into that category at all.
Actually, thinking of this from the perspective of a start-up, it could be cool to instantly demonstrate a use-case of a product (with just a little light editing of a phone screen in post). We spent a lot of money on our product demo videos and now this would basically be free.
Training an embedding/LoRA on the product and using it with the base model, same as is done for image-generation models (video generation models usually often use very similar architecture to image generation models -- e.g., SVD is a Stable Diffusion 2.x family model with some tweaks.)
Now, you may not be be able to do this with Sora when OpenAI releases it as a public product, just like you can't with DALL-E. But that's a limitation of OpenAI's decisions around what to expose, not the underlying technology.
I’d imagine IRL no-tech experiences will be the new ‘escapes’ too.
Maybe I’m too idealistic about the importance of the human spirit/essence…whatever that actually is.
1. Why would Adrej Karpathy leave when he knows such an impressive breakthrough is in the pipeline?
2. Why hasn't Ilya Stuskever spoken about this?
Would we be able to perceive the differences between those and the physical world? I can't help but feel like there is a proof for the simulation theory possible here.
Contrary to the trends in SV, dehumanization of creative professions will result not in productivity boost but in utter chaos and as a result will add more time loss in production process.
I never liked Sam Altman in his Y years, now I know why.
Even with the "blessings" from the "masters" in Davos/Bilderberg, a bad idea is a bad idea. Maybe this will push World ID as a result, but is it necessary?
The current trends in tech are not producing solutions for a professional problem. With rare exceptions, this looks more and more as removing of human input and normalization of a society ruled by AI at any cost. So sad.
I've always been a digital stills guy, and dabbled in video.. as a hobby. As a hobbyist, I always found the hardest thing is making something worth looking at. I don't see AI displacing the pleasure of the art for a hobbyist.
My next guess is the 80/20% or 95/5% problem is gonna be stuff like dialogue matching audio and mouth/face motion.
I do see this kind of stuff killing the stock images / media illustrator / b-roll footage / etc jobs.
Could a content mill pump out plausibly decent Netflix video series given this tool and a couple half decent writers.. maybe? Then again it may be the perpetual "5 years away". There's a wide gap between generating filler content & producing something people choose to watch willingly for entertainment.
The prompts - incredible and such quality - amazing. "Prompt: An extreme close-up of an gray-haired man with a beard in his 60s, he is deep in thought pondering the history of the universe as he sits at a cafe in Paris, his eyes focus on people offscreen as they walk as he sits mostly motionless, he is dressed in a wool coat suit coat with a button-down shirt , he wears a brown beret and glasses and has a very professorial appearance, and the end he offers a subtle closed-mouth smile as if he found the answer to the mystery of life, the lighting is very cinematic with the golden light and the Parisian streets and city in the background, depth of field, cinematic 35mm film."
- This enables everyone to be creators
- Given that everyone's creativity isn't top notch, highest quality will be limited to a the best
- So rest of us will be consumers
- How will we consume if we don't have work and there is no UBI?
It's doubly amazing when you think that the richness of video data is almost infinitely more than text, and require no human made data.
The next step is to combine LLM with this, not for multimodal, but to team up together to make a 'reality model' that can work together to make a shared understanding?
I called LLMs 'language induced reality model' in the past. Then this is 'video induced reality model', which is far better at modeling reality than just language, as humans have testified.
Two things are interesting:
- No audio -- that must have been hard to add, or else it would have been there.
- Spelling is still probably hard to do (the familiar DallE problem)... e.g. a video showing a car driving past a billboard with specified text.
or we are heading towards a skynet-y feature
The monetary value of generic stock content will surely drop and won't be created by professionals anymore. However, that doesn't mean people stop taking pictures of their dog just because they can get midjourney to generate the same thing. Creation for the sake of creation will continue. AI companies will initially reap in a lot of the $ value that used to go to the creators of stock content, but when open source models reach parity the masses will be able to make what's in their mind a reality as casual creators. Hobbyists will still exist and those that become truly great will still rise to notoriety.
You just write the rules of the game and the player input, and let the AI generate the next frame.
Except for live sporting events.
This is why I think megacorps all going to bid for sport league streaming right. That's the only one that AI can't touch.
Yes, I'm still bitter about that.
But I have a problem: I am unable to believe the videos I saw were dreamt by AI. I can feel deeply that I do believe there is some trickery or severe embellishment. If I am wrong, I guess we are at an inflexion point.
I can recall 10+ years ago, we were talking "in hacking groups" about AI because we thought the human brain alone was not good enough anymore... but in a maths/sciences context.
This is neat and all but mostly just a toy. Everything I've seen has me convinced either we are optimizing the wrong loss functions or the architectures we have today are fundamentally limited. This should be understood for what it is and not for what people want it to be.
Wider-Scale coherence is still much better than previous models and has consistently been improving. It's not "visual sharpness at the expense of coherence". At worst, the models are learning wider-scale coherence slower.
Not everything is equally difficult to learn so it follows that some aspects will lag behind others. If coherence weren't improving you might have a point but it is so...
Your vision is hundreds of millions of years in the making.
Its like we keep moving the bar
I have no idea, just guessing...
> The team behind the technology, including the researchers Tim Brooks and Bill Peebles, chose the name because it “evokes the idea of limitless creative potential.”
I wonder if it could be used as a replacement for optical flow to create slow motion videos out of normal speed ones.
I guess what I'm wondering is how "new" the videos are, or how closely do they mimic a particular video in the training set? Will we generate compelling and novel works of art with this, or is this just a very round-about way of re-implementing the YouTube search bar?
Also interesting that some of the examples ignore details in the prompts. No clouds or sun in the sky, no depth of field, their hair isn't blowing in the wind.
We're nowhere near full-automation, these are growing pains, but maybe the canary in the goldmine for the job market. Expect more enthusiasm for UBI or negative tax and the like and policies to follow. Cheap energy is also coming eventually, just slower.
Like you can see some weird artifacts, but take one of these videos, compress it down to a much lower quality and with the loss of quality you might not be able to tell the difference based on these examples. Any artifacts would likely be gone.
Given what I had seen on social media I had figured anything remotely real was a few years away, but I guess not...
I guess we have just stopped worrying about the impact of these tools?
One wonders how you might gain a representation of physics learned in the model. Perhaps multimodal inputs with rendered objects; physics simulations?
Video 7 of 8 on the 2nd player on the page.
> Prompt: The camera rotates around a large stack of vintage televisions all showing different programs — 1950s sci-fi movies, horror movies, news, static, a 1970s sitcom, etc, set inside a large New York museum gallery.
Let's temper the emotions for a second. Sora is great, but it's not ready for prime time. Many people are working on this problem that haven't shared their results yet. The speed of refinement is what's more interesting to me.
All of the AI generated media has this quality where I can immediately tell that it's ai, and that becomes my dominant thought. I see these things on social media and think "oh, another ai pic" and keep scrolling. I've yet to be confused about whether something is ai generated or real for more than several seconds.
Consistency and continuity still seem to be a major issues. It would be very difficult to tell a story using Sora because details and the overall style would change from scene to scene. This is also true of the newest image models.
Many people think that Sora is the second coming, and I hope it turns out to have a major impact on all of our lives. But right now it's looking to have about the same impact that DALL-E has had so far.
How did you rule out survivorship bias?
The bottleneck of creating a separate prompt is very limiting.
Imagine asking for a recipe or car repair and it makes a video of the exact steps. Or if you could upload a video and ask it to make a new ending.
That’s what I imagine multi modal models would be.
https://trends.google.com/trends/explore?date=all&q=disrupt&...
Want to form a trade union I'm your workplace? Best be ready to have videos of you jacking off to be all over the internet.
Videotape a police officer brutalising someone? Could easily have been made with AI, not admissable.
These things will ruin the ability to trust anything online.
You give it data like real time stock data, feed it into Sora, the prompt is "I need a chart based on the data, show me different time ranges"
As you move the cursor, it feeds into sora again, generating the next frame in real time.
- Local/Bespoke high quality video content creation by ordinary Joes: Check. - Ordinary joes making fake porn videos for money: Check. - Reduce cost for real movies dramatically by editing in AI scenes: Check.
A whole industry will get upturned.
I am curious of how optimised their approach is and what hardware you would need to analyse videos at reasonable speed.
To me, it's becoming increasingly obvious that startups whose defensibility hinges on "hoping OpenAI doesn't do this" are probably not very enduring ones.
I also feel a sense of dread too. Imagine the tidal wave of rubbish coming our way. First text, then images and now video can be spewed out in industrial quantities. Will it lead to a better culture? In theory it could, in practice I just feel like we'll be deluged with exponentially more mediocre "content" .
I can see a new market for true end-to-end analogue film productions emerging for people who like film.
I vote for Hothouse, by Brian W Aldiss. So many images need to imagined, like spiders that jump to the moon and back again.
"Sora" is not a video generation technology offered by OpenAI. As of my last update in April 2023, OpenAI provides access to various AI technologies, including GPT (Generative Pre-trained Transformer) for text generation and DALL·E for image generation. For video generation or enhancement, there might be other technologies or platforms available, but "Sora" as a specific product related to OpenAI or video generation does not exist in the information I have.
If you're interested in AI technologies for video generation or any other AI-related inquiries, I'd be happy to provide information or help with what's currently available!
I hope there is at least some cherrypicking here. This also seems like some shots fired at some of the other gen video startups
Sora: plays GTA V
And it is still not perfect. Looking at the example of the plastic chair being dug up in the desert[1] is frankly a bit... funky. But imagine in 5 or even 10 years.
Also, nicely timed to overshadow the Google Gemini 1.5 announcement.
I guess we've all just been replaced.
The amount of power needed to generate this can't be feasible for real time VR today. There's a reason even the company that invented (massive and free) Gmail is charging for its top tier generative AI.
No chance to think "sucks for you, but I'm good here" like so often happens with other issues.
On the other hand, people find the tech very impressive and there are a lot of mind blowing use-cases.
Personally, this opens up the world for me to create video ads for software projects I create, since I have no financial resources or time to actually make videos, I only know how to code. So I find it pretty exciting. It's great for solo entrepreneurs.
The prompt tho. Probably not text. Probably a stream of vibes or something.
There should be an opt out from being subjected to AI content.
This product looks incredible...
Creating these video's in CGI is a profession that can make you serious money.
Until today.
What a leap.
It's still too easy to notice these are all AI rendered.
Looks like a dramatic improvement in video generation but still a miss in terms of realism unless one can apply pose control to the generated videos.
Genuine question I have no idea
The fear I have has less to do with these taking jobs, but in that eventually this is just going to be used by a foreign actor and no one is going to know what is real anymore. This already exists in new stories, now imagine that with actual AI videos that are near indistinguishable from reality. It could get really bad. Have an insane conspiracy theory? Well, now you can have your belief validated by a completely fictional AI generated video that even the most trained eyes have trouble debunking.
The jobs thing is also a concern, because if you have a bunch of idle hands that suddenly aren't sure what to believe or just believe lies, it can quickly turn into mass political violence. Don't be naive to think this isn't already being thought of by various national security services and militaries. We're already on the precipice of it, this could eventually be a good shove down the hill.
I guess it was anticipated.
Now the big question is. As OpenAI keeps pushing boundaries, it's fascinating to see the emergence of tools like Sora AI, capable of creating incredibly lifelike videos. But with this innovation comes a set of concerns we can't ignore.
So i'm worried about getting these tools misused. I'm thinking about what impact could they have on the trustworthiness of visual media, especially in an era plagued by fake news and misinformation? And what about the ethical considerations surrounding the creation and dissemination of content that looks real but isn't?
And, what we should do to tackle these potential issues? Should there be rules or guidelines to govern the use of such tools, and if so, how can we make sure they're effective?
Its why I submitted this. We need some way to attest the authenticity of images.
So i'm worried about getting these tools misused. I'm thinking about what impact could they have on the trustworthiness of visual media, especially in an era plagued by fake news and misinformation? And what about the ethical considerations surrounding the creation and dissemination of content that looks real but isn't?
Why would it?
And the AI is optimizing the video feed purely for that
What would it generate?
You seriously can’t think of a single practical application for generating arbitrary media content on demand?
Brother, have you seen Runway Gen 2, or SVD 1.1? I'm not excited about Sora because I think it looks like Hollywood animations, I'm excited because an open-source 3rd-Gen Sora is going to be so much better, and this much progression in one step is really exciting!
The people who already hate Biden, probably already think he's doing some weird shady stuff, and would point to some conspiracy. The people who like Biden, would accept the alibi.
Ultimately it wouldn't move the needle.
What is concerning, is the technology being used against a regular person, who may not have an alibi.
"porn without consent" - thought crime
"too much porn of whatever you dream of" - yes, conservatives (50% of USA) actually think this is a problem
"spam" - advancing the closed garden model email is heading towards. soon you will simply need government id to make email even though there are plenty of alternative ways to do communication aside from email which was already considered insecure and a bad protocol in 2000. this has nothing to do with AI but they are still acknowledging this absurdity by framing AI as the enabler of that.
"automated social engineering" - just weaponizing the ignorance the bad thought leaders of the industry left us. instead of giving us proper authentication methods, we still have "just send my photo id to these 33 companies, which will ask for it in random ways we dont expect and just have to trust them"
"copyright" - literally not a problem, almost nothing "protected" by copyright matters and the law is just used by aggressive capitalists to shove their products down everyone's throat
"ICBMs being automatically hacked and launched at people" - just stop being bad government and hiring completely uncredible people to implement every mission critical control system while hooking it up to the internet
"racist bias" (or whatever) - this is the dumbest fucking thing i've ever heard of
this website is a perfect snapshot of why tech sucks so hard. its dressed up like cinematic film using a ton of js libs and css hacks or god knows so it can only be viewed smoothly on the latest computer hardware. only on one of the big 3 browsers that each had a trillion man hours of pointless iterations driven by digital graphics marketing companies. and on top of that they have a nice professional tone made by $300K/year PR people. please, sincerely, fuck off.
AGI can’t be far off, that stuff clearly understand a bunch of high level concepts.
Phase 1 (we are here now): While generative AI is not good enough to directly produce parts of the final product, it can already be used to quickly prototype styles, stories, designs, moods, etc. A good chunk of the unnamed behind-the-scenes-people will loose their job.
Phase 2: While generative AI is still expensive, the output quality is sufficient to directly produce parts of / the entire final product. Big production outlets will use it to produce AAA titles and blockbusters. Even actors, directors and other high publicity positions will be replaced.
Phase 3: The production cost will sink further until it becomes attainable by smaller studios and indie productions. The already fierce markets will be completely flooded with more and more quantity over quality. Advertisement will not be pre-produced and cut into videos anymore but become very subtle product placements, impossible for ad-blockers to strip from the product.
Phase 4: Once the production cost falls below the price of one copy of the product, we will get completely customized entertainment products tailored to our personal taste. Online communities will emerge which craft skeletons / templates which then are filled out by the personal parameter sets of the consumers. That way you can still share the experience with friends even though everybody experiences a different variation.
Phase 5: As consumers do not hit any production limits any more (e.g. binge watch their favorite series ad infinitum) and the product becomes optimized to be maximally addictive by measuring their reaction to it, it will become impossible for most human beings to resist. The entertainment mania will reach its peak and social isolation, health issues and economic factors will bring down the human reproduction rate to basically zero.
Phase 6: Human civilization collapsed within one or two generations and the only survivors will be extremely technology-adverse people by selection. AGI might have happened in the meantime but did not have the time to gracefully take over and remodel the human infrastructure to become self sufficient. Instead a strict religion will rule the lands and the dark ages begin anew.
Note that none of this is new, it is just the continuation and intensification of already existing trends. This is also not AGI doomerism as it does not involve a malicious AGI gone rouge or anything like that. It is simply what happens when human nature meets powerful technology.
TLDR: While I love the technology I can only see very negative long-term outcomes from this.
As several others have pointed out, realism of these models will continue to improve, and will soon be economically useful for producing beautiful or functional artifacts - however prompt adherence (getting what you want or intend) of the models is growing much more slowly.
However I think we have a long ways to go before we'll see a decent "AI Film" that tells a compelling story - and this has nothing to do with some sort of naturalistic fallacy that appeals to some innate nature of humans!
It comes down to the dataset and the limits of human creators in their ability to communicate their process. Image-Text and Video-Text pairs are mostly labeled by semi-skilled humans who describe what they see in detail. They are, for the most part, very good at capturing the obvious salient features of an image or a video. "reflections of the neon lights glisten in the sidewalk". However, what you see in a movie scene is the sum total of dozens if not hundreds of influences, large and subtle. Choices made by the actors, camera operators, lighting designers, sound designers, costuming, makeup, editors, etc... Most people are not trained to recognize these choices at all, or might not even be aware that there are choices to make. We (simply) see "Joaqin Phoenix is making awkward small-talk in the elevator with other office workers".
So much of what we experience processes on subconscious and emotional and purely sensory levels, we don't elevate those lower-level qualia to our higher-brain's awareness and label them with vocabulary without intentional training (such as tasting wine, coffee, beer, etc - developing a palate is an act of sensory-vocabulary alignment).
However, despite not raising these things to our intentional awareness, it has an influence on us -- often the desired impact of the person who made that choice in the first place. The overall effect of all of these intentional choices makes things 'feel right'.
There's no fundamental reason AI can't produce an output that has the same effect as those choices, however finding each little choice is like a needle in a haystack. Accurate labeling of the training data tells the AI where to look -- but the people labeling the data are probably not well-versed in all of the little intentional choices that can be made when creating a piece of video-media.
Beyond the issue of the labeling folks being trained in the art itself, there's the problem too of the artists themselves not being able to fully articulate their (numerous, little, snowflake-into-avalanche) choices - or simply not articulating it even if they could. Ask Jackson Pollock about paint viscosity and you'll learn a great deal, but ask about abstract painting composition and there's this ineffable gap that language seems ill-suited to cross. The painter paints what they feel, and they hope that feeling is conveyed to the viewer - but you'd be hard pressed to recreate "Autumn Rhythm (Number 30)" if you had to transmit the information via language and hope they interpreted it correctly. Art is simultaneously vague and specific!
So, to sum up the problem of conveying your intent to the model:
- The training data labels capture obvious or salient features, but not choices only visible to the trained eye
- The material itself is created by human artists who might not even be able to explain all of their choices in words
- You the prompter might not have the vocabulary that captures succinctly and specifically the intended effect
- The end result will necessarily be not quite what you imagined in your mind's eye as a result of all of this missing information
You can still get good results if you tell it to copy something, because the label "Tarantino" captures a lot of detail, even all the little things you and the training data would never have labeled in words. But it won't be yours and - until we have an army of trained artists providing precise descriptions for training data in their area of expertise, and you know how to speak those artists' language - it can't be yours.
We still need nurses, cooks, theater, builders etc.
People to just create lifelike videos of anything they can put their mind to, is bound to lead to the ruining of many peoples' lives.
As many people that are aware and interested in this technology, there is 100x people who have no idea, don't care or can't comprehend it. Those are the people that I fear for. Grab a few pictures of the grandkids off of facebook, and now they have a realistic ransom video to send.
Am i being hyperbolic? I don't think so. Anything made by humans can be broken. And once its broken and out there, good luck.
That's why this technology should not exist.
there’s never been greater distrust of legacy media, and the fact that you can’t trust everything you read on the internet has been a trope for decades
Yeah, that's a problem. Successful societies are built on trust, shared reality and communication. Democracy is a conversation.
The big problem with technology is you can't uninvent technologies that turn out to be net bad. It becomes a perpetual curse once it's invented.
You can’t uninvent something
Sure, now let's talk about knife wounding and acid attacks...
The fundamental issue of human violence still exists.
If we stop developing our tech, China will continue advancing even without leaks from OpenAI, and eventually they will develop better models, even if it takes another decade. Banning something only to cower in fear of ourselves while we watch someone else use it to destroy us seems like poor planning.
Yes, it's a great technical achievement, but I just worry for the future. We don't have good social safety nets, and we aren't close to UBI. It's difficult for me to see that happen unless something drastic changes.
I'm also afraid of one company just having so much power. How does anyone compete?
My fear is the alternative reality that these tools could provide. Given the power and output of the tooling, I could see a future where the "normal" of a society is strategically changed.
For example, many younger generations aren't getting a license at 16. This is for a variety of reasons: you connect with friends online, malls cost money, less walkable spaces, less third places.
If I'm a company that makes money based off of subscription services to my tools, wouldn't it be in my best interest to influence each coming generation?
Making friends and interacting with people is hard, but with our tooling you can find or create the exact friend you want and need.
We can remember now that life is beautiful - but what's to stop from making people think that the life made by AI is most beautiful?
And yeah, I've heard this argument before with video games, escapism, etc. I'm talking more about how easy it is to escape now, and how easy it'd be to spread the idea that escapism is better than what is around you.
In Europe there's no need. Got a licence over two decades ago have never needed to drive. Shops in walking distance, public transport anywhere in the country, convenient deliveries, walkable and cyclable cities.
Meanwhile other places have no freedom from cars, locked into expensive car financing, unable to access basic amenities without a car, and motorists have normalised killing millions of people a year.
The licenses thing is a proxy for what we’re measuring, which is real life socialization.
You’ll see the same trends in Japan, where far fewer people drive than in America. I’d imagine you’d see them in Europe too.
For the most part, the idea of change is rarely inherently bad (even though, IMO, it's natural to inherently resist it) -- and humans adapt quickly to the parts that have negative impacts.
Humans are one of, if not the most, resilient race on the planet. Younger generations not getting licenses, sticking to themselves more, escaping in different ways, etc are all "different" than what we're used to, but to that younger generation it's just a new normal for them.
One day they'll be posting on HN2, wondering whether the crazy technological or societal changes about to come out will mean the downfall of their children (or children's children), and the answer will still be the same: no, but what's "normal" for humankind will continue to change.
As long as they keep having unprotected sex with each other.
Otherwise, you know, humanity is kind of screwed.
Btw, a year or two from now you'll be able to run a more powerful open model locally. So, not they aren't having some outsized amount of power
Yes some people will be scammed as they always have been, such as the recent Hong Kong financial deepfake. But no, millions of people will not keep falling for this. Just like the classic 419 advanced free fraud, it will hit a very small percentage of people.
Invest in web 3.0 now.
IMHO humanity will be fine, decades from now kids will be asking what it was like to live before "AI" like how we might ask an old person what it was like to live before television or electricity.
I'm not optimistic.
Every large animation studio has continually been looking for ways to decrease the number of artists required to produce a film, since the beginning of the field.
You could be the best at using the tools, but I think there could be a point where there is no need to hire because the tools are just that good.
There is zero chance this tech is going to be locked up by a few companies, in a year or two open models will have similar capabilities, I have no idea what this world looks like but I think it’s less of a concern for individuals and more of a concern for the global economy in the short term.
Outside of all of this, yeah we’re either going I have to adapt or die.
Alone, maybe I will be able to launch a unicorn in 2030. It’s just tools with more leverage. The limit is just the computing resources we have, so we’ll have to use computing resources to calculate how much earth resources each of us can use per year, but that seems a usual growth problem.
If AI dystopia is coming, at least it's not here quite yet, so I'll try to enjoy my life today.
I'm seeking lasting meaning; not 'meaning' that dissolves after a season, or at best, at the end of a life.
What I meant by 'accelerat[ing] the realization' is that all of our earthly desires will more readily be fulfilled, and we will see that we still feel empty. AI is like enabling a new cheat code in the game of life, and when you have unlimited ammo the FPS becomes really fun for a moment but then loses its meaning quickly.
Of course you do. Your worldview depends on it.
> I'm seeking lasting meaning
I can get that too. We're the arbiters of meaning in our lives.
> earthly desires will more readily be fulfilled, and we will see that we still feel empty
You misunderstand if you believe that the secular perspective on meaning is to reach it through sensualism, by consuming. That is all that AI can more readily fulfill.
Nothing about the future looks particularly good, other than that medicine is improving. But what’s the point of being alive in such a sanitised, ‘perfect’, instant-dopamine-hits-on-demand kind of world anyway?
Just say to hell with it and bury yourself in an interesting textbook. Learn something that inspires you. It doesn’t matter if ‘AI’ can (or soon will be able to) do it a billion times better than you.
And be kind to those around you.
I've started reading again, because reddit/instagram/etc. has become kind of boring for me? Like, I still go on them to get an instant dopamine hit from time to time, but like you said burying yourself in a textbook just feels so much more rewarding.
I’m sure such groups already exist, but maybe not specifically with this goal in mind.
Learning for its own sake really is the answer to lasting happiness… for some of us, anyway.
Since 529 CE!
If you’re just talking about the idea of becoming a monk: yes, I very much like the idea of becoming a modern, digitally-enabled monk.
But none of these feelings are new, just different problems manifesting the same.
No mountains here + no garden for me, Yosemite is thousands of kilometers away. The sea has no waves. It's -5 degrees.
Which is "fine", but it means that indoor activities are more prevalent, including computers.
And I'm sure for some other people their options are way more restricted.
The key though is to avoid becoming a cult.
I doubt much of what we know today will turn out to be wrong. Maybe our abstractions will turn out to have been naive or suboptimal, but at least they're demonstrably predictive. They're not just quackery or mysticism.
You did say ‘wrong’, though, not ‘considered wrong’.
You have 3 baskets with 5 apples each, you remove 7 apples from each basket, remove 5 baskets and you have -2 baskets with -2 apples each thus therefore you have 4 apples left all without the involvement of trees, like Jezus!
Yeah, i can see people laugh at that.
It's probably the same feeling farmers had in the beginning of the 20th century when they started seeing industrialized farming technologies (tractors, etc). Sure, farming tech eliminated tons of farming jobs, but they have been replaced by other types of jobs in the cities.
It's the same thing with AI. Some will lose their jobs, but only to find different types of jobs that AI can't do.
The problem is the transition into that new world.
I can't imagine anything changing our culture's insistence that personal responsibility in employment means zero responsibility for employers, policy makers, or society at large. That is, short of a large scale armed rebellion, or maybe mass unionization.
don't worry; AI drones will deal efficiently with both those forms of terrorism and malinformation
Bear in mind that a substantial portion of people (perhaps 30%) don't feel satisfied unless they see someone else worse off. We are not an inherently egalitarian species.
> jobs that AI can't do.
Sure, by definition, you’ve described the set of jobs that won’t be replaced by AI. But naming a few would be a lot more useful of a comment. It’s not impossible to imagine that that set might shrink to being pretty much empty within the next ten years.
No but it’s also not impossible to imagine the opposite. AI beat humans at chess decades ago but there are more humans generating income from chess today than there were before Deep Blue.
Chess players get paid because it’s entertaining for others to watch.
So your argument only shows that we can expect work as a form of entertainment to survive. Outside of YouTube, where programmers and musicians and such can make a living by streaming their work live, this is a minuscule minority.
The strongest interpretation of what you’re saying seems to be that we’ll end up in a world where everything (science, engineering, writing, design) is a sport and none of it really matters because ultimately it’s ‘just a game’. Maybe so… but is that really something to look forward to?
Now upgrade AI to do every job better than humans so that there are no drudgery jobs. What money are they going to spend?
Not too long ago, people would come and visit the first family in the village who had installed running water, because it was a new and exciting thing to see. And yet people don't wake up every day excited to see water coming from their kitchen tap.
Imagine a breakthrough not only in AI but also robotics, allowing restaurants to replace the entire staff (chefs, cooks, waiters, etc) with AI-powered robots. Then I believe that higher-end restaurants will STILL be employing humans, as it will be perceived as more expensive, more sophisticated, therefore worth a premium price. What if robot cooks cook better and faster than human cooks? Then higher-end restaurants will probably have human cooks supervising robot cooks to correct their occasional errors, thereby still providing a service superior to cheaper restaurants using robot cooks only.
Not remotely comparable. Farming is a backbreaking job, many were happy to see it going away. This is taking over the creative functions. Turns out what Humanity is best at, is menial labor?
So the downside is we have lives devoid of meaning. The upside is we live in a scarcity free paradise where all diseases have been cured by superhuman ai and we can all live doing whatever we want.
Anyhow, what makes you think the AI or whomever controls it has any use for a bunch of useless eaters?
But when one is 30+ years old, or even 40+ years old, it's hard to completely switch careers, especially when you're also dealing with the fact that it's not because you were bad at your job. Rather, a machine was made to replace you and you simply can't compete with a machine.
It's evolution, of course, but it is a stressful process.
Nearly all humans work for money (aka, just basic stuff) and not because they’re passionate about their work. It is just a sad situation all around
I doubt whether this is true. Lots of hype, but no tangible improvement to show for chronic conditions for common people.
The systems that deliver the medical care might be, however (and indeed observably locally are, in many cases).
You are mistaken. To realize that you will have to look back several decades and read the literature of those times, of what is left. Now note I'm taking about chronic illnesses (diabetes, cancer etc) not acute ones like an infection etc. The medical practitioner of yesteryear did not have the fancy diagnostic tools that we have today, but several of them appear to me to be sharp observers.
If you make claims that bold no one should even bother to read on.
The actual issue, which is the only worthwhile thing you wrote about, is cost and availability.
The vaccines were not the result of medicine getting "better" - they just happened to have a solution for the right thing at the right time, which is fortunate (and we're lucky that it worked, because there was no guarantee of that beforehand) but if the pandemic hadn't happened, what advances would we be discussing? What advances are actually making medicine better aside from once-in-a-hundred-year worldwide emergencies?
- fast acting analog insulins that are metabolized in 2-3 hours instead of 6-7
- insulin pumps that automatically dose exactly the right proportion of insulin
- continuous glucose monitoring system that lets you see your BG update in real time (before, it was finger sticks 4-5 times a day; before that, urine test strips where you pee on a stick to get a 6 hours delayed reading (!))
- automated dosing algorithms that can automatically correct BG to bring it into range
In aggregate, these amount to what is closer than not to a functional cure for type 1 diabetes. 100 years ago, this was a fatal condition.
The world is way bigger than technology and the Internet. It hasn’t really gone anywhere
Fresh air and sunlight are important though.
And that sand takes a very, very long time with lots of big brains to figure out how to manipulate at the nanometer level in order to give you a "beep boop"
It's not like Intel could decide tomorrow to spin up a fab and immediately make NVIDIA and TSMC irrelevant. They're the next closest thing given they make chips, have GPU technology, and also foundry experience and it's still multiple years of effort if they chose that direction.
Your statement is a lot like saying "poker has predictable odds" and yet there is still a vast ocean of poker players.
But AI is different than previous waves, like search engines and social networks. You can download a model on a stick. You can run it on a CPU or GPU, even a phone. These models are easy to work with, directly in natural language, easy to fine-tune, faster, cheaper, and private under your control. AI is a decentralizing technology, will empower everyone directly, it's like open source and Linux in that it puts users in control.
on the flip side, we'll just generate VR wilderness in the near future and nobody will care what's real or not
i thought the point of tv was to sit back and be entertained, usually through some form of storytelling. personally, i don't want to have any part in the creation. if anything custom content would be annoying, because i'd lose the only social aspect of tv (discussing with others)
Just another tool.
Until full automated agent that is able to carry out a task from start to finish without human intervention, there is something for us to do I guess.
You are not alone : https://youtu.be/h3-va0umXTY?t=383
PS: youtube.com at 6:23, "Leonardo DiCaprio,,Julia Butters in Once Upon a Time in... Hollywood --break"
> Nothing about the future looks particularly good, other than that medicine is improving.
How do you reconcile your thoughts with what the CEOs of these AI companies keep telling us? I.e. "the present is the most amazing time to be alive", and "the future will be unimaginably better". I'm paraphrasing, but it's the gist of what Sam Altman recently said at the World Government Summit[1].
Are these people visionaries of some idealistic future that these technologies will bring us, or are they blinded by their own greed and power and driving humanity towards a future they can control? Something else?
FWIW I share your thoughts and feelings, but at the same time have a pinch of cautious optimism that things might indeed be better overall. Sure, bad actors that use technology for malicious purposes will continue to exist, but there is potential for this technology to open new advancements in all areas of science, which could improve all our lives in ways we can't imagine yet.
I guess I'm more excited about the possibilities and seeing how all this unfolds than pessimistic, although that is still a strong feeling.
Three ways:
* It's the job of CEOs to advocate the benefits of what they're doing.
* Those things might be true, for them.
* Those things might be true, from a global perspective, even if there are some people who are worse off. White-collar workers might just be those people worse off.
It gives me some hope, which fades rapidly as soon as I remember the vested interests such people have.
1) My wife trained as a typesetter on a photo typesetting machine. That was already replacing typesetters working with lead, and the people sorting the used lead, and working with inks etc. They still needed a past-up artist and more. Eventually the GUI based computer arrived, with PageMaker, Quark, Indesign etc. These days she is super productive with a massive online icon library available, full printing and distribution capabilities. Able to do a job that could have involved half a dozen or more people previously.
Are those people unemployed now? Not really (we are talking 1.5 generation now, so not the same people). Unemployment levels are low, and the workforce is significantly larger, with men and women working. The working (outside of the home) part of the population has gone up significantly over a few generations, despite all the new enabling productive tech.
What I see is a lot of visually higher quality work being delivered, but often with the same core content. So productivity has increased, but you get a glossy new shiny report in a PDF, instead of a photocopy of a typewritten page. (Yes, I do simplify. But I think you get the gist.)
2) I started as a system admin, the a systems analyst, worked through project manager, etc. until I was leading startups. In the space where I work now, circular economy and food production, there is so much work to do, that any AI support we can get is welcome. But as the work is innovative, new and not done before, most often the AI tools aren’t that useful, yet. That may change, but with a society that needs to replace a significant part of the infrastructure and processes to achieve a long term sustainable society, I actually don’t worry the AI tools will take my job or any of my colleagues job away. There is plenty to do. I have enough new things in front of me that I could probably keep a whole big venture fund occupied for a long time.
One can adapt to needing to learn new technology, but one cannot adapt to an algorithm out performing you
Sure a part of the population is slow to adapt and therefore at a disadvantage. But the others, like his wife, adapted.
The idea is that this wave of automation will be no different than other times this has happened to us in the past.
I've never totally understood this binary moment when AGI does "everything" better. How can one even define everything?
Our AI partner could be the most intelligent mathematician or researcher. That's great then we can bounce ideas off of them and they can help us realize our professional / creative ambitions.
Sure if our goal is to maximize profit then maybe we can outsource the decisions to an AI agent.
You can get a computer to create infinite remixes of songs. I haven't seen that replacing music producers doing the same.
When everyone can just press a button and have better music automatically generated, based on their exact preferences inferred from their DNA or an fMRI brain scan, what are your creative ambitions?
I'm obviously not talking about today's limited (public) AI, but far into the future, like in 5 years.
* whether or not that is actually the case is irrelevant
So much about music appreciation is about knowing the artist and for example knowing that what they sing about is shaped by their personal history allowing you to identify with it.
A lot about the appreciation of art is the process and the intention behind the artwork. The final image or sond is just one part.
I find it very difficult to define "better" when we're talking about art.
That's the easy stuff! That's the stuff ML has been successful doing for a decade or more.
"Now, you can typeset everything in your office" - early Macintosh ad.
needs to replace a significant part of the infrastructure and processes to achieve a long term sustainable society
AI infra is extremely energy intensive and not sustainable by any metric.Sure you have big data centers full of them but that's already happening with other hardware in other businesses too.
- It emits 30% of climate emissions
- 30% of food produced is gone in losses and waste
- People in middle to high income countries have significant obesity and related problems (some countries have either malnourished people or obese people, and less in between)
- Pollution from our agriculture and aquaculture is killing the ocean near land
- Our intensive agriculture is threatening biodiversity
- We have lost up to 70% of insects in many industrial nations
- Essentially all (90%?) of ocean fish stocks are overfished or at capacity. Even a supposedly rational and environmentally aware nation like Sweden can stop the overfishing. The cod stock has collapsed and now the herring is going too
- The soils are being destroyed or depleted
- The phosphor and nitrogen cycle is broken (fossil fuels or resources that are mismanaged)
I could go on. It is well documented.
I work on circular food production, where we really care about putting together highly efficient nutrient loops and make sure they work locally/regionally. A mix of tech (automation, IT, climate control, etc), agriculture, horticulture, aquaculture, insects etc. As part of this there are very interesting completing pieces with ways of getting the nutrients in creative ways (new food tech), dealing with animal disease (new tech), combined with sensors, ML and just plain old common sense, that can make a huge impact. If we just think through the process a bit more and take responsibility for the externalities, which really are starting to bite.
Much is still overhyped in foodtech imho, specifically the stuff which claims silver bullets without proper circularity. Which is detrimental to the real solutions as investors like simple superscalable solutions, and the simple solutions are mostly not sustainable. (There are of course exceptions).
I'm just trying to not let it get in the way of appreciating the world. I'm planning to travel to mainland Europe sometime next year (gap year). SpaceX has reignited spaceflight, and there's so much cool stuff going on in that space. Science marches on, with a steady stream of interesting discoveries.
And programming is great - for now. It feels slightly strange spending a week writing a project that may be finished with a single prompt in a few year's time, but it's enjoyable.
Maybe I'm overreacting? I've grown up in a pretty calm period, with the west in a clearly dominant position. Maybe this is, paradoxically, a return to normality?
― Vladimir Ilyich Lenin
In the US, let's keep that in mind.
I can see all of my plans for world domination coming together right in front of my eyes. A few years ago I was absolutely certain I’d die without achieving my dream of becoming God Emperor of a united Planet Earth.
We have seen this in small scale of social media that ones self esteem.
We will see a new set of problems that would be much deeper. Videos and image that make you believe false reality, reliance of GPT will generate false knowledge.
False reality problems have started popping up everywhere. It is going to be much deeper. I think we are in for a really crazy trip
So there are a lot of things to be depressed about before you get depressed about a little increase in misinformation and idiocy on the interwebs. I mean... things like polio and the measles are literally back to fuck with us because people are so fucking stupid they think vaccines are a bad thing.
It'll be fine.
If anything the optimist in me is hoping that all this "AI" generated content is going to make the internet so useless that our society (well the part that doesn't believe the earth is flat and that Bill Gates has mini clones in the vaccines) finally get away from it. In my region of Denmark our local police posts their immediate updates on twitter, which was fine when everyone could see them, not so great now that you need an account. I very rarely care about what they post, but around new years a fireworks container blew up near here, and I had to register (and then later delete) a twitter account to figure out if I had to worry about it or not. It'd be nice if the impending doom of fake content is going to move our institutions and politicians away from big tech SoMe platforms and it just might if they become useless.
https://en.wikipedia.org/wiki/CARES_Act
EDIT: For the downvoters, yes components of CARES was in fact inspired by UBI:
https://www.cnbc.com/2020/03/13/andrew-yang-aoc-free-ubi-cas...
https://www.businessinsider.com/coronavirus-aoc-demands-univ...
UBI has three basic properties that the CARES act fulfills none of
1. Covers cost of living for some basic standard (debatable, but should include food, water, and shelter at minimum)
2. Is available to everyone without onerous requirements or means-testing (IE is "universal")
3. Carries a reasonable expectation of continuity such that people can plan around continuing to have it
The CARES act was an emergency measure that absolutely zero people expected or intended to be permanent, it was laden with all the means-testing and bureaucratic hurdles that unemployment generally carries, and it very clearly did not provide adequate support for quite a lot of people
It's meaningless to call something a "test" when it carries none of the properties that proponents of a policy claim would make it desirable. The only perspective from which the comparison even makes sense is from that of someone who's not considered it seriously and come up with a strawman to argue against it (IE something like "UBI is the government gives people some money")
It also seems worth mentioning that I really don't buy the highly political claim that some people seem to view as self-evident: that people remained unemployed longer because they got extended unemployment benefits, rather than as a result of the massive economic shock that prompted that decision in the first place
Many people, rightfully, (over-)react to the American caricature of Christianity (mega churches, Kenneth Copeland, etc.) as the definition of what it is (that's arguably the deception hinted at in the Bible), but reading/trusting the raw word—what's referred to as "sola scriptura"—is remarkably helpful in navigating what's taking place.
> to the American caricature of Christianity
This cannot be overestimated.
As does the I Ching
These are all just Rorschach tests, why choose one of the most corruptible and corrupted approaches?
> why choose one of the most corruptible and corrupted approaches?
Because when the non-prophetic elements of it are applied to life, all of the anxiety, fear, and dread you feel evaporates. It's only when you view it through the lens of a "church" or "leader" (read: group) that it loses its meaning.
I've read the I Ching and it lacks a religious/church element which leads to the conclusion you've had. It's not until people take it and turn it into something it isn't that it loses its value.
Arguably, Christianity, due to its claims, has become weaponized. Interestingly, this very outcome is prophesied in the Bible (which, personally, cements my faith in it what it prescribes).
…said every doomsday preacher since the Bible was written.
Correct, which is why I avoid religion (in the institutional church sense). I'm a bit of an odd duck because I came to the Bible after having been a practicing Buddhist for several years and generally being unexposed to Christianity (save for a lukewarm exposure to Jesuit Catholicism) or any religion growing up.
Having lived a mostly-secular life and only later (at age ~30) coming to Christianity, I can confidently say that in regard to reality, it's taught me that it's highly subjective. What most people consider as "reality" is just the interpretation of what they see that keeps them from losing their mind. For some, reality is being an unhinged hedonist, for others it's planting a garden, and for others it's generally just "trying to be nice and getting along."
Personally, God/Christ (and by extension, what's recorded in the Bible) is the interpretation of reality that makes the most sense to me. In practice/study, I've found that it maps 1:1 with what I see while also filling in the blanks on things I can't explain (e.g., the ability for the human body to heal itself, the pace/behavior of nature, or humanity's unrelenting drive to destroy what it doesn't/refuses to understand).
I guess it’s nice that you believe that, but the truth of the matter is that you are about as close to traditional right-wing mainstream Christianity as it’s possible to conceivably get. Like, if I were to imagine the archetypal Christian hypocritical sinner… it would be you.
Just wait.
These huge updates are interestingly timed though - same day?
One company having that much power is a different matter, and I address it by looking at how we can distribute GPT training through decentralized and open platforms.
It won't be me or you. Whoever it is, they will not share any of the economic upsides of AI with the public unless they are legally forced -- zero, zip, zilch, nada. Even then, they will keep the lion's share for themselves, and they will use their surplus to shape society to their advantage.
So yes, many millions of us have a big problem to worry about, especially considering how much struggling there already is now.
But even if it stays in private hands, one company monopolizing a technology and keeping it expensive/out-of-reach is generally not how technological innovation works. There is generally intense competition between providers, with each aggressively cutting prices to capture market share.
Pure theology
We're already witnessing this with the creation of textual and graphical content through ChatGPT. It's now possible to generate various types of text content and a wide range of graphics at the cost of 10 cents for the dozen ChatGPT API calls. And the work is completed in a few minutes, as opposed to several hours. This represents a several orders of magnitude increase in per capita productivity for these specific tasks. As AI technology advances, the scope of applications benefiting from such productivity boosts is expected to widen, which means human civilization will experience a revolutionary increase in productivity, and with it, resource abundance.
Often with text, the GPT is more of an assistant/advisor/proofreader for me, rather than a stand-alone creator of quality content.
Sometimes it works really well for emails. Like for producing responses to formal communications. It can save cut the time to respond from 10 minutes to 30 seconds.
I for one, would be overwhelmed. In the meantime I will be passionate and joyful about the things I like regardless of whether AI can do them a million times better. I have fun doing it.. while the AI is.. just AI.
Right! I keep saying, that at least we have to kickoff the process. Not even the legislative process, but convincing the public that we'll need it eventually (alternatively a whole different system worldwide, but that will be even harder). Will take a long time anyway.
- paulg
As a side note I shudder to think how many nightmare fuel cursed videos the researchers must have had to work through to get this result. Gotta applaud them for that I guess.
> No one weaves now, and that's fine.
Did horses find new jobs when we moved to steam power? Leave aside the odd horse show and fairground ride. By the numbers, what do you think happened?
They found alternate employment as pack horses in WW1. The problem was solved after that.
This stuff is going to change media and reality so much. Best to get involved in local groups.
Name one, and see if that holds up in 5 years...
Being personable, mediating conflict, leadership in general. Also cooking, baking and masonry.
More climate change, war, microplastics in our body and now extreme joblessness ?
If I woke up and I saw a headline that said OpenAI has developed and AI which told us how to sequester huge amounts of cO2 then I’d be excited and agree.
Exactly. I'm sick of people advocating changes as "progress" until we get some fundamental baseline sampling of humanity's well-being. When "are you depressed?" "do you contemplate suicide?" "are you exhausted?" go up for 10 years around the globe, then people will look like lunatics saying this is "progress" and maybe we'll have a better conversation about where progress actually is.
We had a half-assed lockdown for a few months where most people just kind of stayed indoors and saw noticeable environmental improvements world-wide. An unaligned AGI can easily conclude the best way to fix these problems is to un-exist all humans.
We will be just fine.
Yeah, progress has gotten us nowhere /s
So practically massive destruction of the biosphere so people can sit there smug on there computers in the air-con.
Anyway you reinforced my point. Progress is a good idea, I'm not sure we all have the same ideas about what good progress looks like. Lifting the rest of the developing world by selling them arms and fossil fuels so I can sit on my computer in my room reading smug comments is probably not good progress.
When Facebook came out very few considered it an existential risk. Turns out, it has immense power over elections. Elections have consequences for the well being of billions of people on the planet. Not to mention it might negatively impact the mental health of its users (a large chunk of the human population).
Most of my programming job is tightly coupled with the business processes and logistics of the company I work for, AI will not replace me there.
Also I'm not convinced this is sustainable, I'm thinking this will be like GCI where the first iron man film looked phenomenal but where huge demand + the drive to make it profitable will drive down the quality to just above barely acceptable levels like the CGI in current marvel blockbusters.
Wondering what the runway on this statement is.
"I still haven't seen any text written by AI that doesn't contradict itself a few paragraphs later"
"I still haven't seen any picture made by AI that doesn't look like an abstract nightmare"
"I still haven't seen any picture made by AI that includes hands with the right number of fingers"
"I still haven't seen any videos made by AI that aren't janky and uncanny"
"I still haven't seen any videos made by AI that move me"
The transition from the first to the last took less than five years.
Early attempts to do animation by morphing sort of worked, and were prone to some of the same problems that 2D generative AI systems have. The intermediate frames between the starting and ending positions did not obey physical constraints.
This is a good problem to work on, because it leads to a more effective understanding of the real world.
When I use Unity I write ten lines of code and the tool generates probably 50k. Ever looked into the folder of a modern frontend project after typing one command into a terminal? I've been 99% dependent on code generation for ages.
Maybe some movie you've watched has been spun up by a Sora-like platform based on a prompt that itself was AI-generated from a market research report. Stephen King said that horror is the feeling of walking into your house and finding that all of your furniture has been replaced by identical copies - finding out that all of the media everybody consumes has actually been generated by non-human entities would give me the same feeling
Yes it matters to me a great deal. But there's a reason Stephen King made that observation a long time ago. All the actors in a modern Marvel movie look like they've been grown in some petri-dish in a Hollywood basement and all the lines sound like they come from LLMs for the last fifteen years. There's been nothing recognizably human in mass media for decades. 90% of modern movies are asexual Ken doll like actors jumping around in front of green screens to the demands of market research reports already.
I'm not saying the scenario isn't scary, I'm saying we've been in that hellscape for ages and the particularly implementation details of technologies used to get us there ("AI" in this case) don't interest me that much. And in the same vein, an authentic artist can surely make something human with AI tools.
Exactly. People just aren't seeing this. You don't even have to limit the fake memories to real people. Don't have a girlfriend? Generate videos and photographs of you and your dream girl traveling the world together, sharing intimate moments, starting a family. The possibilities are so exciting. I think the people who hate this idea are people who already have it all. They're not like me and you.
As for social safety nets: if this affects people as heavily as you think (on an unprecedented, never before seen level), the US will almost certainly put _something_ into place and add some heavy taxes on something like this. If tens of millions of Americans are removed from the work force and can't find other work because of this, they'll form a really strong voting block.
Also consider that things are never perfect. We've had wars around the world for a notable amount of time. Even the US has been in places we shouldn't be for a serious chunk of the last century, but things have worked out. We have a ton of news and access now so we're just more aware of these things.
Hopefully that perspective helps a bit. HN and social media can have "doomer" tones quote a bit. Hopefully some perspective can help show that this may not be as large a change as we think.
Or maybe I'm an idiot, as some child comments may point out shortly.
By definition, we don't have data for events we haven't seen before. So instead I reason as well as I can:
Consider the set of all jobs a human being could do. Consider the set of all jobs an AI system could perform as well as a human being but more cheaply. Is the AI set growing, and if so, how quickly?
Prior technology is generally narrow and dumb: I cannot tell my cotton gin to go plant cotton for me, nor can I ask it to fix itself when it breaks. Therefore I take on a strategic role in using and managing my cotton gin. The promise of AI systems is that they can be general and intelligent. If they can run themselves, then why do I need a job telling them what to do?
"Computer" used to be a profession, where people would sit and do multiplication tables and arithmetic all day [1]. Then computing machines came along and put all those people out of work, but it also created entire new categories of jobs. We got software engineers, computer engineers, administrators, tons of sub-categories for all of those, and probably dozens more categories than I can think of.
I think that there's a very high likelihood with the current jobs that humans do better than computers, most will be replaced by cheaper AI labor. However, I don't see why we should assume that set of things that humans do better than computers is static.
This is not as nebulous of a set as it sounds because it has real human boundaries: there are limits to how fast we can learn, think, communicate, move, etc. and there are limits to how consistently we can perform because of fatigue, boredom, distraction, biological needs like food or sleep, etc. The future is uncertain, but I don't see why an AI system couldn't push past these boundaries.
I also struggle to think about all this, but I imagine if you can flip a switch and everything produced and consumed in the economy could be done in half the time, is that a good or bad thing? If we keep flipping that switch and approaching a point where everything is being produced with almost no human effort, does it become bad all of a sudden?
Somehow we'd need to distribute all this production, I'm not sure how it would work out, but just going from what we have now to half or 25% of effort needed is probably an improvement, at least I'd take that.
AI has no spark, no drive, no ambition, no initiative, no theory of mind, and it's not clear to me that it will ever have these things. Right now, it's just a hammer that can build 100 houses a second, but who needs 100 slightly wonky houses?
Just add a cost function.
We can no longer equate intelligence with humanity. Humans are just one kind of intelligence.
I was interpreting the parent comment as saying the spark of consciousness only needed a cost function.
Personally, I disagree that our current neural nets are accurate representations of what goes on in the human brain. We don’t have an agreed upon theory of consciousness, yet ML businesses spread the idea that we have solved the mind and that current LLMs are accurate incarnations of it.
More than the functionality of ai replacing current human jobs, I worry what we will lose if we stop wondering about the universe in between our ears thinking we know everything there is to know.
Also, 100 slightly wonky houses will sell like hot cakes if each one costs less than 1/100th of a not-slightly-wonky house. People will buy 100 of them instead of 1 and just live in a different one every day/hour so they always notice the novel parts instead of the wonky parts. We've had mass manufacturing for centuries and they always prevail when the trade-offs are acceptable.
So, it is completely alike what a lot of humans are like, at least at their jobs?
Many people (including myself) have bought into the narrative that history will repeat here and things will be better eventually, but not how much has to break first, and it's used as a hammer by OpenAI and probably every innovator who disrupted.
They advertised "Safety" but no "Economic Impact" analysis because the latter is less scary and requires difficult predictive work, the former is just narrow legalese defined by 80-year-old congressman they have to abide by to "release" v1.0. There is at-least a Congressional Budget office(CBO) where the 80-year-olds work, flawed as it maybe...
2. I don't think what they have can be protected all that well. Others will catch up.
There is some usefulness to those feelings - this announcement will probably have an impact on your life soon enough. But you cant let every button push and distant threat pull you down can you.
Also remember, life has its own ways: as far as you know, it could also be the beginning of the best days of your life.
“Glad did I live and gladly die, And I laid me down with a will … Here he lies where he longed to be; Home is the sailor, home from sea, And the hunter home from the hill.”
I suppose it's going to be something like "You're asking for more purpose in your life? Sorry, I can't let you do that, Dave."
Dall-E was crazy and then suddenly people were doing the same thing on consumer hardware with an open model within a year.
Filmmakers being able to bring their vision to life using generative models is going to create such a huge expansion of the market.
What people don't realize is that long term these advances are a death knell for mega-corps, not for individuals.
Why do I need to kiss Weinstein's ass to get my movie made if I can do it with a shoestring budget and AI and have the same assistance to create marketing materials, etc. I need a lot less money to break even and can focus on niche markets aligned with my artistic vision instead of mass appeal to cover costs plus the middlemen involved in distribution and production.
I’m sure some people would be more ok to work a shitty job with the hope that they might make it as an artist.
Now that it’s becoming more and more obvious they will not be an artist that’s better than AI, what do they have left to hope for ?
But I agree, there's going to be a "war" around this and if could go wrong, if we are not careful.
Also I take issue with your argument about 'none of the jobs we do now' existing through most of history. Farming, construction, fighting, bookkeeping, cooking, transport, security are all jobs that have been around as long as people have lived in settlements.
Sure, you could point to the long history of nomadic hunting and gathering prior to that, but that's like expanding your argument back to the origin of cellular life or forward to the heat death of the universe in order to make your interlocutor's arguments look insignificant on a cosmic scale. It's not a helpful contribution to addressing the real challenges of the present.
For every one musician that's able to pay the bills, there's 1000 equally talented musicians that can't even get noticed.
For what it’s worth, I think we’re going to see a slide in quality. Maybe there will be a niche for some. But, I think companies will settle for 70% quality if it means eliminating 100% of a full-time position.
This sounds nice, but having worked with many artists in the past a lot of them do it because they're good at it, it's enjoyable enough, and it pays their bills so they can eat.
Telling them, "You're now free to make the art you really wanted to make!" doesn't bring much comfort when you're taking away their ability to put food on the table.
The issue is that those jobs that got automated to "become a craft again" have mostly vanished, except for high-end stuff. Some examples: shoe making, artisan furniture, tailors, watchmakers. Unless you are the best of the best these are hobbies now not something you make money from.
Nowadays most people make money in bleak half-automated jobs (e.g. construction, factory workers) or in white collar jobs sitting in front of a computer in some cubicle doing some mind numbing task for a megacorp.
I'm usually hyped about technological advancement, but very bleak about AI. I think it will just bring more sublte propaganda for state actors, more subtle advertising for megacorps, the dieing of creative jobs like graphic artists or actors is just a sad sideeffect (these will still exist, but only as high end -- we will always have real AAA actors, but the days of extras on movie sets are counted -- lots of the Hollywood protests were because studios started doing contracts for noname actors that stated that the studio will regain rights of the actor's digital likeness)
> Nowadays most people make money in bleak half-automated jobs (e.g. construction, factory workers) or in white collar jobs sitting in front of a computer in some cubicle doing some mind numbing task for a megacorp.
And all the while they enjoy abundance of shoes, furniture, clothes and watches with value/price ratio absurdly high by standards of most of human history.
The craft is an activity, kind of an art by itself. Many find it enjoyable.
The destination is the journey, dude!
I made a twitter thread[1] with weird metal cybertrucks using Midjourney a couple days ago. I personally enjoyed the process and do not have the talent nor the time to do that without generative AI. There are people who do have that talent, but honestly I doubt anyone else would've put in the time.
I think you might have it a little backwards. For most people, the fun part is "making a movie", not "watching hundreds and hundreds of hours of footage picking between 10 different shots". That's the drudgery, and that's the part generative AI can eliminate.
I think you might have it a little backwards. For most people, the fun part is "making a movie", not "watching hundreds and hundreds of hours of footage picking between 10 different shots". That's the drudgery, and that's the part generative AI can eliminate.
No, that's the craft, and solving problems where the continuity doesn't line up, or production had to drop shots, or the story as shot and written sucks in some way, is where the art comes in.
The drudgery is things like ingesting all the material, sorting it into bins, lining up slate cues, dealing with timecode errors, rendering schedules, working your way through long lists of deliverables and so on. You have literally confused the logistics part with the creative act.
But many of those things involve a lot of drudgery, and the drudgery is what these "AI" solutions are best at. If you want to go above and beyond and craft the perfect shot, that opportunity would still be available to you. Why would it not?
When we invented machines that make clothes, did that reduce the number of jobs in the clothing industry? When we got better and better at it, did that make fashion worse? No. If you want a machine made suit for $50, you can find one. If you want a handmade suit for $5000, you can find one.
Tech like this expands opportunities, it does not eliminate them. If and when it gets to the point where Sora is better at making videos than a human in every conceivable dimension, then we can have this discussion and bemoan our loss. But we're not even close to that point.
And with your suit example, you're looking at it from the point of view of consumer choice (which is great) without really looking at the question of of how people in the clothing/textile industry are affected. It's difficult to find longitudinal data at the global level, but we can look at the impact of previous innovations (from outsourcing to manufacturing technology) on the US clothing market; employment there has fallen by nearly 90% over 30 years: https://www.statista.com/statistics/242729/number-of-employe...
The usual response to observations like this is 'well who wants to work in the clothing industry, those people are now free to do other things, great opportunity for people in other parts of the world etc.', but the the constant drive to lower prices by cutting labor costs or quality has big negative externalities. Lots of people that used to make a living thanks to their skill with a sewing machine, at least in the US, are no longer able to monetize that and had to switch to something else; chances they were less skilled at that other thing (or they'd have been doing it instead) and so suffered an economic loss while that transition was forced upon them.
Luddism is never the answer.
Scratch that; luddism is the answer for people who don't actually care about humanity as a whole (but frequently pretend they do) and just want their hobby or their job or their neighborhood to stay the same and for everyone else to stop ruining things. But for the rest of the world, increasing technological efficiency means more people get more things for less. This is good actually.
I see these AI models as lowering the barrier to entry, while giving more power to the users that choose to explore that direction.
Meanwhile with AI, given the same model and inputs - including a prompt which may include the names of specific artists "in the style of x" - one can reproduce mathematically equivalent results, regardless of the person using it. If one can perfectly replicate the work by simply replicating the tools, then the human using the tool adds nothing of unique personal value to the end result. Even if one were to concede that AI generated content were art, it still wouldn't be the art of the user, it would be the art of the model.
You asked an AI to make something for you. Thats not really making it yourself. Its like hiring someone to create something for you.
Disclaimer: This post was generated using an llm guided by a human who couldn't be bothered explaining why you're wrong.
The feeling is mutual.
Impressive that you can dismiss an entire genre of art as trivial mindlessness
I'm merely pointing out why it's stupid.
(edit to give some body to my comment above:
Hosting a great dinner party is hard work and requires coordination between food, decor, seasonality, people attending, etc. It is akin to a director coordinating the parts of a film. So I do think hosting a good dinner party can count as artistic expression.
I don't know the parent comment's intended reading, but I was reacting to the idea that typing a Sora prompt makes someone a good artist. If the parent means instead that AI allows people to coordinate multiple media in a broader expression that was not possible otherwise, then I fully agree.
)
I didn't say "AI art makes you a great artist" I said it helps you make great art
Tools allow people to express themselves at a higher aesthetic level without needing the extreme technical skills.
GenAI is a tool that lets creators of one medium expand to other mediums without much effort. Like having transcripts auto-generated for a visual podcast, just in the other direction. Low budget (or amateur) poems/songs can turn into short videos; or replace generic album art with better quality generic album art.
The draw will be the primary medium, the rest will just be an extra bonus.
I've updated my comment to explain my view, which I think closely aligns with yours.
When cameras were harder to use, you had national geographic taking you all over the world to photograph different locations because only you and a handful of people knew how to take a picture properly. Now you just hire a local person with a camera to go take the picture you want since it’s much easier to use a camera.
You had people doing photography for ads, now stock photography will do for most brands.
You had people buy high school portraits, I am not sure if people buy those anymore, but a picture of what you looked like in high school is worth a lot less when you can take a selfie every other day.
I don't see it as a negative per se, the thing is most people won't have the decency to keep all that shit for themselves and, say, share just the best 1% they produce. They will flood their social networks and the rest of us will have to sweep through the crap for our daily dose of internet memes.
Blockbuster movies depend to a large extent on the pedigree and abilities of their cast. For the big studios, these models are therefore quite useless apart from bringing dead actors alive again. If publishing material created from living actors without isn't illegal already, in a few years it will be.
This might actually save the movie industry and force it to improve the quality of its output. There will be a huge indy scene of movie makers using models that can only compete via the content of the movies they produce. The realism of the characters won't matter because everyone can have those now. The current big studios will be forced to make very good use of human actors to compete with them though, and become innovative again.
Now with such AI tools, you can write scripts, create art work, crate footage, record voice overs and dialogues. All of this means less need for labor - creative that will not only cause huge employment in the sector but also lead to protests, it already happened last year in Hollywood, it's going to get louder and louder unless we put regulations to prevent job disruptions.
I like going to theater or opera. Even for famous pieces, the performance will be slightly different and unique every time. Imperfect, but with changing and nevertheless accomplished actors, singers, musicians, and dancers. Many people feel the same and that's why they watch live performances of singers, bands, and DJs.
There is no moat. This will all be commonplace for everyone soon, including with a rich open source community.
OpenAI won't let you do nudity or pop culture, but you can bet your uncle that models better than "Sora" will be doing this in just a few months.
> We are giving the enjoyable parts of life to a computer. And we are left with the drudgery.
No. This means that the tens of thousands of people working in entertainment building other people's visions can now be their own writers, actors, and directors.
This is a collapse of the Hollywood studio system and the beginnings of a Cambrian explosion of individual creators.
An AI capable of this would likely also be able to e.g. flip burgers and bring us to Fully Automated Luxury Communism.
Most “good” art isn’t just what you see, it’s also the story behind it. Why was it made? What is the story of the artist? What does it make you feel?
AI might allow more people to tell some of those stories they may have lacked the raw skills to tell before. And for those who have the skills, they can make exactly what they envision, without being limited by some of the randomness in the AI. I think there will always be a place for that, and at the top of the market, that’s what people want.
My point is that I am an aspiring artist, who is waiting tables and I invest all my spare cash and time to get better at my craft and hopefully allow my craft to support me financially.
Any hope of financial benefit coming from my craft is quickly taken away by Dall-e. This has nothing to do with how much I enjoy my art.
Are we though? People still do plenty of things out of interest or hobby, despite it being fully automatable?
e.g. blacksmithing or making certain homemade things?
While these are non digital things, why can't we apply the same thing here?
Some people still hand write assembly out of the novelety and interest of it. Despite there being better tools or arguably better ways of writing code.
Yesterday I asked a local llm to write a python script to have a several multimodal llms rank 50,000 images generated by a stable diffusion model. I then used those images to train a new checkpoint for the model and can now repeat the process ad infinitum.
In the olden days of 2020 I would have had to hire 5000 people each working for a day to do the same.
Medium term is ... less good.
these advances are a death knell for mega-corps, not for individuals.
Certain mega-corps. We've been down the email, finance, and social media road and know what people actually use. It's centralized corp infra.I think we'll soon see a suburban mom in middle America with a part-time penchant for storytelling make a blockbuster video game mostly by herself.
The only original creativity will be in creating new formats and new kinds of experiences - which will mostly mean inventing new kinds of AI.
Everything made in an existing format will either be worthless or near as.
Same applies to software dev. Far more quickly than most people expect, it will also apply to AI dev.
And to hit scale first might be enough of a moat?
Of what market? Certainly not film production. I have my doubts about whether it will expand the market for films, in the economic sense. The lower the cost of producing and distributing a film, the lower the monetary value people will place on it.
Look how most music artists are no longer able to survive on royalties, and a few massive streaming companies have an astounding profitable oligopoly of consumers' music interest. Yes, many pre-streaming publishers were exploitative or unethical, but I'm not convinced that it was to a greater degree than the current market leaders. Consider also that the streaming revolution steamrolled many, perhaps most, indie record labels that supported niche genres; some live on but are no longer able to sustain physical output and a reduced to being digital marketing companies.
Now, people will continue to tell stories and entertain others, so technology like this will be good for people with an artistic vision who can't easily access publishers for whatever reason. It will certainly allow people to pursue bold artistic visions that would not otherwise be economically feasible - exotic locations, spectacular special effects, technically complex perspective moves. Those are good things; I worked in the film industry for a long time and have several unproduced scripts that I'd like to apply this technology to, so I'm not rejecting it.
However, more content doesn't necessarily translate into more economic activity; I think it very likely that visual media will be further devalued as a result. People who have spent years or a lifetime developing genuine craft will be told to abandon it in favor of giving suggestions to a computer system, and those who don't will be laughed at or suspected of fakery, because fakery is so widespread these days (Relevant recent example: https://news.ycombinator.com/item?id=39379073). The easier it becomes to make something, the less value the market will assign to it; rational from the abstracted perspective of pure price theory, disastrous in real life.
Increasingly, we seem to be tilting towards a Huxley-esque dystopia of stunning and infinite-feeling virtual worlds to which we can escape on demand, and an increasingly shitty real world marked by the brutal economic logic of total resource and information exploitation. Already a stock rejoinder to complaints about the state of things is that humanity is on paper richer than ever before, to the point that bums have smartphones and anyone can afford an xbox. I have a homeless neighbor who's living in his car, spending his dying years watching YouTube on his phone to fall asleep because he's lonely. Technically this is an expansion of the market, but I don't think it's a good outcome.
The torrents will be filled with convincing full-length films which are maybe not as good as the originals, but still very watchable.
We will have infinite chances to get better films than Indiana Jones and the Kingdom of the Crystal Skull.
*: except the ones starved for compute resources of course
Anybody can publish anything with the web. Have publishing houses disappeared? Have scientific journals disappeared?
The insanity is doing the same thing over and over and thinking "this time it will work".
https://www.technologyreview.com/2020/11/12/1011944/artifici...
AI turns skilled labor into cheap ones, supposedly, right? It's a massive enabler for mega-corps. Not a death knell.
In a world where attention is scare I sadly think big corps and power brokers will still play a large role. Maybe not though.
Advances like this are necessary steps to get us there.
People will be displaced in the short term, as they have for every other large scale advance... Cars, assembly line and so on. Better to focus on progress and helping those disaffected the most along the way
AI is going to be massively deflationary. How useful is UBI when the cost of goods and services approaches zero due to automation?
With that said, I can imagine the federal reserve will then helicopter in money to everyone in order to reach its 2% inflation target, which kinda sounds like UBI.
There are winners and losers, but it's absurd to think that we should avoid progress to protect jobs.
My preferred method for ensuring a just transition during times of technological progress (and to eliminating involuntary unemployment generally) is a Federal Job Guarantee http://www.jobguarantee.org/
Personally, the advances in AI have just made my job easier and allowed me to get more done. I don't see that trend changing either.
Industrialisation and computers/automation took away massive amount of jobs while globally improving people lives, this may possibly (maybe not) do the same.
If in the future, anybody can write the book, create a photograph or a motion picture or an music album with just few words describing what they have in mind, this will be a tremendous productivity improvement and will unleash an overflow of human creativity.
I like to compare it to what Jobs said about computer, they are the "bicycle of the mind" [1].
[1] https://medium.learningbyshipping.com/bicycle-121262546097
Open-source will catch up in <6 months. Note that Meta will ship llama3 anytime. So will Mistral.
I am a PM, and switched to becoming a builder. Enjoy learning, keep building. What people take time to realize is that building things is a habit. As you combine that with ongoing learning, you will enjoy the process and eventually build something to earn a living.
Unfortunately I have nothing to showcase yet.
Touch grass, tend your garden, play with your kids, drink a beer, bake a pie, write a poem, take a walk, carve a sculpture, play a board game, mend a sweater, take a breath. Relax.
we will finally free ourselves from mediocre humans being the bottleneck for everything
If you want a safety net move to Scandinavia
OSS so far has been an effective tool as a check and balance against big corp.
Whether people are employed or not is a policy decision of the central bank and not related to how good AI is.
I am hoping our society and civilisation doesn't implode as everyone becomes unemployable as the economy and social contract collapses
To me, it seems guaranteed that will be drastic changes. There will be many attempts at new ways of organizing society with successes and failures along the way. Not out of altruism or desire to share but out of self-interest of those who collect the power afforded by AI and automation.
Realize that it's a choice to respond to things this way. This feeling comes from a certain set of assumptions and learned responses.
Remember that people are bad at predicting the future. Look at the historical track record of people predicting the implications of technological advancement. You'll find that almost nobody gets it right. Granted, that sometimes means that things are worse than we expect, but there are also many cases where things turn out better than we expect. If you're prone to focusing on potential negatives, maybe you can consciously balance that out by forcing yourself to imagine potential positives as well.
Try to focus on things you personally have control over. Why worry about something that you can't change? Focus on problems that you can contribute to solving.
Industrial scale technology might ruin us though, so you might have some point, mostly I'm referring to climate change which is for sure the greatest existential threat imaginable right now. However it seems technology might bail us out here too, nuclear and renewables.
We undoubtedly have reaped immense benefits from the industrial revolution for example- that doesn't mean I'd have any interest in living through it or that it was executed in a way that prioritized the people who lived during those times.
The more freely available the tech is, the easier it is to reproduce, the less the average joe is going to be locked out of the benefits.
My entire career and everything good in my working life has involved open source software and I'm sure that it will continue to be the case.
material:
- water: too little => thirst, too much => drown
- heat: too little => freeze, too much => burn
- food: too little => starvation, too much => obesity
spiritual:
- courage: too little => cowardice, too much => foolhardiness
- diligence: too little => slothfulness, too much => workaholism
- respect: too little => disregard, too much => idolatry
etc.
life is a balance
Not in my thirties and almost nothing I worried about has come true. I mean, tomorrow we might get wiped out by a runaway technological singularity, but I could've spent the last 30 years of my life worrying a lot less too..
I went to cognitive behavior therapy, for me it was like someone opened up my mind and showed it to me on a screen, it was a mirror into my head. It was amazing how it felt like I could rewire my thought patterns over the course of a few months.
The main takeaway from it all however, was the mantra: _thoughts are not facts_.
If you can realize that your thoughts are not objective truths, you will be much better off in almost every aspect of your life, because after living this mantra for many years, putting it to the test constantly, I know it's solid.
Later on I read a lot of Buddhist philosophy which matched incredibly well with the therapy because a lot of Buddhist thinking and meditation practice is quite similar in it's approach. This sort of reinforced the validity of the CBT because I realized wise people have known about seeing things in an objective light for millennia, which was validating for me and helped me continue on the introspective path.
Basically, we're all hallucinating in one way or another, almost all of the time, and that is ok, just be aware of that. When we're worried about the future, we're worried about something which doesn't yet exist, which is actually crazy.
Of course it doesn't mean we should just ignore long term problems, no one advocates for that. But we shouldn't assume we know the outcome in advance because that often causes stress.
Warning: I think that for most westerners, it's "safer" to get into something like CBT, Buddhism comes with some IMO very confronting ideas for a lot of people where as CBT is much more user friendly for westerners.
The tens of thousands of people working in entertainment building other people's visions can now be their own writers, actors, and directors. And they'll find their own fans.
Studios will go away. Disney will no longer control Star Wars, because your kids will make it instead. In fact, the very notion of IP is about to drive to zero.
And OpenAI won't own this. They won't even let you do "off book" things, and that's a no-go for art. Open source is going to own this space.
There are other companies with results just as mature. They just didn't time a press release to go head to head with Gemini.
Edit: Putin on AI, 2017 https://youtu.be/aJELcvjREgk?t=29
"Whoever takes the lead in this area will be the ruler of the world."
"тот кто станет лидером в этой сфере будет властелином мира"
I think in the initial years there'll be some major incidents where a fake thing gets major attention for a few days until it's debunked, but the much larger issue will be the inverse.
Only some people can make Star Wars (the pinnacle of independent filmmaking if you read Lucas's biography). It has nothing to do with the tools.
IP in the arts is how artists get paid.
I can assure you that no one in the creative industry feels liberated by these tools. Do you realise that just because you are good at lighting, you don't want to be an actor and make a movie? No, you like to be good at lightning, work with others who are good at what you do, and create a great work of art together.
AI imagery only knows what exists. It's tough to make it do innovative technical effects and great new lightning. "oh my god, stock video sites are dead" Yes, exactly; stock, by definition, is commoditised.
It's terrible news for the people being replaced. Their training and decades of experience is their competitive advantage and livelihood. When that experience becomes irrelevant because anyone can create similar quality work at the push of a button, they're suddenly left with nothing of value in a world flooded with competition.
Why do people always say this / think that saying this is helpful? Try saying to someone with ADHD, "realize that you are choosing not to get your chores done today. You're choosing not to get out of bed on time. You're choosing to show behavior that your peers describe as 'lazy'. This will keep happening as long as you let it!"
So what if you have the ability to choose whether you are depressed or not? Not everyone got the same choice. Not everyone still has that choice.
I don't really expect another solution, but this always kind of bothers me when I see people saying everything is a choice.
With neurodivergence and mental disorders, what you see as "choice" can end up not being a choice at all.
It's an easy thing for people to say when they don't really want to help others.
What I posted is what I have personally found to be the most useful advice in overcoming self-destructive mental habits.
I'm glad a one-time, one-line quip worked for you, but in my experience, positive mental habits are built over time, through support and continuous practice.
That's making a lot of assumptions about my personal history that you couldn't possibly know anything about.
>but in my experience, positive mental habits are built over time, through support and continuous practice.
I agree, and I don't think anything I said implies otherwise.
If you are responding to people's problems with common one-liners, it can be interpreted as belittling someone. It could be interpreted as an attempt to over-simplify or attempt to make them feel they are "inferior" to see and solve their issues, when their issues are to them, much larger than a random one-line quip.
So so so much of this about. Not just disorders too.
You might read my comment as trying to claim that my disorders define me and that because I have these disorders I can afford to give up on this stuff because 'it's hopeless'. Truth is I've been trying to get past this for damn near a decade at this point and it's not nearly as easy as you make it out to be, and that's why I say that I don't have the same choice you think I do.
I didn't even know I had ADHD until a year or so ago, I'd just routinely lose the ability to do the stuff I love and I'd have to go find something else to do instead. Depression would stem from all the things I knew I loved but that I could no longer motivate myself to do. In fact I was probably even worse off before I knew about this because I thought that I was just doing something wrong, not being controlled by an invisible menace that most other people don't even know exists
I don't mean to be hostile or to impose that it can't be as easy as you're describing. I just don't think that it's right to say it's always just a choice how you react.
I have tons of completely involuntary reactions caused by primarily trauma, but I can't control them. They do things like force me literally out of consciousness with overwhelming guilt and/or sadness. That's not a choice. I didn't choose that. That's completely autonomous!`
It's easy to say "don't worry" if you haven't been affected by events like this. I feel it's stronger for society to say "I don't know what will happen, but we'll work through it together."
But reality is as reality is and nobody is owed a desk job. These are very exciting times with what type of society could be built with this tech, human inefficiencies are responsible for a lot of suffering that we can might be ab;e to stamp out soon.
We are building technology, to suggest no agency is helpful in avoiding any feeling of responsibility or guilt — perhaps rendering your comment within the realms of waxing philosophical.
Who better to worry about this than the people of hacker news?
From a pure mental health standpoint, sure, it’s solid advice but I think it’s narrowed the context of the broader concern too much.
An alternative to learned helplessness of “nothing you can do” is to encourage technologists to do the opposite.
Instead of forgetting about it, trying to put it out of your mind, fight for the future you want. Join others in that effort. That’s the reason society has hope — not the people shrugging as people fall by the wayside.
Depression mediation by agency feels more positive, but I don’t have a lot of experience tbh. Just a view that we, technologists, shouldn’t abdicate responsibility nor encourage others to do so.
That culture, imo, is why a large section of tech workers, consumers and commentators see the industry in a bad light. They’re not wrong.
EDIT: to add, “what problems can I personally solve” also individualises society’s ability to shape itself for the better. “What problems can I personally get involved in solving”, “what communities are trying to solve problems I care about” is perhaps the message I’d advocate for.
Cat's out of the bag. There is no legislation that will stop this. Not unless/until it has some obscene cost and AI gets locked down like nuclear weapons. But even then, it's just too simple to make these things now that the tech is known.
I sure don't know the answer but we just don't know what's coming next. Gonna have to wait and see.
As an individual, I think the experience would be dreadful
Like billionaires didn't already scare a lot of people.
For example, I’m getting text messages all day long from random politicians asking for money. If you told people 50 years ago that one day we’d be carrying devices where we could be pinged with unwanted solicitations all day and night, they might have imagined an asphyxiating nightmare. But in reality, it’s mainly a nuisance.
The point is that your brain makes all kinds of emotional predictions about the future, but they aren’t really very useful and if you’re experiencing depression or anxiety, I can guarantee they are biased predictions.
Also I think there will be more junk out there with how easy it will be to mass produce garbage.
We need to push back generally against our top heavy economy. More local and regional focus would help.
For most of human history, people didn’t have constant access to art work or videos. Things were fine. Maybe instead of watching manufactured shit on social media, go see a live theatre production. Seek genuineness.
You can live a great life without ever seeing a picture or video.
People will not loose their jobs, because you still need someone to input prompts for 10 hours straight in order to get a piece of video you want.
Natural language is not perfectly precise to get exactly what you want from a model like this, and the results remain kinda random. Instead of making video in a traditional tool you will be spinning AI roulette until it generates your desired result. And even then you will probably want to edit it.
Tools advance, but tools remain tools.
I don't think anything will fix or change this, definitely not UBI, the situation is a fundamental part of the human condition. I share your dread and fear that I will not be able to compete, even if my life improves by all other measures.
The one thing that would dramatically change my calculus is medical advances that significantly push back death and aging.
One reason is, readiness of tech does not mean it’s being applied.
Another is just like one OpenAI came out of no where, others will too. It’s normal to be focused on a few things to lose sight of that.
Gemini realistically does some impressive things.
What can we do?
Building the tech is important but applying it well for actual adoption is still wide open for the average persons use.
It does seem to mean that what we think might take 5y probably will take 1y in 2024, if not less like 2023. So think 10x, and 10x again as the real goal.
The forces of blind or cynical techno-optimism accelerating capitalism may feel insurmountable, but the future is not set in stone. Every day around the world, people in seemingly hopeless circumstances nevertheless devote their lives to fighting for what they believe in, and sometimes enough people do this over years or even decades that there’s a rupture in oppressive systems and humanity is forever changed for the better. We can only strive to live by our most deeply held values, within the circumstances we were placed, so that when we look back at the end of our lives we can take comfort in the fact that we did the best we could, and just maybe this will be enough to avert the inevitable.
I just dread, *shrug*. You don't have to be depressed or doomed, it all comes from premature predictions.
That said, we are surely at the phase similar to that one right before the internet, if not electronics. I genuinely don't understand people who write AI off as yet another Bitcoin or "just an enhanced chatbot". They'll have to catch up on an insanely complex area (even ignoring rocket science behind all that) which will do without them, and right now is the opportunity to jump the ship early.
It's only my nobody's opinion, but I can't see the way in which _that_ could fail. I find it incredibly stupid to think so and to just live your life as if nothing happens. If you're not the ruling class or a landlord, f...ing learn.
Current AI wave is a corporate funded experiment desperate to find something compelling beyond controlled demos to economically recoup the deepening hole in their balance sheet. The novelty has begun to wear off, the innovation has started to stagnate and the money running out. The only money making innovation left to be seen is in creating more spam. Thats where I see this wave headed.
OpenAI has proven it's a shit company with rotten fundamentals playing with a shiny new toy. They will crash and burn spectacularly. As many before have done in various fields.
My reaction after using any AI tool from the last couple years to do anything meaningful ends with just a big facepalm.
Its a grift[1]
...and our future lies in the hand of venture capitalists, many of whom have no moral compass, just an insatiable hunger to make ever larger sum of money.
Every industrial revolution and its resulting automation has brought not only more jobs but also created a more diverse set of jobs. Therefore also new industries are created. History rhymes, the ruling fears in such times have always been similar. Claims are being made but without any reasonable theories, expertise or provable facts (e.g. Goldman Sachs unemployment prediction is absolute bs). This is even more true when such related AI matters are thought about in more detail. Furthermore, even though employing tens of millions of people probably, only a few industries like content creation, movie etc. are affected. The affacted workforce of these industries is highly creative, as they are being paid for their job. The set of jobs today is big, they won't become cleaning staff nor homeless.
This technology has also to proof itself (Its technical potential is unlimited but financially limited by the size of funds being invested, and these are limited)
Transition to the use of such tools in corporations could take years, depending on the type and size and other parameters. People underestimate the inefficiencies that a lot of companies embody - and I am only talking about the US and some parts of Europe here. If a company did their job for 2 decades the same way, a sudden switch does not happen overnight. Affected people have ways to transition to other industries, educate themselves further and much more. Especially as someone living in the west, the opportunities are huge. And in addition, the wide array of different variables about the economy and the earth, and everything its differing societies are, comes into play: Some corporations want real videos made by real people; Some companies want to stay the way they are and compete using their traditional methods; Corporations are still going to hire ad agencies - ad agencies whose workflow his now much more efficient and more open to new creative spheres which benefits both customer and themselves. They list could go one endlessly.
Lots of people seem to fear or think about the alleged sole power OpenAI COULD achieve. But would that be a problem, would "another Alphabet" be a problem? Hundreds of millions of people benefited and are benefiting today from their products. They have products that are reliable and work (This forum consisting of tech experts is a niche case, nearly all people don't care at all if data on them is being used for commercial purposes). Google had a patent guaranteed monopoly on search. But here we have: an almost non patented or patentable market, an open source community, other companies of all sizes competing, innovation happening and much more. It is true that companies like OpenAI have more funds available to spend than others, but such circumstances have always driven competition and innovation. And at the end of the day, customers are still going to use the best product they have decided to be so.
I know I may be stating the obvious but: The economy and the world is a chaos system with a unpredictable future to come.