DALL-E presentation also looked cool and everyone was stoked about it. Now that we know of its limitations and oddities? YMMV, but I'd say not so much - Stable Diffusion is still the go-to solution. I strongly suspect the same thing with Sora.
They're literally taking requests and doing them in 15 minutes.
But all to be said, it is no less impressive after this new demo
Sarah is a video sorter, this was her life. She graduated top of her class in film, and all she could find was the monotonous job of selecting videos that looked just real enough.
Until one day, she couldn't believe it. It was her. A video of of her in that very moment sorting. She went to pause the video, but stopped when he doppelganger did the same.
Sure, for people who want detailed control with AI-generated video, workflows built around SD + AnimateDiff, Stable Video Diffusion, MotionDiff, etc., are still going to beat Sora for the immediate future, and OpenAI's approach structurally isn't as friendly to developing a broad ecosystem adding power on top of the base models.
OTOH, the basic simple prompt-to-video capacity of Sora now is good enough for some uses, and where detailed control is not essential that space is going to keep expanding -- one question is how much their plans for safety checking (which they state will apply both to the prompt and every frame of output) will cripple this versus alternatives, and how much the regulatory environment will or won't make it possible to compete with that.
Strictly to prompting, probably, just as that is the case with Dall-E 3 vs, say, SDXL.
The thing is, there’s a lot more that you can do than just tweaking prompting with open models, compared to hosted models that offer limited interaction options.
and i still prefer Dalle-3 to SD.
there are mini-people in the 2060s market and in the cat one an extra paw comes out of nowhere
Most fictional long-form video (whether live-action movies or cartoons, etc) is composed of many shots, most of them much shorter than 7 seconds, let alone 60.
I think the main factor that will be key to generate a whole movie is being able to pass some reference images of the characters/places/objects so they remain congruent between two generations.
You could already write a whole book in GPT-3 from running a series of one-short-chapter-at-a-time generations and passing the summary/outline of what's happened so far. (I know I did, in a time that feels like ages ago but was just early last year)
Why would this be different?
I partly agree with this. The congruency however needs to extend to more than 2 generations. If a single scene is composed of multiple shots, then those multiple shots need to be part of the same world the scene is being shot in. If you check the video with the title `A beautiful homemade video showing the people of Lagos, Nigeria in the year 2056. Shot with a mobile phone camera.` the surroundings do not seem to make sense as the view starts with a market, spirals around a point and then ends with a bridge which does not fit into the market. If the the different shots generated the model did fit together seamlessly, trying to make the fit together is where the difficulty comes in. However I do not have any experience in video editing, so it's just speculation.
In particular, looking at the video titled "Borneo wildlife on the Kinabatangan River" (number 7 in the third group), the accurate parallax of the tree stood out to me. I'm so curious to learn how this is working.
[Direct link to the video: https://player.vimeo.com/video/913130937?h=469b1c8a45]
https://youtube.com/watch?v=P1IcaBn3ej0
From a few years ago, where the game is rendered traditionally and used as a ground truth, with a model on top of it that enhances the graphics.
After maybe 10-15 years we will be past the point where the entire game can be generated without obvious mistakes in consistency.
Realtime AI dialogue is already possible but still a bit primitive, I wrote a blog post about it here: https://jgibbs.dev/blogs/local-llm-npcs-in-unreal-engine
The DK1 I could wear for like 1 minite before feeling sick, so they are getting better ...
I am prone to sea sickness. Maybe it is related.
Yeah, but I mean who knows why. I know some people can't, my GF is one of them.
I've often wondered if im ok with it because im used to the object on head stuff (like 25 odd years of motorcycle riding/ergo helmet wearing) and close up, high fov coverage fast past gaming? (I play on a 32" maybe 70 cms from the eyes give or take.)
> I am prone to sea sickness. Maybe it is related.
I'd think it might be given my understanding of why illness in many is triggered. It's odd because I never got sick from it, but i've seen others get INCREDIBLY ill in two different ways.
1. My GF tried to use simple locomotion in a game and almost vomited as an immediate reaction
2. A friend who was fine at first, but then randomly started getting very slowly ill over a matter of like an hour, just getting more and more nausea after the fact.
It's unfortunate, because due to lack of bad feelings/nausea/discomfort etc, I love VR. I equally from those around me can see no real path forward for it as it stands today though because of those impacts and limitations.
That being said, maybe they get smaller, lighter, we learn to induce motion sickness less, I dunno. I'm not optimistic.
If/once they get it working though, society will shift fast.
There’s an XR app called Brink Traveler that’s full of handcrafted photogrammetry recreations of scenic landmarks. On especially gloomy PNW winter days, I’ll lug a heat lamp to my kitchen and let it warm up the tiled stone a bit, put a floor fan on random oscillation, toss on some good headphones, load up a sunny desert location in VR, and just lounge on the warm stone floor for an hour.
My conscious brain “knows” this isn’t real and just visuals alone can’t fool it anymore, but after about 15 minutes of visuals + sensory input matching, it stops caring entirely. I’ve caught myself reflexively squinting at the virtual sun even though my headset doesn’t have HDR.
For games like call of duty or other hyper realistic games it very likely will be.
A large part of fighting games is the style.
The cost difference of just making bespoke art and tuning an AI system to generate it for you may not be worth it (at least right now.)
AI researcher at MSFT barely have more insights about OpenAI than you do reading HN.
In the same month, they were also using GPT4 in public - before OpenAI.
And they had access to GPT4 in 2022 (which was when they decided to create Bing Chat, now called Copilot).
All the current GPT4 models at MSFT are also finetuned versions (literally Creative and Precise mode runs different finetuned versions of GPT4). It runs finetuned versions since launch even...
OpenAI is likely limited by how fast they are able to scale their hiring. They had 778 FTEs when all the board drama occurred, up 100% YoY. Microsoft has 221,000. It seems difficult to delegate enough headcount to all the exploratory projects of MSFT and it's hard to scale headcount quicker while preserving some semblance of culture.
I don't think what you're saying is correct though, either. All the early news outlets reported 49% ownership:
https://en.wikipedia.org/wiki/OpenAI#:~:text=Rumors%20of%20t...
https://www.theverge.com/2023/1/23/23567448/microsoft-openai...
https://www.reuters.com/world/uk/uk-antitrust-regulator-cons...
https://techcrunch.com/2023/01/23/microsoft-invests-billions...
The only official statement from Micorosft is "While details of our agreement remain confidential, it is important to note that Microsoft does not own any portion of OpenAI and is simply entitled to share of profit distributions," said company spokesman Frank Shaw.
No numbers, though.
Do you have a better source for numbers?
Good luck generating anything similar to an 80s action movie. The violence and light nudity will prevent you from generating anything.
These video clips just generic stock clips. You cut cut them together to make a sequence of random flashy whatever, but you still can't do storytelling in any conventional sense. We don't appear to be close to being able to use these tools for the hypothetical disruptive use case we worry about.
Nonetheless, The stock video and photo people are in trouble. So long as the details don't matter this stuff is presumably useful.
I'm not supporting it in any way, I think you should be able to generate and distribute any legal content with the tools, but just giving a possible motive for OpenAI being so conservative whenever it comes to ethics and what they are making.
They're trying to be all-around uncontroversial.
Worth noting that Google also has Phenaki [0] and VideoPoet [1] and Imagen Video [2]
[0] https://sites.research.google/phenaki/
Or Meta will do it for them.
Don't get overly excited until you can actually use the technology.