StoryDiffusion: Long-range image and video generation
storydiffusion.github.io
storydiffusion.github.io
But I guess it took me multiple minutes to find these problems, watching each video clip many times, rather than having any of them jump out at me. So, it's not like literally full consistent object persistence, but at a casual viewing it was very persuasive.
Maybe people who shoot or edit video frequently would notice some of these problems more quickly, because they're more attuned to looking for continuity problems?
Here we're looking at video, of high quality individual frames, where the inconsistencies are maybe clear and maybe not - but compared to craion (around the time of dalle): https://i.ytimg.com/vi/lcoitxKbw_0/maxresdefault.jpg it's wild how that's changed. And this capability was a vast improvement over things before (at least ones that weren't fixed goals, the GAN approach to faces in headshots was very lifelike before this)
I’m no video editor but I noticed straight away that The characters’ eyes and hair tend to change, sometimes dramatically as they turn their head. Also, the head movement tends to be jerky or abrupt especially in the middle of the turn.
It seems to me that eventually these systems are going to have to be grounded in some hard truths about our world. Like, there are things called objects, objects can be distinct, objects can have relationships between other objects, etc. Then the generative network would have to generate data around these priors. Or maybe they already have that, I don't know how they work.
It's also telling that most of the shots do their best to hide hands - whenever they are visible, they are obviously broken.
What about the woman with glasses? Her face literally "jumps"[1] Same with this guy's hands[2]
Interesting, we notice that [1] has "sora" in the name though I think it is a reference to the main image on sora[3]
Not sure if the gallery is weird to anyone else, but it doesn't exactly show new images and the position indicator is wonky.
The thing that makes me most suspicious is seeing the numbers on these demos. 1, 2, 4 (terrifying to me), 5, 65, 66, 68, 72, 73, 83, 85, 86 (is this Simone Giertz? Vic Michaelis?). The part that is tough about evaluating generative models is the cherry picking for demonstrations. You have to do it or people tear your work apart but also in doing so you give a false impression of what your work can actually do.
IMO it has gotten out of hand and is not benefiting anyone. It makes these papers more akin to advertising than communication of research. We talk about integrity of the research community and why we argue over borderline works but come on, if you can get a better review by more samples, you can get better reviews by paying more, not by doing better work. A pay to play system is far worse for the integrity of ML (or any science) than arguing over borderline works.
Edit: I think it is also a bit problematic that this is posted BEFORE the arxiv link or GitHub goes live. I'd appeal to the HN community to not upvote these kinds of works until at least the paper is live.
[0] https://storydiffusion.github.io/MagicStory_files/longvideo/...
[1] https://storydiffusion.github.io/MagicStory_files/longvideo/...
[2] https://storydiffusion.github.io/MagicStory_files/longvideo/...
Not even in the same ballpark. Even when things are wrong in Sora it seems like the imagery is still very crisp. If I watched these videos for 5 minutes I know I would get a headache.
I am constantly surprised by how well it copes with my typos, grammatical errors and generally poor spelling.
In the "early days" of GPT-4 I tried testing it as a way to get around poor transcription for an in-car voice assistant. It managed: "I'm how dew yew say... Freud?" => Turn up the temperature... which was nonsense most people would stare at for a long time before making any sense of.
I suspect that some grammar and spelling issues may be the authors themselves. For example "A Asian Man": "a" instead of "an" is a common mistake for many Asian languages due to not having similar forms in their languages. So considering consistent article errors, I expect this to be an issue from the authors. Not sure the "M" capitalization. Similar things with "The man have breakfast", "They have launch at restaurant", "They play in (the) amusement part."
Considering the comics have similar types of error (the squirrel one clearer) I'd chalk it up to language barrier instead of the process. Though LeCun is not wearing gloves on the moon, and well...
"I made a new thing", go to the repo, COMING SOON. Or this, here's the paper, no we won't show our work.
The video of two girls talking seems so natural. There are some artifacts but the movement is so natural and clothes and other things around are not continuously changing.
I hope it does become open source, which i suspect it won't because it's coming from byte dance.
These artifacts are an improvement over current state.
We have been conditioned to only react to hype and "news", rather than analyze reality and see the danger.
I complained to amazon, and they said since I hadn't purchased the book they couldn't do anything. So I bought the book, complained, and returned it. The chapters devoted to the details of the specific air fryer model were either very general (almost quotes of product description on amazon), or just plain wrong.
What I thought I would get would be like the magic lantern books about specific camera models. Instead it was auto-generated pages of nonsense.
I don’t think this form of generative AI needs to become a source of spam, carefully designed platforms can let people enjoy their niche content without making them feel isolated
Is being used to create spam is not the same as needs to be spam, and we mostly just need platforms that leverage generative AI natively to bridge the gap.