Text2Video-Zero Code and Weights Released by Picsart AI Research (12G VRAM)
github.com
github.com
https://github.com/Picsart-AI-Research/Text2Video-Zero/blob/...
https://web.archive.org/web/20230329124453/https://github.co...
Check line 45
But hey, still pretty good for the early days. Maybe need to figure out how to prompt engineer it and tune it. Seems like it's very heavily based on existing image models? Wondering how easy it is to adapt to other image models. I think I need to read the paper. https://imgur.com/a/h3ciJNn
perhaps I dont understand this tool
It was a dachshund. Doing some sort of a dance in the air. Nothing resembling a flip.
Why are people downvoting this? The examples on GitHub are clearly doing the action described (unlike the dachsund above) and don't have the wild deformations that the "Spongebob hugging a cactus" or "Will Smith eating spaghetti" animations have.
Can't wait to try it though.
Seems fair to choose good results when you try to demonstrate what the software "can" do. Maybe a mention of the process often being iterative would be the most honest, but I think anyone familiar with these tools assume this by now.
To be fair, metrics for vision are limited and many don't know how to correctly inturpret them. You both want to know what's the best images the network can produce a well as the average. Reviewers, unfortunately, just want to reject (no incentive you accept). That said, it can still be a useful tool because you know the best (or near) that the model can produce. This is also why platforms like HuggingFace are so critical.
I do think we should ask more of the research community but just know that this is in fact normal and even standard. Most research judge by things like citations and just reading the works themselves. The more hype in a research area the more noisy the signal for good research. (This is good research imo fwiw)
If only people would just listen to your sage advice!
This kind of sentence makes it come across as if you are on a high horse talking down to the unwashed masses. You might want to avoid this.
They told you to unblock adblock to upload now.
PS: Hello from my high horse
My horse is also really high. We've been smoking weed together.
https://github.com/libredirect/libredirect
I don't know think it will work on mobile chrome, but for others like me that can't take imgurs bloat anymore, its great.
The weights are the fruition of training. They make the model actually useful.
- Model is the code
- Data is the text/images used for traing
- Weights are the training results
For example Lucene, models will be the java library, data is text data like wikipedia and weights are the lucene index. if you have all the 3 you can start searching right away, if you have model+data you have to generate the index which can take a lot of time, training/indexing take more than searching or using the model. if you have just he model you need to get your own data and run training on it
(LLaMa also released the weights for researchers and they leaked online, but the weights are not open-source.)
When a company releases a model including the weights, you can download their pre-trained weights and run inference with them, without having to train by yourself.
Image models work well at least: https://huggingface.co/docs/diffusers/optimization/mps
It is even possible to run them in the browser: https://stablediffusionweb.com/
corridor digital handled it by training their model on specific images of specific people - and so they effectively said "video of the panda called phil that we have trained you on images of"
Clearly this is not possible here - so I am missing how they got it close
If a workflow cannot support iterating on every detail in isolation with psychotic levels of control, it won't even be adopted by the VFX industry.
Kids at home will be making Star Wars, not Industrial Lights and Magic.
People who don't mind having output that's an aesthetic amalgam of whatever it has already ingested won't mind using these sorts of tools to generate work. For a movie studio that lives and dies on presenting imagery so precisely and thoughtfully crafted that it leaves a lasting impression for decades, I doubt it will be anything more than a tool in the toolkit to smooth things along for the people who've made careers figuring out how to do that.
I think there's two reasons this sort of thinking is so prevelant. A) Since developers are only really exposed to the tools-end of other people's professions, they tend to forget that the most important component is the brain that's dreaming up the imagery and ideas to begin with. Art school or learning to be an artist is a lot more about thinking, ideas, analyzing, interpreting, and honing your eye than about using tools... people without those skills who can use amazing generative tools will make smooth, believable looking garbage in no time flat from the comfort of their own living rooms. Great. B) Most people, especially those from primarily STEM backgrounds, don't really understand what art communicates beyond what it physically represents and maybe some intangible vibe or mood. Someone who really knows what they're talking about would probably take longer to accurately describe an existing artistic image than would be reasonable to feed to a prompt. Once again, that will be fine for a lot of people-- even small game studios, for example, that need to generate assets for their 3rd mobile release this month, but it's got miiiiles to go before it's making a movie scene or key game moment that people will remember for decades.
I'd hope so, but there are a lot of films churning out the next Michael Bay Transformers-type movie which also get the vast majority of VFX work and make the most revenue that upending the VFX industry might remove a lot of jobs.
Why do people say stuff like this? How do you know what machines will ever be able to do, can you tell the future? And more fundamentally, do you believe humans are more than a biological machine? If you don't (and thus aren't a physicalist) then sure, you can make statements like this because you'd believe that humans have something fundamentally different that no machine can replicate, but if you do think so, you'd believe that anything humans can do, machines will eventually be able to do, even if it takes a long time.
At first you seemed to be insinuating that AI will "upend vfx" because a lot fof movies are "like transformers" and now you're saying "never is a long time and what if AI becomes like a human"
What makes you think industrial light and magic or anyone in the vfx thinks this way? It has been one of the most competitive and rapidly changing industries in the last 40 years.
I'm completely done with this pointless pedantic argument against your lack of understanding.
And if you're done with this argument, as you've mentioned a few times over threads, I'm not sure why you continue to reply.
We're not even close to knowing what every part of the brain does, or even have a complete model of how individual biological neurons work, let alone know how we would replicate them, let alone have the potential for doing so in my career.
https://nautil.us/the-big-problem-with-big-science-venturesl...
Might this happen in a really, really long time? Maybe? We're sure as hell not going to do it with a cluster of GPUs. The actual functions of the neurochemical parts that drive emotion are one of the parts we know least about.
So I guess I'm in this for the long haul. What other aspect of this topic can you make bold declarations about without actually having checked if they're correct?
--------
EDIT: I once again can't reply because HN probably has better sense than I do.
The context of this conversation is the VFX industry. Your assertion was that people who work in the VFX industry shouldn't be surprised if AI takes their jobs. You've moved the goalpost continents from that argument, instead arguing that you're still theoretically correct because in some distant future, science will prevail in mimicking or creating biological machines that can perceive emotion.
Please re-read what I said. In the very first comment I made where I said "Maybe a subsequent generation that is vastly more precise, but this isn't even ballpark." And then after that I said that when we get to the stage where machines can do this, it'll philosophically be actual intelligence and not artificial intelligence. Pretending that's in the scope of this conversation is just daft. And then in this very comment where I mused that it might happen in a really long time. We both know that's not relevant to the current VFX industry and you're just not capable of admitting you're wrong.
Any other topics you'd like to pretend the discussion is about so you can pretend you're still correct?
You are again missing my entire point, which is to not say things like something is "never" going to happen if we already observe it happening. I'm not sure how many times I can repeat this point over and over. If you disagree that brains are not biological machines or that they can be replicated by humans through technology, just tell me now and we can stop, that is a fundamental difference that is unreconcilable in just an HN conversation.
An AI model doesn't get tired of receiving direction and doing retakes.
It's a matter of time.
[0] https://old.reddit.com/r/StableDiffusion/comments/1244h2c/wi...
Good for anything that is actually open source and doesn't close off research, like Text2Video.
You want to upend the VFX industry, you probably want text-to-3d-model and text-for-script-for-existing-animation-systems, that still supports either of those being finetuned manually with existing tools.
Where are you getting this idea?
And in 3 months, it will be upending the vfx industry...
Until the AI knows what you want and like by analyzing your browser and viewing history. There is a part in the Three Body Problem trilogy where the aliens create art that's on par if not better than human-made art, and humans are shocked because the aliens do not even resemble any sort of human intelligence at all.
No, art is whatever the artist wants to call art, and it's whatever people want to find meaning in. Sure, you can't equate it to producing images, but the vast majority of people won't really care, if they can see some cool images or videos then that's what they'll want. This is the same argument that was used when Stable Diffusion released, yet the inexorable wheel of progress turns on.
Your non-definition of art has no relevance to this discussion. You assume this is all very simple because you have no idea what you're talking about. You're engaged in the sort of discussion that inspired the dunning kruger research.
> the vast majority of people won't really care
If the vast majority of people won't care, why do all major movies and games spend tens or hundreds of millions of dollars on VFX when they could get lower quality versions of the same exact imagery for 1/20th the budget like SyFy productions do? It provides a huge ROI because people care. If your taste and/or perception isn't sophisticated enough for it to matter, you can't just assume that's the case for everybody else. Truthfully, you almost certainly do have the perceptual sophistication to care-- you just haven't spent much time breaking down all of the incredibly important details that escape your notice that heavily influence your impression of the end product.
There's no shame in not knowing something, but there is shame in the hubris of trying to explain that thing to subject matter experts.
----
EDIT: I can't reply to your comment but it doesn't matter. You're clearly way outside of your depth and I have no interest in playing along to help you avoid having to confront that. Bye.
Not really. Don't assume you're the only artist on this site, there are others too. Why don't you provide a valid definition of art then? Then we'll use that one for this discussion.
> If the vast majority of people won't care, why do all major movies and games spend tens or hundreds of millions of dollars on VFX when they could get lower quality versions of the same exact imagery for 1/20th the budget like SyFy productions do?
That's...my point. This is basically what I said (or perhaps what I meant to say, if that didn't come across clearly). Then you said,
> Creating images is a vastly different process than creating art though creating art might involve it. If you think art is merely customizing imagery things to suit people's preferences, you don't understand art.
I understood this to mean that you see something beyond creating images (ie, fancy VFX) and that there is some deeper meaning of "art." My point is that as long as people see pretty, expensive pictures on the screen, many of them won't care about some ideal artistic merit. See how much money Avatar or Transformers make over generally more highly-regarded artistic films like Everything Everywhere All At Once which didn't make nearly as much.
Either way, saying someone is "way out of their depth" simply because they disagree with your preconceived notions really isn't a way to communicate, if so, just shut yourself off to any opinions at all. To onlookers, it just seems like an excuse to not engage in meaningful discourse.
I'm saying you're way out of your depth because you're treading the philosophical ground most committed artists are sick of within a few years of adolescence, and your assertions about the effect and creation of visuals in entertainment show even less understanding than that. I've had a billion conversations like this with developers, engineers, etc. so enamored with their own ability to reason that they incorrectly assume it gives them universal expertise.
I'm well beyond the age where I feel the need to explain basic aspects of something I have professional expertise in because someone insists their lack of knowledge is just as valid as my hard-won education and experience. So if you want to keep arguing about it, go ahead. You're just going to do it by yourself.
* lol. if you don't think there's a need to edit wonky in then you haven't seen midjourney.
Besides, humans have the power of boredom and so it's not possible to create "perfect" images that would hold their attention forever.
(Although a lot of "recommendation algorithms" are actually spam fighting, not engagement promoting.)
Depends on how well they work, as you say, TikTok is a big one where I doubt many people want their feed in a chronological manner instead. Technical people especially seem to dislike algorithmic feeds but for the layperson, they don't really mind it (indeed, they don't even think about it at all, based on personal experience some of the people I've talked to don't even know that chronological feeds are even a thing).
There will be a learning curve.
Do I run pix2pix at the end?
Also can I somehow have more frames and set the GIF speed?
To create a single frame with Stable Diffusion 4-5B parameters, 512x512, 20 iterations takes 5-30 minutes depending on your CPU. On any modern GPU it's only 0.1-20 seconds!
Similarly with LLMs, to produce one token with a transformer of 30-50 layers 7-12B parameters you will wait several CPU minutes while it takes few seconds on a Pascal-generation GPU and tiny fraction of second on Ampere.
Depends a lot on the cpu. Are you specifically talking about Text2Video or SD in general? IIRC, last time I tried SD on my CPU (10 core 10850k, not exactly cutting edge) it did take less than one minute for more than 20 iters. This was about 4-5 months ago, things might have gotten better.
The GPU (even a vintage 1070) was faster still of course.