Stable Video Diffusion
stability.ai
stability.ai
I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue jays" would be near "toronto" or "cn tower". The improvements in scale and speed (image -> now video) are impressive, but given how incredibly able the image generation models are, they simultaneously feel crippled and limited by their lack of editing / iteration ability.
Has anyone come across a solution where model can iterate (eg, with prompts like "move the bicycle to the left side of the photo")? It feels like we're close.
https://emu-edit.metademolab.com/
https://ai.meta.com/blog/emu-text-to-video-generation-image-...
Also, maybe you can't edit post facto, but when you give prompts, would you not be able to say : two blue jays but no CN tower
I feel like we're close too, but for another reason.
For although I love SD and these video examples are great... It's a flawed method: they never get lighting correctly and there are many incoherent things just about everywhere. Any 3D artist or photographer can immediately spot that.
However I'm willing to bet that we'll soon have something much better: you'll describe something and you'll get a full 3D scene, with 3D models, source of lights set up, etc.
And the scene shall be sent into Blender and you'll click on a button and have an actual rendering made by Blender, with correct lighting.
Wanna move that bicycle? Move it in the 3D scene exactly where you want.
That is coming.
And for audio it's the same: why generate an audio file when soon models shall be able to generate the various tracks, with all the instruments and whatnots, allowing to create the audio file?
That is coming too.
The main issue is going to be having the right dataset. You basically need to record user actions in something like blender (ie: moving a model of a bike to the left of a scene), match it to a text description of the action (ie; "move bike to the left") and match those to before/after snapshots of the resulting file format.
You need a whole metric fuckton of these.
After that, you train your model to produce those 3d scene files instead of image bitmaps.
You can do this for a lot of other tasks. These general purpose models can learn anything that you can usefully represent in data.
I can imagine AGI being, at least in part, a large set of these purpose trained models. Heck, maybe our brains work this way. When we learn to throw a ball, we train a model in a subset of our brain to do just this and then this model is called on by our general consciousness when needed.
Sorry, I'm just rambling here but its very exciting stuff.
It can learn anything you have data for.
Heck, we do it with geospatial data already, generating segmentation vectors. Why not 3D?
Not in theory, but the level of complexity is way higher and the amount of data available is much smaller.
Compare bitmaps to this: https://fossies.org/linux/blender/doc/blender_file_format/my...
But we haven't gotten diffusion working well for text/code, so generating long files is a problem.
I'm not experienced enough to validate their claims, but I love the choice of languages to evaluate on:
> Python, Bash and Excel conditional formatting rules.
A 3D scene is vastly more complex, and the way you consume it is tangential to the rendering of it we use to interpret. It is a collection of arbitrary data structures.
We’ll need a new approach for this kind of problem
> A 3D scene is vastly more complex
3D scenes, in fact, are also data, numbers and tokens. (Well, numbers, but so are tokens.)
Not at all the same as fixed sized arrays representing images.
In fact, the data structures of a 3D scene can be serialized as text, and a properly trained text gen system could generate such a representation directly, though that's probably not the best route to decent text-to-3d.
Serializing 3D models as text is not going to work for negligibly non trivial circumstances.
I agree with this philosophy - Teach the AI to work with the same tools the human does. We already have a lot of human experts to refer to. Training material is everywhere.
There isn't a "text-to-video" expert we can query to help us refine the capabilities around SD. It's a one-shot, Jupiter-scale model with incomprehensible inertia. Contrast this with an expert-tuned model (i.e. natural language instructions) that can be nuanced precisely and to the the point of imperceptibility with a single sentence.
The other cool thing about the "use existing tools" path is that if the AI fails part way through, it's actually possible for a human operator to step in and attempt recovery.
I'm always confused why I don't hear more about projects going in this direction. Controlnets are great, but there's still quite a lot of hallucination and other tiny mistakes that a skilled human would never make.
More on the Blender file format: https://fossies.org/linux/blender/doc/blender_file_format/my...
The only trick is that there has to be enough Blender Python code to train the LLM on.
While it generates a lot of code that initially makes sense, when you use the code, you get a jumbled block.
But also I suspect there just isn't that much openscad code in the training data, and the semantics are different enough to python or any of the other languages that are well-represented that it struggles.
That pretty much matches my experience working with NN's and LLM's
LLMs are good at writing copy that sounds accurate and creative enough, and there are known techniques to improve that (such as generating an outline first, then generating each section separately). If you then give them a list of templates, and written examples of what they are used for, the LLM is able to pick one that's a suitable match. But this is all just probability, there's no real creativity here.
Earlier this year I played around with trying to have GPT-3 directly output an SVG given a prompt for a simple design task (a poster for a school sports day), and the results were pretty bad. It was able to generate a syntantically coreect SVG, but the design was terrible. Think using #F00 and #0F0 as colours, placing elements outside the screen boundaries, layering elements so they are overlapping.
This was before GPT-4, so it would be interesting to repeat that now. Given the success people are having with GPT-4V, I feel that it could just be a matter of needing to train a model to do this specific task.
Almost all of the generative 3D models you see are actually generative image models that essentially (very crude simplification) perform something like photogrammetry to generate a 3D model - 'does this 3D object, rendered from 25 different views, match the text prompt as evaluated by this model trained on text-image pairs'?
This is a shitty way to generate 3D models, and it's why they almost all look kind of malformed.
You wouldn't need any models to learn from. But my intuition is that RL is still quite weak, and that the model would flounder after learning to mimic background color and placing a few spheres.
Probably because they aren't as advanced and the demos aren't as impressive to nontechnical audiences who don't understand the implications: there’s lots of work on text-to-3d-model generation, and even plugins for some stable diffusion UIs (e.g., MotionDiff for ComyUI.)
For single 3D object the biggest dataset is ObjaverseXL with 10M samples
For full 3D scenes you could at best get ~1000 scenes with datasets like ScanNet I guess
Text2Image models are trained on datasets with 5 billion samples
Got a public dataset?
I've not worked in VFX for a while, but when I did the modeling departments at multiple studios had giant libraries of completed geometries for every project they ever did, plus even larger libraries of all the pieces and parts they use as generic lego geometry whenever they need something new.
Every 3D modeler I know has their own personal libraries of things they'd made as well as their own "lego sets" of pieces and parts and generative geometry tools they use when making new things.
Now this is just a guess, but do you know anyone going through one of those video game schools? I wager the schools have big model libraries for the students as well. Hell, I bet Ringling and Sheridan (the two Harvards of Animation) have colossally sized model libraries for use by their students. Contact them.
So we could have something to convert AI-generated image output into 3D scenes without having to explicitly train the "creative" AI for that.
Probably much more viable, because the quantity of 3D models out in the wild is far far lower than that of bitmap images.
The question is whether the 99% of the audience would even care...
As for "the internet", there will always some small part of it which will obsess and/or laught over anything, doesn't mean they represent anything significant - not even when they're vocal.
Perhaps a more computationally expensive but better looking method will be to pull all objects in the scene from a 3D model library, then programmatically set the scene and render it.
However, 3D is just one approach to rendering visuals. There are so many other styles and methods how people create images, and if I understand correctly, we can do image-to-text to analyze image content, as well as text-to-image to generate it - regardless of the orginal method (3d render or paintbrush or camera lens). There are some "fuzzy primitives" in the layers there that translate to the visual elements.
I'm hoping we see "editors" that let us manipulate / edit / iterate over generated images in terms of those.
In the end diffusion technology can make a more realistic image faster than a rendering engine can.
I feel pretty strongly that this pipeline will be the foundation for most of the next decade of graphics and making things by hand in 3D will become extremely niche because lets face it anyone who has worked in 3D it's tedious, it's time consuming, takes large teams and it's not even well paid.
The future is just tools that give us better controls and every frame will be coming from latent space not simulated photons.
I say this as someone who had done 3D professionally in the past.
Diffusion models aren't LLMs (they may use something similar as their text encoder layer) and they simulate their training corpus, which usually isn't selected solely for physical fidelity, because that's not actually the single criteria for visual imagery outside of what is created by diffusion models.
Emu can do that.
The bluejay/toronto thing may be addressable later (I suspect via more detailed annotations a la dalle3) - these current video models are highly focused on figuring out temporal coherence
This is not the flex you think it is. You don't have to like sports, but snarking on people who do doesn't make you intellectual, it just makes you come across as a douchebag, no different than a sports fan making fun of "D&D nerds" or something.
Rather than projecting your own hangups and calling people names, try instead assuming that they're not trying to offend you personally and are just using common vernacular.
Do the parameters think that Jazz musicians are mormon? Padres often surf? Wizards like the Lincoln Memorial?
Also, I originally tried to get the 3 characters in the image to be generated simultaneously, but eventually gave up as DALL-E had a hard time understanding how I wanted them positioned relative to each other. I just generated 3 separate characters and positioned them in the same image using Gimp.
Nearly all of the available models have this, even the highly commercialized ones like in Adobe Firefly and Canva, it’s called inpainting in most tools.
See video: https://www.adobe.com/max/2023/sessions/project-stardust-gs6...
Yeah. They're not "videos" so much as images that move around a bit.
This doesn't really look any better than those Midjourney + RunwayML videos we had half a year ago.
>Has anyone come across a solution where model can iterate (eg, with prompts like "move the bicycle to the left side of the photo")? It feels like we're close.
Google has a model called Phenaki that supposedly allows for that kind of stuff. But the public can't use it so it's hard to say how good it actually is.
As for your last question yes that exists. There are two models from Meta that do exactly this, instruction based iteration on photos, Emu Edit[0], and videos, Emu Video[1].
There's also LLaVa-interactive[2] for photos where you can even chat with the model about the current image.
[0]: https://emu-edit.metademolab.com/
I can’t wait to see what people do with this once controlnet is properly adapted to video. Generating videos from scratch is cool, but the real utility of this will be the temporal consistency. Getting stable video out of stable diffusion typically involves lots of manual post processing to remove flicker.
And that’s just for diffusion focused approaches like ours. There are probably other techniques from the token flow or nerf family of approaches close to breakout levels of quality, tons of talented researchers working on that too.
I have not, so I went to https://civitai.com/ which I guess is what you're talking about? But I cannot find a single video there, just images and models.
https://www.youtube.com/watch?v=3WWy98ylLT4
https://www.youtube.com/shorts/1vqOjYWEF84
https://www.youtube.com/shorts/jOIb9QbrhZ8
https://www.youtube.com/shorts/C3F_YI84TXA
https://www.youtube.com/shorts/4IqJHozY4F0
https://www.youtube.com/shorts/h3OmBLlm5-g
https://www.youtube.com/shorts/ZT7tuIgSDRk
https://www.youtube.com/shorts/WnUYbsOMyvs
https://www.youtube.com/shorts/BKKqX2aMlSg
The inconsistencies are what's most interesting in these videos in fact
Go there, in the top right of the content area it has two drop-downs: Most Reactions | Filters
Under filters, change the media setting to video.
Civitai has a notoriously poor layout for finding/browsing things unfortunately.
I ask as a noob in this area.
There’s been open source AI/ML for 20+ years.
Nothing comes close to the massive milestones over the past year.
Take AlexNet - the major "oh shit" moment in image classification.
It had an absolutely mind-blowing number of parameters at a whopping 62 million.
Holy shit, what a large network, right?
Absolutely unprecedented.
Now, for language models, anything under 1B parameters is a toy that barely works.
Stable diffusion has around 1B or so - or the early models did, I'm sure they're larger now.
A whole lot of smart people had to do a bunch of cool stuff to be able to keep networks working at all at that size.
Many, many times over the years, people have tried to make larger networks, which fail to converge (read: learn to do something useful) in all sorts of crazy ways.
At this size, it's also expensive to train these things from scratch, and takes a shit-ton of data, so research/discovery of new things is slow and difficult.
But, we kind of climbed over a cliff, and now things are absolutely taking off in all the fields around this kind of stuff.
Take a look at XTTSv2 for example, a leading open source text-to-speech model. It uses multiple models in its architecture, but one of them is GPT.
There are a few key models that are still being used in a bunch of different modalities like CLIP, U-Net, GPT, etc. or similar variants. When they were released / made available, people jumped on them and started experimenting.
SDXL is 6.6 billion.
The availability of GPU compute time. Up until the Russian invasion into Ukraine, interest rates were low AF so everyone and their dog thought it would be a cool idea to mine one or another sort of shitcoin. Once rising interest rates killed that business model for good, miners dumped their GPUs on the open market, and an awful lot of cloud computing capacity suddenly went free.
Emad Mostaque and his investment in stable diffusion, and his decision to release it to the world.
I'm sure there are others, but those are the two that stick out to me.
- Unsupervised learning techniques, e.g. transformers and diffusion models. You need unsupervised techniques in order to utilize enough data. There have been other unsupervised techniques in the past, e.g. GANs, but they don't work as well.
- Massive amounts of training data.
- The belief that training these models will produce something valuable. It costs between hundreds of thousands to millions of dollars to train these models. The people doing the training need to believe they're going to get something interesting out at the end. More and more people and teams are starting to see training a large model as something worth pursuing.
- Better GPUs, which enables training larger models.
- Honestly the fall of crypto probably also contributed, because miners were eating a lot of GPU time.
You're right that conditional generation start to blur the lines though.
While you're right about GANs, diffusion models as transformers as transformers are most commonly trained with supervised learning.
Unsupervised is a confusing term as there is always an underlying loss being optimized and working as a supervision signal, even for good old kmeans. But generative models are generally considered to be part of unsupervised methods.
Exactly. The growth in the next decade is going to be unimaginable because now governments and MNCs believe that there realistically be progress made in this field.
Of course those two releases didn't fall out of the sky.
Not that more advances don't happen with sustained hype, just there's some sort of tipping point involving usefulness based either on improvement of the thing in question or it's utility elsewhere.
The main utility will me misinformation
Right now, AnimateDiff is leading the way in consistency but I'm really excited to see what people will do with this new model.
Since videos are rarely consumed raw, what if this becomes a pipeline in Blender instead? (Blender the 3d software). Now the video becomes a complete scene with all the key elements of the text input animated. You have your textures, you have your animation, you have your camera, you have all the objects in place. We can even have the render engine in the pipeline to increase the speed of video generation.
It may sound like I'm complaining, but I'm just ask making a feature request...
What do you mean by “advocating”? The iPhone has had a LiDAR camera since 2020.
I don't think it's possible to invent copyright-like rights.
Copyleft is an example of someone successfully inventing a copyright-like right by bootstrapping off existing copyright with a specially engineered contract.
1) You and I invent our own private "copyright" for data (which is not copyrightable)
2) Everything is fine until my wife walks up to my computer and makes a copy of the data. She's not bound by our private "copyright." She doesn't even know it exists, and shares the data with her bestie.
And... our private pseudo-copyright is dead.
Also: Licenses are not the same as contracts. There are times when something can be both, one, or the other. But there are a lot of limits on how far they reach. The output of a program is rarely copyrightable by the author (as opposed to the user).
As you agreed to in our contract, you now need to compensate me for the damage caused by your failure to prevent unauthorized third-party access. Of course you're free to attempt to recover the sum you have to pay me from your wife.
> The output of a program is rarely copyrightable by the author (as opposed to the user).
The author of the program can make it a condition of letting the user use the program that the user has to assign all copyright to the author of the program, kind of like "By uploading any User Content you hereby grant and will grant Y Combinator and its affiliated companies a nonexclusive, worldwide, royalty free, fully paid up, transferable, sublicensable, perpetual, irrevocable license to copy, display, upload, perform, distribute, store, modify and otherwise use your User Content for any Y Combinator-related purpose in any form, medium or technology now known or later developed." https://www.ycombinator.com/legal/
1) You have a $1T product.
2) My wife leaks it, or a burglar does. I am a typical consumer, with say, a $20k net worth.
You have two choices:
1) Sue me, recover $20k, and be down $1T (minus $20k, plus litigation fees), and get the press of ruining the life of some innocent random person
2) Not sue me. Be down $1T (including the $20k) .
And yes, the author of a program can put whatever conditions they want into the license: "By using this program, you agree to transfer $1M into my bank account in bit coin, to give me your first-born baby, to swear fealty to me, and to give me your wife it servitude." A court can then read those conditions, have a good laugh, and not enforce them. There are very clear limits on what a court will enforce in licenses (and contracts), and owning the output of a program, and barring exceptional circumstance, courts will not enforce them:
https://www.lexology.com/library/detail.aspx?g=eb52567a-2104...
This is why programmers should learn basic law, not treat is as computer code, and consult lawyers when issues come up. Read by a lawyer, a license or contract with an unenforceable clause is as good as having no such clause.
It seems to me that the cases in the article you linked involved the author of the program arguing that their copyright automatically extended to the output without any extra contractual provisions concerning copyright assignment, so I don't think they can be used as precedent regarding the enforceability of such clauses.
I think it is quite likely a court would find that unconscionable.
At the end of the day, a license is a legal contract. If you agree that an image which you produce with some software will be GPL'ed, it's enforceable.
As an example, see the Creative Commons license, ShareAlike clause:
> If you remix, transform, or build upon the material, you must distribute your contributions under the same license as the original.
you can put whatever you want in a contract, doesn't mean it's enforceable
> In enterprise organizations (meaning those with >250 PCs or >$1 Million US Dollars in annual revenue), no use is permitted beyond the open source, academic research, and classroom learning environment scenarios described above.
First off, you have copyright law, which grants monopolies on the act of copying to the creators of the original. In order to legally make use of that work you need to either have permission to do so (a license), or you need to own a copy of the work that was made by someone with permission to make and sell copies (a sale). For the purposes of computer software, you will almost always get rights to the software through a license and not a sale. In fact, there is an argument that usage of computer software requires a license and that a sale wouldn't be enough because you wouldn't have permission to load it into RAM[0].
Licenses are, at least under US law, contracts. These are Turing-complete priestly rites written in a special register of English that legally bind people to do or not do certain things. A license can grant rights, or, confusingly, take them away. For example, you could write a license that takes away your fair use rights[1], and courts will actually respect that. So you can also have a license that says you're only allowed to use software for specific listed purposes but not others.
In copyright you also have the notion of a derivative work. This was invented whole-cloth by the US Supreme Court, who needed a reason to prosecute someone for making a SSSniperWolf-tier abridgement[2] of someone else's George Washington biography. Normal copyright infringement is evidenced by substantial similarity and access: i.e. you saw the original, then you made something that's nearly identical, ergo infringement. The law regarding derivative works goes a step further and counts hypothetical works that an author might make - like sequels, translations, remakes, abridgements, and so on - as requiring permission in order to make. Without that permission, you don't own anything and your work has no right to exist.
The GPL is the anticopyright "judo move", invented by a really ornery computer programmer that was angry about not being able to fix their printer drivers. It disclaims almost the entire copyright monopoly, but it leaves behind one license restriction, called a "copyleft": any derivative work must be licensed under the GPL. So if you modify the software and distribute it, you have to distribute your changes under GPL terms, thus locking the software in the commons.
Images made with software are not derivative works of the software, nor do they contain a substantially similar copy of the software in them. Ergo, the GPL copyleft does not trip. In fact, even if it did trip, your image is still not a derivative work of the software, so you don't lose ownership over the image because you didn't get permission. This also applies to model licenses on AI software, insamuch as the AI companies don't own their training data[3].
However, there's still something that licenses can take away: your right to use the software. If you use the model for "commercial" purposes - whatever those would be - you'd be in breach of the license. What happens next is also determined by the license. It could be written to take away your noncommercial rights if you breach the license, or it could preserve them. In either case, however, the primary enforcement mechanism would be a court of law, and courts usually award money damages. If particularly justified, they could demand you destroy all copies of the software.
If it went to SCOTUS (unlikely), they might even decide that images made by software are derivative works of the software after all, just to spite you. The Betamax case said that advertising a copying device with potentially infringing scenarios was fine as long as that device could be used in a non-infringing manner, but then the Grokster case said it was "inducement" and overturned it. Static, unchanging rules are ultimately a polite fiction, and the law can change behind your back if the people in power want or need it to. This is why you don't talk about the law in terms of something being legal or illegal, you talk about it in terms of risk.
[0] Yes, this is a real argument that courts have actually made. Or at least the Ninth Circuit.
The actual facts of the case are even more insane - basically a company trying to sue former employees for fixing it's customers computers. Imagine if Apple sued Louis Rossman for pirating macOS every time he turned on a customer laptop. The only reason why they can't is because Congress actually created a special exemption for computer repair and made it part of the DMCA.
[1] For example, one of the things you agree to when you buy Oracle database software is to give up your right to benchmark the software. I'm serious! The tech industry is evil and needs to burn down to the ground!
[2] They took 300 pages worth of material from 12 books and copied it into a separate, 2 volume work.
[3] Whether or not copyright on the training data images flows through to make generated images a derivative work is a separate legal question in active litigation.
Not necessarily; gratuitous licenses are not contracts. Licenses which happen to also meet the requirements for contracts (or be embedded in agreements that do) are contracts or components of contracts, but that's not all licenses.
The same way you can easily violate any "non-commercial" clauses of models like this one as private person or as some tiny startup, but company that decide to use them for their business will more likely just go and pay.
So it's possible to ignore license, but legal and financial risks are not worth it for businesses.
One smart thing they did was they'd check the online job listings and if a firm advertised for needing AutoCAD experience they'd check their licenses. I knew firms who got calls from Autodesk legal the DAY AFTER posting an opening.
> An image isn't GPL'd because it was produced with GIMP.
That's because of how the GPL is written, not because of some limitation of software licences.
It makes me think of the difference between ancestral and non-ancestral samplers, e.g. Euler vs Euler Ancestral. With Euler, the output is somewhat deterministic and doesn't vary with increasing sampling steps, but with Ancestral, noise is added to each step which creates more variety but is more random/stochastic.
I assume to create video, the sampler needs to lean heavily on the previous frame while injecting some kind of sub-prompt, like rotate <object> to the left by 5 degrees, etc. I like the phrase another commenter used, "temporal consistency".
Edit: Indeed the special sauce is "temporal layers". [0]
> Recently, latent diffusion models trained for 2D image synthesis have been turned into generative video models by inserting temporal layers and finetuning them on small, high-quality video datasets
[0] https://stability.ai/research/stable-video-diffusion-scaling...
So this example was posted an hour ago, and it's jumping all over the place frame to frame (somewhat weak temporal consistency). The author appears to have used pretty straight-forward text2img + Animatediff:
https://www.reddit.com/r/StableDiffusion/comments/180no09/on...
Fixing that frame to frame jitter related to animation is probably the most in-demand thing around Stable Diffusion right now.
Animatediff motion painting made a splash the other day:
https://www.reddit.com/r/StableDiffusion/comments/17xnqn7/ro...
It's definitely an exciting time around SD + animation. You can see how close it is to reaching the next level of generation.
Also, can someone benchmark it on m3 devices? It would be cool to see if it is worth getting on to run these diffusion inferences and development. If m3 pro can allow finetuning it would be amazing to use it on downstream tasks!
I’m the background section of the research paper they mention “temporal convolution layers”, can anyone explain what that is? What sort of training data is the input to represent temporal states between images that make up a video? Or does that mean something else?
I was working on similar idea few years ago using this paper as reference and it was working extremely well for consistency also helping with flicker. https://arxiv.org/abs/1811.08383
A good resource for the "instead" case: https://unit8.com/resources/temporal-convolutional-networks-...
The "also" case is an example of 3D convolution, an example of a paper that uses it: https://www.cv-foundation.org/openaccess/content_iccv_2015/p...
It's crazy to see this level of progress in just a bit over half a year.
[1]: https://epiccoleman.com/posts/2023-03-05-deforum-stable-diff...
Surely it will be written using machine vision and llms !
Back in the mid 90s to 2010 or so, graphical improvements were hailed as photorealistic only to be improved upon with each subsequent blockbuster game.
I think we're in a similar phase with AI[0]: every new release in $category is better, gets hailed as super fantastic world changing, is improved upon in the subsequent Two Minute Papers video on $category, and the cycle repeats.
[0] all of them: LLMs, image generators, cars, robots, voice recognition and synthesis, scientific research, …
Many more examples, of course.
Big quality improvement over Marathon 2 on a mid-90s Mac, which itself was a substantial boost over the Commodore 64 and NES I'd been playing on before that.
Whenever I saw anybody calling those graphics "photorealistic", I always had to roll my eyes and question if those people were legally blind.
Like, c'mon. Yeah, they could be large leaps ahead of the previous generation, but photorealistic? Get real.
Even today, I'm not sure there's a single game that I would say has photo-realistic graphics.
Looking just at the videos (because I don't have time to play the latest games any more and even if I did it's unreleased), I think that "Unrecord" is also something I can't distinguish from a filmed cinematic experience[0]: https://store.steampowered.com/app/2381520/Unrecord/
Though there are still caveats even there, as the pixelated faces are almost certainly necessary given the state of the art; and because cinematic experiences are themselves fake, I can't tell if the guns are "really-real" or "Hollywood".
Buuuuut… I thought much the same about Myst back in the day, and even the bits that stayed impressive for years (the fancy bedroom in the Stoneship age), don't stand out any more. Riven was better, but even that's not really realistic now. I think I did manage to fool my GCSE art teacher at the time with a printed screenshot from Riven, but that might just have been because printers were bad at everything.
IMO, though, the lighting in the indoor scenes is just not quite right. There's something uncanny valley about it to me. When the flashlight shines, it's clearly still a computer render to my eyes.
The outdoor shots, though, definitely look flawless.
On a more serious note, I don't think Roger Deakins has anything to worry about right now. Or maybe ever. We've been here before. DAWs opened up an entire world of audio production to people that could afford a laptop and some basic gear. But we certainly do not have a thousand Beatles out there. It still requires talent and effort.
As well as marketing.
Geordi: "Computer, in the Holmesian style, create a mystery to confound Data with an opponent who has the ability to defeat him"
Like, at this point, what are the technical counters to the assertion that our world is a simulation?
Let's steel-man — you mean 3D VR. Let's stipulate there's a headset today that renders 3D visually indistinguishable from reality. We're still short the other 4 senses
Much like faith, there's always a way to sort of escape the traps here and say "can you PROVE this is base reality"
The general technical argument against "brain in a vat being stimulated" would be the computation expense of doing such, but you can also write that off with the equivalent of foveated rendering but for all senses / entities
How about this theory is neither verifiable nor falsifiable.
It's not a very useful endeavour to worry about, but it can be fun to speculate about what might give rise to testable hypotheses and what that might tell us about the world.
If it would mean something drastic to you, I would be very curious to hear your preexisting existential beliefs/commitments.
People say this sometimes and its kind of slowly revealed to me that its just a new kind of geocentrism: its not just a simulation people have in mind, but one where earth/humans are centered, and the rest of the universe is just for the benefit of "our" part of the simulation.
Which is a fine theory I guess, but is also just essentially wanting God to exist with extra steps!
First off, there are zero technical proofs that we are in a sim, just a number of philosophical arguments.
In practical terms, we cannot yet simulate a single human cell at the molecular level, given the massive number of interactions that occur every microsecond. Simulating our entire universe is not technically possible within the lifetime of our universe, according to our current understanding of computation and physics. You either have to assume that ‘the sim’ is very narrowly focussed in scope and fidelity, and / or that the outer universe that hosts ‘the sim’ has laws of physics that are essentially magic from our perspective. In which case the simulation hypothesis is essentially a religious argument, where the creator typed 'let there be light' into his computer. If there isn't such a creator, the sim hypothesis 'merely' suggests that our universe, at its lowest levels, looks somewhat computational, which is an entirely different argument.
Very true, but to me this view of the universe and one's existence within it as a sort of second-rate solipsist bodge isn't a satisfyingly profound answer to the question of life the universe and everything.
Although put like that it explains quite a lot.
[Edit] There is also a sense in which the sim-as-a-focussed-mini-universe view is even less falsifiable, because sim proponents address any doubt about the sim by moving the goal posts to accommodate what they claim is actually achievable by the putative creator/hacker on Planet Tharg or similar.
Kind of like how video games won't render the full resolution textures when the character is far away or zoomed out.
I'm sure I'm not the first person to have thought this.
To an extent...
PS: Video is 2 years old, but still really impressive.
It's purely a religious question. When humanity invented the wheel, religion described the world as a giant wheel rotating in cycles. When humanity invented books, religion described the world as a book, and God as a it's writer. When humanity invented complex mechanism, religion described the world as giant mechanism and God as a watchmaker. Then computers where invented, and you can guess what happened next.
To be clear, this remains wicked, wicked, wicked exciting.
It was the same for image generation, where one needed to produce text prompts to create the image, and stuff like img2img and Controlnet that allowed things like controlling poses and inpainting, or having multiple prompts with masks controlling which part of the image is influenced by which prompt.
The input eventually becomes meanings mapped to reality.
I do not think so as the chance of constructing a fleshy eldritch horror is quite high.
There is a market for everything!
After all fine-tuning wouldn't take that long.
Diffusion models for moving images are already used to a limited extent for this. And I'm sure it will be the use case, not just an edge case.
The LICENSE is a special non-commercial one: https://huggingface.co/stabilityai/stable-video-diffusion-im...
It's unclear how exactly to run it easily: diffusers has video generation support now but need to see if it plugs in seamlessly.
They can be hacked into a Jupyter Notebook but it's really not fun.
I do wonder if there have been any codec studies that measure power usage with respect to RAM
I think this will really open new ways and new doors to creativity and creative expression.
I have googled for it but mostly just get low quality web tools.
Instance One : Act as a top tier Hollywood scenarist, use the public available data for emotional sentiment to generate a storyline, apply the well known archetypes from proven blockbusters for character development. Move to instance two.
Instance Two: Act as top tier producer. {insert generated prompt}. Move to instance three.
Instance Three: Generate Meta-humans and load personality traits. Move to instance four.
Instance Four: Act as a top tier director.{insert generated prompt}. Move to instance five.
Instance Five: Act as a top tier editor.{insert generated prompt}. Move to instance six.
Instance Six: Act as a top tier marketing and advertisement agency.{insert generated prompt}. Move to instance seven.
Instance Seven: Act as a top tier accountant, generate an interface to real-time ROI data and give me the results on an optimized timeline into my AI induced dream.
Personal GPT: Buy some stocks, diversify my portfolio, stock up on synthetic meat, bug-coke and Soma. Call my mom and tell her I made it.
For example, the man in the cowboy hat seems he is almost gagging. In the train video the tracks seem to be too wide while the train ice skates across them.
Let me know what you think of it! It works best on landscape images from my tests.
Don't get me wrong, this is insanely cool, but it's still a long way from good enough to be truly disruptive.
All of Hollywood falls.
As long as people can "clock" content generated from these models, it will be treated by consumers as low-effort drivel, no matter how much actual artistic effort goes in the exercise. Only once these systems push through the threshold of being indistinguishable from artistry will all hell break loose, and we are still very far from that.
Paint-by-numbers low-effort market-driven stuff will take a hit for sure, but that's only a portion of the market, and frankly not one I'm going to be missing.
CGI in films used to be obvious all the time no matter how good the artists using it, now it's everywhere and only noticeable when that's the point; the gap from Tron to Fellowship of the Ring was 19.5 years.
My guess is the analogy here puts the quality of existing genAI somewhere near the equivalent of early TV CGI, given its use in one of the Marvel title sequences etc., but it is just an analogy and there's no guarantees of anything either way.
weird logic circles yall keep making to justify your beliefs, i mean the world is very easy like you just described if you completely strip all nuance and complexity
people used to believe at the start of the space race we'd have mars colonies by now because they looked at the rate of technological advancement from 1910 to 1970, from the first flight to landing on the moon; yet that didn't happen because everything doesn't follow the same repeatable patterns
Second, I literally wrote the same point you seem to think is a gotcha:
> it is just an analogy and there's no guarantees of anything either way
If anything, a deluge of cheap AI generated movies is going to lead to a flight to quality. The big studios will be more powerful because they will reap the productivity gains and use traditional techniques to smooth out the rough edges.
People have been amenable to low grade garbage movies for a long, long time. See Adam Sandler's back catalog.
Actually, when processing power catches up, I'm expecting a movie engine with well-defined characters, scenes, entities, etc., so people will be able to share mostly text-based scenarios to watch on their hardware players.
Oh wait... No.
All it has done is create an environment where indy games are now assumed to be trash unless proven otherwise, making getting traction as a small developer orders of magnitude harder than it has ever been because their efforts are drowning in a sea of mediocrity.
That same thing is already starting to happen on youtube with AI content, and there's no reason for me to expect this going any other way.