Stability.ai – Introducing Stable Video 3D
stability.ai
stability.ai
We know that a single image of an object physically can't cover all the sides of it, so it's all guesswork in AI. This is totally fine for certain scenario, but in lots of other cases, it's trivial to have multiple images of the same object, and if that offers higher fidelity, it's totally worth it.
I'm aware there are many algorithms or AI models that already do that. I'm asking about Stability's one specifically because if they have impressive Single Image result, surely their multi-image results would also be much better than state-of-the-art?
[edit]Managed to generate by tweaking the script to generate less frames simultaneously. 19.5GB peak VRAM usage, 1 min 25 secs to generate at 225 watts.[/edit]
Traditional CPU-bound physics/simulation models have typically wanted all the RAM they could get; the more RAM the more accurate the model. The same is true for AI models.
I can max out 24GB just using spreadsheets and databases, let alone my 3D work or anything computational.
Stable diffusion, stable video, text models, audio models, I never had issues with anything yet
There's a lot of open weights activity around 7B/13B models which the 4090 will run with ease. But you could can run those OK on much cheaper cards like the 4070Ti (which is of course why they're popular).
And there's a lot of open weights activity around 70B and 8x7B models which are state-of-the-art - but too big to fit on a 4090. There's not much activity around 30B models, which are too big to be mainstream and too small to be cutting edge.
If you're specifically looking to QLoRA fine-tune a 7B/13B model a 4090 can do that - but if you want to go bigger than that you'll end up using a cloud multi-gpu machine anyway.
Nvidia have held back the majority of their cards from going over 24GB for years now. It's 2024 and my laptop has 96GB of RAM available to the GPU but desktop GPUs that cost several thousands just by themselves are stuck at 24GB.
This is like Intel and their refusal to support ECC memory; when AMD does on nearly all Ryzens.
—
Note: your laptop is probably using a 64-bit memory bus for system RAM. For GPUs, the 4090 is 384-bit. That takes up a lot more die area for the bus and memory controller.
I guess my point is, rather than give the cards more RAM, the gaming cards should just be priced cheaper.
5090 likely won't have more than 32 GB, if even that much.
0. https://manifold.markets/Tenoke/how-much-vram-will-nvidia-50...
https://github.com/vosen/ZLUDA
That's a binary level wrapper. Of course there's also ROCm HIP at the source level, and many other things, such as SYCL
[1] https://www.tomshardware.com/pc-components/gpus/chinese-work...
No matter how much vram you have, there's something that doesn't fit :)
I have 3090 Ti and I can run Q4 quant 33b models at 30t/s with 8k context. A 4090 would allow me to do the same but with ~45t/s, both inference speeds are more than fast enough for people so 3090 is the usual choice. In my tests on runpod, H100 with 80GB memory is around the same speed as 3090, so slower than a 4090.
Odd statement. I don't really know what you mean by that. Perhaps 'math _works_, code should too' ?
I would definitely agree that it _should_ work.
I'm of the belief that no one should _have to_ publish (e.g. to graduate, get promotions, etc) in academia, and that publications should only occur if they're believed to be near Novel prize worthy, and fully reproducible by code with packaging that should last and work in 10 years, from data archives that will exist in 10 years.
But it seems I have been outvoted by the administration in academia.
Hence, we get this "ai that doesn't run" phenomenon
Do you want to publicly fund researchers only for the industrial research partner's benefit?
My main point was that there is a lot of noise in scientific journals that are caused from pressures in academia that are requirements if publishing. If these are removed, then the quality of work published increases and quantity decreased.
There are other places to post work that is derivative and non-novel like blogs. The field of biology has an immense amount of work that is mostly observational without strong conclusions or predictivity. A tabulation of observation should definitely be put out by a lab, and it should be much sooner with far less pressures than today, such as the typical dance of putting the data in during publication. The SRA is one example of a place to share data. If the typical way to work was put all data immediately onto a public repo, sometimes comment on it in ways that have been seen before on blogs and other classes below scientific journals, and then if something truly substantial comes out of it (a novel model that is analytical and highly predictive of cell behavior in all situations for example) then publish.
It could alleviate the noise from the signal. LLMs is one case where the noise is very strong in that many papers are simply 'we fine tuned an llm'.
That said, I certainly think that researchers can do more to make their code and data more accessible. We have the tools to do so already but the incentives are often misaligned.
> Odd statement. I don't really know what you mean by that. Perhaps 'math _works_, code should too' ?
It was a typo. I mean "Everything should run in it" as in most LLM should be able to run at least quantized in 4090
Looking forward to experimenting with this.
I think a good hobbyist application for this would be something like modelling figurines for games, which is already a pretty popular 3D printing application. This would allow people with limited modelling skills to bring fantastical, unique characters to life “easily”.
Another way of looking at it, 3D artists often begin projects by taking reference images of their subject from multiple angles, then very manually turning that into a 3D model. That step could potentially be greatly sped up with an algorithm like this one. The artist could (hopefully) then focus on cleanup, rigging, etc, and have a quality asset in significantly less time.
There are other steps to 3D printing in general, though; a super rough outline:
- Model generation
- "Slicing" - processing the 3D model into instructions that the 3D printer can handle, as well as adding any support structures or other modifications to make it printable
- Printing - the actual printing process
- Post-processing - depending on the 3D printing technology used, the desired resulting product, and the specific model/slicing settings, this can be as simple as "remove from bed and use" to "carefully snip off support structures, let cure in a UV chamber for X minutes, sand and fill, then paint"
As I said before, this AI model specifically would cover 3D model generation. If you were to use a printing technology that doesn't require support structures, and handles color directly in the printing process (I think powder bed fusion is the only real option here?), the entire process should be fairly automatable - a human might be needed to remove the part from the printer, but there might not be much post-processing to do.
The rest of your desired workflow is a bit more nebulous - I don't know how you would handle "scanning what teens are doing on instagram", at least in a way that would let you generate toys from the information; generating and posting the advertisement shouldn't be too hard - have a standardish template that you fill in with a render from the model, and the description; printing on demand again is possible, though you'll likely need a human to remove the part, check it for quality and ship it. You could automate the latter, but that would probably be more trouble than it's worth.
On finding out what teens want, that part is somewhat easy-ish, I guess you'd need a couple of agents, one that is scanning teen blogs for stories and then converting them to key words, then another agent that takes the key words (#taylorswift #HaileyBieberChiaPudding #latestkdrama etc) into Instagram, after a while your recommend page will turn into a pretty accurate representation of what teens are into, then just have an agent look at those images and generate difs of them. I doubt it would work for a bunch of reasons, but it's an interesting thought experiment! Thanks!
What they show in the demo: https://i.imgur.com/9bZNTcd.jpeg
What comes out of the 3D printer: https://i.imgur.com/MZrzsfh.png
When I see a demo where they are showing wireframes I know it’ll be good enough.
It's proposing a solution to the author's observation that everyone is doing it in second order fashion and missing a significant amount of necessary data.
The implication is that rather than doing it the hard way via the already-obtained 2nd order dataset, it'll be easier to get a new dataset, and getting that dataset will be significantly easier that it was to get the second-order dataset, as you don't need to worry about aesthetic variety as much as teaching what level of detail is needed in the mesh for it to be "real"
There aren't a bajillion high-quality 3D models of everything, but there are an unbounded number of high-quality 3D models of some things, due to the existence of procedural mesh systems for things like foliage.
You could, at the very least, train an ML model to translate images of jungles into 3D meshes of the trees composing them right now.
Although I wonder if having a few very-well-understood object types like these, to serve as a base, would be enough to allow such a model to deduce more generalized rules of optics, such that it could then be trained on other object categories with much smaller training sets...
I’d imagine it’d require a ton of tagging, although I have a good idea of how I could leverage existing APIs to tag it mostly automatically by generating three still image thumbnails of the content, then feeding that through CLIP, and verifying that all two or three agree on what it’s an STL of, and manually tag the ones that fail that test.
(I dream of the day when this can be used to automatically create paper-craft templates.)
For games at the very least you need to consider polygon budget, getting reasonably good UVs, and generating materials which fit into a PBR shader pipeline, at least if it's going to work with rendering pipelines as we know them today (as opposed to rendering neural representations directly, which is a thing people are trying to do but is totally unproven in production).
So yes, there might be a wooden frame in the middle of that window, but does it match the math on both angles of it? Doubt it.
You could set something up on RunPod or AWS, but I doubt it's worth the effort.
It does look like SV3D is not a part of the API currently, but only a matter of time I imagine.
EDIT> Does "single image inputs" mean more than one image?
Does this output an actual 3D mesh? Or does it only output a 3d-looking rendered animation?
So can it actually output a 3d model? Or just images of what it thinks the object would look like from other angles?
OTOH, with their shift to a less open licensing structure, community tooling probably won’t emerge with the same level of energy.
Basically, cool looking video no one watches. I say that as a huge fanboy/artist myself. It is like Christmas every other day right now. All this VC money being set on fire to make better Avante-garde film tools is just wonderful. A dream come true.
Or making 3d game assets from objects you have around. Imagine: take your phone, go around town, into shops, into churches, come back, press a button, get huge library 3d assets to populate your game.
Or, something like this, for IKEA: couple of photos of a room --> extract objects --> let user re-arrange furniture. The room could be either the user's room or an IKEA showroom.
You can do it with existing tools, but this kind of technology reduces it to pressing a couple of buttons.
How would it handle other objects? (People, fabrics, buildings, plants, mountains, mechanical parts, etc)