Toon3D: Seeing cartoons from a new perspective
toon3d.studio
toon3d.studio
Non-photo-realistic (NPR) 3D art goes back a surprisingly long way in animations. I rewatched the 1988 Disney cartoon "Oliver and Company" recently, and I was surprised to see that the cars and buildings were "cel-shaded" 3D models. I assumed that the movie had been remastered, but when I looked it up, I found out that it was the first Disney movie ever to make heavy use of CGI[0] and that what I was seeing was in the original. The page I found says:
"This was the first Disney movie to make heavy use of computer animation. CGI effects were used for making the skyscrapers, the cars, trains, Fagin's scooter-cart and the climactic Subway chase. It was also the first Disney film to have a department created specifically for computer animation."
References ----------
https://m.youtube.com/watch?v=mix9rStOqoI
Now I am curious to watch it
Tron came out 1982, six years before Oliver & Company.
[1] https://filmschoolrejects.com/tron-costumes-glowing-effect/ Thanks legions of Taiwanese animators (:
Sounds pretty heavy to me.
>Eleven minutes of the film used "computer-assisted imagery" such as the skyscrapers, the taxi cabs, trains, Fagin's scooter-cart, and the climactic subway chase
I think Tron wins in terms of CGI
https://youtu.be/P5hHV2torG0?t=126
Wait, you're telling me that computers have enabled us to have fewer artists and thereby replacing artists for a long time now?!
Just like pretty much every industry out there?!
And that it's widely accepted so long as people get their cheap plastic goods from China?!
And that the current outrage won't even be remembered in 20 years?!
Which tells, AI hatred don't necessary come from what pro-AI thinks where it comes from, people potentially just find AI art rage inducing.
Like, not even specific technical aspect of AI is bad or could use improvements. It just sits at the wrong side of the uncanny valley, and arguments clump around that.
That's the real problem with generative AI.
I remember seeing this blog write up on what 3D animators do to make things look acceptable. Like make a character 9 feet tall because when the camera panned them, they looked too short at their "real" in-system height. Or archway doors that are huge but at the perspective shot, look "normal" to us. Or having a short character stand on an out-of-scene blue box to make them having a conversation with a tall character not look silly due to an extreme height difference? Or a hallway that in real life would be 1,000 feet long but looks about 100 in-world because of how the camera passes past it, and how each door on that 1,000 foot hallway is 18 feet high, etc.
I wonder if shows like Futurama used those tricks as well, so when you sort of re-create the 3D space the animators were working in by reverse engineering like this, then you see the giant doors and 9 foot people and non-Euclidian hallways, etc. Just because it looks smooth as the camera passes it, doesn't mean that actual 3D model makes sense at other perspectives.
[0]: https://www.gameinformer.com/b/news/archive/2013/11/20/the-t...
Maybe similar kind of things could be an application of this.
Another possible use-case might be a game development studio developing a license game based on a 2d cartoon, but making the game 3d. They could use this as a tool for visualization while planning and developing, to iterate quickly and to reference how the original 2d could translate into 3d.
Team of 2d artists draw the desired vehicles for the cartoon from two or three angles. Software like this makes a usable 3d model of it.
Even if AI has a place in the 2D to 3D part of a pipeline, surely you'd still want the 2D artwork be unambigously representative of what the 3D asset should look like, rather than providing self-contradictory input data and praying that the AI can magically make it make sense.
If you do want numerous 2D artworks which share a realistically defined 3D space then that can easily be done by making a very rough 3D scene and then painting over it, you don't need any AI for that.
While rough, these do look better than some implementations of the artwork for cartoon games.
On the other hand you are probably better off only using the depth prediction and filling any voids in using image generation instead of this mapping process.
I could see it being used to create a "scratch track" that human animators animate on top of. An aid to tweening.
I tried a crude and terrible version of something like this a few years ago, but not just inconsistent spaces without a clear ground truth - purely abstract non-space images which aren't supposed to represent a 3D space at all. Transform an abstract art painting (Kandinsky or Pollock for example) into a explorable virtual reality space. Obviously there is no 'ground truth' for whatever 'walking around inside a Pollock painting' means - the goal was just to see what happens if you try to do it anyway. The workflow was:
1. Start From Single Abstract Art Source Image
2. SinGan to Create Alternative 'viewpoints' of the 'scene'
3. 3d-photo-inpainting (or Ken Burns, similar project) on original and SinGan'd images (monocular depth mapping, outputs a zoom/rotate/pan video)
4. Throw 3d-photo-inpainting frames into photogrammetry app (Nerf didn't exist yet) and dial up all the knobs to allow for the maximum amount of errors and inconsistency
5. Pray the photogrammetry process doesn't explode (9 times out of 10 it crashed after 24 hours, brutal)
I must have posted an example on Twitter but I can't find the right search term to find it. But for example, even 2019 tier depth mapping produced pretty fun videos from abstract art: https://x.com/jonathanfly/status/1174033265524690949 The closest thing I can find is photogrammetry of an NVIDIA GauGAN video (not consistent frame to frame) https://x.com/jonathanfly/status/1258127899401609217
I'm curious if this project can do a better job at the same idea. Maybe I can try this weekend.
Well just in case it wasn't obvious, Toon3D, the project being discussed, is doing that. Part of the workflow is asking the user to indicate correspondences between geometry in different images, and each image is processed individually to create blocks of geometry you can toggle on or off visually.
Older projects:
https://github.com/sniklaus/3d-ken-burns
https://github.com/vt-vl-lab/3d-photo-inpainting
I believe there are some NeRF variants that do something like this from a single image as well, but I haven't personally tried any.
In the end (from my superficial) understanding, the problem with porting anything into VR (say in Unity in which you can walk around an object) is the important of creating a clean mesh. The 3D model that tools such as OP (I haven't dived deep into it yet) is these are point cloud in 3D space. They do not generate a 3D mesh.
Going from memory from tools I came across during my research, there is tools like this https://developer.nvidia.com/blog/getting-started-with-nvidi..., again, this does not generate a mesh. I think it is just a video and not something you can simply walk around in VR.
My low key motivation was to make a clone/model like what Matterport and sell it to real estate companies. Major gap in my understanding - the cause of me to loose steam is - I was not sure how are they able to automate the step to generate clean mesh from bunch of photos from a camera. To me, this is the most labor intensive part. Later, I heard there are ML model that is able to do this very step, I have no idea on this tho.
Are you saying it can take a point cloud 3D representation into a fully working and clean 3D mesh for VR?
My understanding is that essentially these technologies take some images as input, and then train a model, where the model is learning the best way to render the imagines into a model in some sense. I think for gaussian splats, it represents images as sort of "blobs" in space, and each image has the same set of blobs that have to be used from some perspective to render the image, hence by positioning the splats such that each image is rendered correctly, you can reproduce the scene.
This training is currently very expensive and has to be done for each model, but produces an output that can be explored in real time.
I think the photogrammetry approaches used by matterport et all are older and require much higher quality input data, whereas the newer approaches can work with much less and lower quality data.
https://github.com/3DTopia/OpenLRM (They mention NeRF as inspiration but it seems original paper it was based on decided to use visual transformers. the opensource version seems to use meta's dino as one of key components)
Or if a 2D artist could sketch a couple of poses and automatically get a well structured 3D model and textures?
I think there's been a lot of concern in the industry about the impact AI and similar tools will have on artists, but it seems like it's possible to imagine a future where machine learning based systems work more directly with an artist rather than rendering based on language etc.,
I don't know how I feel about all the moral arguments about AI training etc.,. I think to me more concerning is how it could impact people more so than how it was trained. Even if a perfectly "ethically" trained model learned to produce perfect art and artists became a niche field, I think it could still be a bad outcome for civilization as a whole because I think there's value in humans producing art, and in having a society where it's (at least somewhat) of a sustainable field.
Otoh, I think it's amazing that people can produce the kinds of images using image models, so I'm not sure. Ideally we'd be able to support people in what they want without needing their to be a market for it, but the world's not ready for that.
However, the "messy" reconstructions of 3D space seen in these videos did make me think of the recent hype over LLMs.
That is, the representations have a clear link to the "truth" or "facts" of the underlying material, but are in no way accurate enough to be considered useful as source material for further use.
An algorithm that needs to look at the lines and try to figure out a real-world scenario that correlates to that representation might be trying to create something that could never exist in any coherent form.
So, not hyperbole.
Someone with an extreme sensitivity for kindness can easily be seen as a curmudgeon by others, ie, after long years of disillusionment with the human race, or after a traumatic experience, or simply because of how they look.
Some people might be good at being kind 'in the moment', while others need to reflect - and the second kind can be a 'bigger', more encompassing, more effective or beneficial kindness.
And many (all?) of the people who give the most of themselves without hope for any reward genuinely care nothing for external validation or recognition - meaning we don't often hear about them or recognize them.
One could garner a reputation as an absolute arse, while accomplishing fantastically beneficial changes in the world. And conversely, a man could get a reputation as a folksy down-to-earth guy who you'd love to have a beer with, even as he sets the planet on a course to perpetual war. Cough.
Quoting Miyazaki, which was not especially harsh given they showed him a naked mutant zombie crawling across the ground using it's head and arm as legs while constantly trying to arch it's butt toward the camera.
> Every morning, not recent days, but I see my friend who has a disability.
> It's so hard for him just to do a high five (waves hand showing difficulty)
> His arm with stiff muscle reaching out to my hand (demonstrates body stiffness)
> Now thinking of him, I can't watch this stuff and find it interesting
> Whoever creates this stuff has no idea what pain is, or whatsoever. I am utterly disgusted.
> If you really want to make creepy stuff, you can go ahead and do it
> I would never wish to incorporate this technology into my work at all
> I strongly feel that this is an insult to life itself.
(room sits in silence awkwardly)
From a Review by Matthewmatosis [1]:
> Not long after setting out, I found myself staying in a quiet place, just moving Slugcat around various obstacles as smoothly as I could. [...] What was happening on screen looked like an animal testing its limits so as to build survival skills. It was then that I knew that this system was a resounding success.
> The hand-drawn images are usually faithful representations of the world, but only in a qualitative sense, since it is difficult for humans to draw multiple perspectives of an object or scene 3D consistently. Nevertheless, people can easily perceive 3D scenes from inconsistent inputs!
It is difficult for human artists to maintain perfect geometrical consistency. But that is NOT why 2D animation of 3D scenes is geometrically inconsistent! The reason is that artists stylize 3D scenes to emphasize things for specific artistic reasons. This is especially true for something surreal like SpongeBob. But even King of the Hill has stylized "living room perspectives," "kitchen perspectives," etc. The artists are trying to make things look good, not realistic. And they aren't trying to make humans reconstruct a perfect 3D image - they are trying to evoke our 3D imaginations. It's a very different thing.
Pixar and other high-quality 3D animation studios intentionally distort the real geometry of their scenes for cinematic effect: a small child viewed from an adult's perspective might be rendered with a freakishly long neck and stubby little torso, because the animators are intentionally exaggerating visual foreshortening to emphasize the emotional effect of a wee little child. A realistic perspective would be simply boring. These techniques are all over the place in Pixar movies - it's why their films look so good compared to cheaper studios, who really are just moving a virtual camera around a Euclidean 3D space.
I don't want to comment on the technical details. But it really seems like the authors missed the artistic mark.
In addition to all the noise and haze -- so the intermediate frames wouldn’t be usable alongside the originals -- the start and end points of each element hardly ever connect up. Each wall, door, etc flies vaguely towards its destination, but fades out just as the “same” element fades in at its final position a few feet away.
It’s a lovely idea, though, and it would be great to see an actually working version.
Gaussian Splatting is, in my opinion, simply the wrong tool for geometrically inconsistent images even if you manually annotate a bunch of keypoints. Another thing is that the spherical harmonics color representation makes it easy for the model to "cheat" when there are relatively few views, i.e. even when the Gaussian is completely geometrically wrong, it can still show the right color in the directions of the views. Perhaps they should have just disabled the spherical harmonics thing (i.e. making each Gaussian the same color regardless of which direction you're looking at it from), since most cartoons have flat-ish shading anyway.
Furthermore, they didn't attempt any sort of photometric calibration or color estimation between the different views. For example the paintings each show the building in a very different lighting condition, and it seems they made no attempt to handle that at all, leading to a very ugly reconstruction.
Finally, this method requires significant amounts of human work to do the manual annotation for a very subpar result, making us wonder what the whole point of it is. It would seem to me that diffusion models like Sora or Veo could do a much better job if you just want to interpolate between different views. It isn't much different from image inpainting, which diffusion models excel at.
I find this premise unsound. The reason is less that it is difficult but more that it is undesirable - in this medium.
I'm wondering if it's at all useful in understanding / improving AI's ability to infer semantic meaning from even real images in a variety of scenarios? Like the ability to re-interpret an interpreted construction (drawing) of a scene.
One area of application may be helping machines better understand hand drawn human input?
Is this widespread? My sense is that most mainstream TV animation that isn't obviously CGI is still drawn in 2D, with 3D work if used at all being relegated to backgrounds and the like.
I feel like this type of thing best applies in the kind of domain they're already in— TV shows with hundreds of hours of content that a machine can comb through looking for reference images to synthesize into these models.
What do I do with this data exactly? Not really following the instructions from README
Do I need a hefty GPU to run this? Doesn't say anything about hardware.
What am I going to get as a result? Will it generate a 3d model or "point clouds" ?
Do I need multiple inputs (from different angles) through the labeler?
What is the depth estimator being used here (this im most interested in especially its able to detect ground from multiple angles) ?
Guess I'm just really lost here but super eager to use this. We do have a real world application to use this.
There are many others: Kuula, Cupix, iStaging, EyeSpy360... Real estate companies use them a lot, e.g. to create a virtual tour for prospective buyers.
It doesn't seem like the OP comes even close to this though.