Video Vectorization
vectorly.io
vectorly.io
I guess these days we have animated SVG, and https://lottiefiles.com/ is getting some traction - but these require you to export in a specific format of course, you can't just convert/trace a bitmap movie with these. And SVG or Lottie aren't designed for longer/streaming vector animations, and they don't carry synchronised or streaming audio - Flash did all of those things.
Vectorisation of bitmap images does have some artifacts, as is evident in the Simpsons demo on this website - when possible, you should export in a vector movie format directly from the vector animation software.
It is kind of depressing that we don't have an open standard for vector movies (with sound), over a decade after Flash was killed. Sometimes it feels like technology stops or moves backwards.
The idea of a vector codec isn't new though, and prior attempts have been made, though I don't remember enough details to find a reference.
Just for anyone who might not be familiar, "vector" in this context refers to the graphics you would generate in a tool like Inkscape, where you can define complex geometric shapes with just a few datapoints. "Raster" graphics are what Gimp works with, where each pixel is specified (effectively) individually.
The problem is that you could have those concepts in a video codec, but then you'd be in some sort of weird hybrid world. A general encoder would almost certainly never use most of those transforms and primitives. Meanwhile, the decoder would need to understand both to properly render anything. I'd imagine you'd frequently run into cases where the decoder fails simply because whoever wrote it didn't want to be bothered supporting the full set of features for the steam (since encoders don't employ that full feature set).
Then there is the whole danger of having a turing complete video stream format. Last thing you'd want is for someone to publish a bitcoin miner on a youtube video :)
https://www.meetup.com/SF-Video-Technology/events/zpltdrybck...
There are already vector-graphics runtimes like WebGL and SVG that are more than capable of rendering "video", even within the html video tag, and through modern video streaming architectures like Dash and HLS (will share the link to the talk, and demo links for these things next week).
Those aren't "codecs" in the traditional sense, and I think there is an open question as to whether a "codec" is even necessary. Scrimba (https://scrimba.com/) uses "HTML" as a video codec just fine in production, and it works perfectly, and there's no "codec" per se behind it.
That said - in 2020, web architectures already exist in a way that you could easily make a "video codec" for vector graphics, and some standardization would help - thought not specifically necessary for adoption in the way it would be for a "regular" video codec which doesn't just enjoy native vector graphics runtimes like SVG or WebGL.
If you needed a file format, I honestly think that Lottie is the best option for a "file format" for vector-video, since it's open source, and mostly just based on the old Flash standard (https://www.adobe.com/devnet/swf.html)
What's missing from Lottie for it to be a "video" format would be easier integration into video streaming architectures like HLS or DASH, and that's honestly something I'd like to do as an open source project - essentially a way for video players to "play" lottie video files as one of the video options (like, on top of 1080p, 720p and 480p versions, you have the vector version as well)
From the perspective of Vectorly, we clearly understand that there's a lot of skepticism around the idea of a new "codec", and we wanted to avoid the idea that we are actually building a new codec.
Our preferred framing is this: There are already "codecs" for vector graphics that are open, and as well established as H264 is. We're just working on a converter / transcoder from raster to vector, which admittedly will always have artifacts of some kind - though you could just as easily take source vector files and stream them and transcode them to a vector codec (SVG or WebGL) without doing that kind of raster to vector conversion that we're proposing.
I think you're right about current web tech being able to support this. Lottie extended to an open source vector "video" format would be awesome, with HLS/Dash streaming and especially audio (streaming in sync) too. Hope you can find a sponsor for such an open source project
1. https://www.pouet.net/prod.php?which=99 https://www.youtube.com/watch?v=J2r7-ygXOzo
2. https://www.pouet.net/prod.php?which=100 https://www.youtube.com/watch?v=tGetanBEKK8
Kind of related, this series of articles on the vector encoder for Another World [3].
1: http://www.mos6502.com/friday-commodore/9-fingers-the-infamo...
2: https://chriscummins.cc/s/genetics/
3: https://fabiensanglard.net/another_world_polygons/index.html
> "State of the Art was traced by hand with a Genlock overlay and tool I developed. In 9 fingers the process was automatic, my program controlled the videoplayer, digitized one picture, traced it, and skipped to the next frame. For me the equipment at that time was expensive, about 150 Euros for the videoplayer (used the prize money from State of the Art), since it had to show de-interlaced pictures."
But it was so long ago I don’t have the link, I only responded to you to help verify that it was a real thing and you’re not dreaming
Or maybe we all are dreaming 🧐
Ok that got meta quickly
I also swear that I've already seen a vector/cartoon optimized codec, but my google-ability are lacking.
https://files.vectorly.io/demo/v0-2-simpsons-250kbps/index.h...
(Edit: I just notice the full-screen button on the video, I really think they should be emphasizing that, since vector video should be able to stay smooth at extreme resolutions)
https://files.vectorly.io/demo/resizeable/index.html
We're still not comfortable putting demos on our front page, for all the bugs everyone's mentioned, and other issues.
Not even sure who posted this on HN, but grateful for the attention / feedback!
There's a relatively low-quality cut on YouTube: https://www.youtube.com/watch?v=Z3ehuHRnC4Q
> Video Codec for Classical Cartoon Animations with Hardware Accelerated Playback
> We introduce a novel approach to video compression which is suitable for traditional outline-based cartoon animations. In this case the dynamic foreground consists of several homogeneous regions and the background is static textural image. For this drawing style we show how to recover hybrid representation where the background is stored as a single bitmap and the foreground as a sequence of vector images.
The idea of using prior knowledge about the nature of the content to decide on an encoding scheme makes intuitive sense to me, though I'm not a codec person so I don't know how feasible it would be to make these ideas into the hardware-accelerated codecs we know from other methods.
Of course, these methods would make the most sense when used directly by the animation studios during export, not as an afterthought. But I'll take what we can get.
By the way, the corpus of Sýkora's works [2] is really really impressive in my opinion. He gave a talk in my institute while I was researching methods around neural style, and his take on parametric models, paired with the quality (and speed!) of his results, really left a mark, if not to say they made me seriously question wtf I was doing there. His work is strictly tailored to a professional animation / video production setting, so it seems extremely applicable compared to the toy-like nature of neural style methods. That is not to say he doesn't know about those. His team's recent papers actually fruitfully combine the two.
One good example is https://asciinema.org , which plays back terminal sessions. The text in the "video" is selectable!
You can certainly leverage that, but it also increases the resource demands (and thus power demands) on the client.
Try remoting into a machine over RDP and playing a youtube video occupying a large portion of the screen. You may be surprised to find that it plays nearly perfectly, even over non-ideal network conditions (wifi). Try this same exercise with VNC and you will experience a frame maybe once every other second.
I am not sure exactly the heuristics involved, but RDP is certainly switching between modes of operation based on what kind of visual information is on the screen.
The best thing to do is just use a modern video codec and make sure it works well with text and sharp edges.
In a world where videos are still sent around as multi-megabyte gif files, and audio clips are still distributed with a random slideshow on YouTube, I think a lot of users aren't so bothered about efficiency - they just want the simplest thing that works.
Lots of music is most efficiently stored as MIDI, yet how many songs on iTunes are midi?
For video, raster is king because it works for everything.
While storing music as its component instruments is indeed size efficient, converting existing music into MIDIs using some ML/AI algorithm is not yet effective. Vectorly seems to convert _existing_ raster videos into vector ones using some combination of computer vision algorithms.
Switching to H265 or AV1 is way more compelling in practice because it works for everything, it's hardware accelerated, the decoders are off-the-shelf, and the quality is significantly improved. A modern GPU can both encode and decode H265.
Random link with some pricing information for Africa [1] claims prices are up from $0.50/GB.
[1] https://kenyanwallstreet.com/mobile-data-pricing-2020-report...
- Our videos use about 1/100 bandwidth of a video. This matters for rendering (fans running when watching HD?) and to people who live in places with shitty ISPs like in Germany
- It is rendered and thereby crisp as a Pringle
- No codec needed as we render it as html
- When you pause a video you actually have the text there, not pixels. For programming that allows you to copy and change things!
- We store the context it was recorded in (dev environment) so you can change and run code in the "video" itself. This changes the pedagogy you can do in a video to be more interactive and hands on.
(As an aside, MIDI is a useful format in its own right, and is still alive and well in the music-making world, even if from a technical perspective it's sort of outdated in 2020.)
Also would be nice to use when you have bad internet connection speed to watch e-learning material animations.
For e-learning you may need a hybrid M3u like Playlist approach with video for the presenter and vector graphics for the screen casts.
Manga videos would probably also compress well.
Children animated videos.
What if you would reduce the color space vectorize of ordinary video to for example 8 colors and smooth out the noise to make large flat surfaces could you compress it with vectors?
Here are two random screen captures from anime airing this season:
https://pbs.twimg.com/media/EdP8gc4WAAAbkWy?format=jpg&name=...
https://pbs.twimg.com/media/EdPwGmyXoAIjlGw?format=jpg&name=...
Suffice it to say that vectorizing this content would be difficult and the vector representation wouldn't be particularly small.
Using html gives us some additional benefits in that we can combine rendered with rasterized content. We also get access to a lot more advanced functionality through the browser context.
>In practice however, DRM, streaming , analytics and ad placement also require javascript logic to function in web runtimes, so in real word settings web-video playback can and does use a non-trivial amount of CPU time.
I'm a little skeptical of this claim. No idea how much CPU is used for DRM, but I can't imagine it's on the order of multiple percentage points.
GPUs happen to be very good at drawing shapes. So there's probably a lot of room for improvement, probably even enough that these kinds of videos could be faster than decoding "normal" video (for which modern CPUs and GPUs have dedicated silicon).
This doesn't mean that I'm not entirely skeptical of this, especially because this probably will fail spectacularly for animation that isn't just vector graphics, which is most animation.
In comparison, video acceleration on most modern GPUs is done using dedicated silicon that typically has its own performance budget (so it doesn't eat into the GPU time to run shaders or your compositor).
Guessing very roughly based on the figures in https://raphlinus.github.io/rust/graphics/gpu/2020/06/13/fas... (and Pathfinder is integrating some of the techniques there), I’d estimate that piet-gpu and soon Pathfinder, running on Intel integrated graphics, could comfortably draw each frame of this Simpsons sample video in well under one millisecond (for an effective maximum of 6% GPU/CPU usage), probably even under a tenth of a millisecond (0.6%).
(Admittedly you could easily create a more complicated video which would be much more draining, e.g. paris-30k frames would struggle to hit 120fps on integrated graphics. Generally speaking it’s easier to make pathological cases in vector encodings than raster. But let’s forget about that.)
At that point, the hardware separation of video decoder and other graphics stuff (which is a strong point in general, because it’s so computationally expensive) is much less important.
You speak of Pathfinder and Slug as being “very complicated”, and indeed they are; but video decoding is hardly any less complicated, and much more computationally expensive.
Playing back vector video is simply much less expensive, as it’s doing much less work. Assuming content like the sample videos here, I would expect vector video to be immediately competitive with raster video codecs in power efficiency and performance, once you shift the bulk of the vector rendering from CPU to GPU.
Slug is not applicable to this scene -- it makes several core operating assumptions (being able to pre-process curve data into a friendlier format, and its shape bounding boxes are well-defined and pre-processed. These hold for font / text assets but not real-time glyphs)
Nationalization of costs, privatization of profits.
I'm going to need to see some evidence of that! I'm pretty sure H.265 or AV1 could encoder better quality videos at that bitrate.
>https://files.vectorly.io/demo/v0-2-simpsons-250kbps/index.h...
There isn't a raster version to compare to, but that looks noticeably worse than what I'd expect from a raster version. There's a lot of artifacting when there's motion, and the linework looks.... off.
>https://files.vectorly.io/demo/khan-20kbps/index.html
The khan academy one looks much better, although there's still some minor artifacting, eg. when the mouse comes close to the "O" in "O_2" changes a bit.
Looking at https://www.youtube.com/watch?v=a1nmZq1KEHk at 240p I think their Simpsons rendition looks better. But that's probably because of the low resolution. Makes me wonder how well H264 would perform with the same bit rate but a higher resolution.
I think it'd be best to have a 1 to 1 comparison between the two right on the vectorly site. I'm thinking that they'd have done that if they believed they were already ready to outperform, but that's a bit of speculation.
That being said, they're just starting and I'm sure they could find some fat to squeeze out of the data stream.
https://caseantiques.com/item/lot-738-the-simpsons-animation...
Look at all that line detail. It creates subtle gradients in the final film, which don't readily transfer to vector graphics.
In this case, I think with some better pixel-based filtering before vectorization, they could've gotten a better result.
However, the much bigger problem is that the vast majority of animation uses paintings for backgrounds that aren't just solid lines and colors.
No patent should be granted for this technique on general. It is known from late 80s. Maybe some details.
Also maybe the patent office shouldn't grant the patent because "use ANN to fit x" is also obvious. Again, only for details.
I did a similar thing for the same motivation a while back (compress Khan Academy videos to a miniscule amount of data), and spent an hour or two on it. http://funny.computer/cloud/Endless/khanvx/player.html and came away thinking it would be exceptionally difficult to work in a general case, but it might be a small useful subset of a larger video codec.
This is not the first time people have tried to do vector video codecs, either. Remember the VSV project from University of Bath? https://www.youtube.com/watch?v=LKaixODJEmo
The idea is so obvious that I would be astounded if this company gets anywhere. I'd wager many research teams already attempted this and were never heard from again.
Also note that video compression is pretty impressive these days. A typical 2 hour 1080p movie compresses down to a handful of GiB. Compare that to a typical 1080p action game which is easily ten times that big, because storing all the meshes and textures takes a lot of space, it turns out.
https://en.wikipedia.org/wiki/Discrete_cosine_transform#3-D_...
Not sure what the pros/cons this kind of approach might bring, or if it's already being used in some areas. I find it really hard to believe that 3d modeling software ecosystem has not experimented with the idea of lossy compression. Does the loss of high frequency information cause more adverse outcomes in 3d models than in 2d models? Seems like this could be useful in games or other creative areas where mathematical precision of the final compressed artifact is not essential.
You wouldn't have any conversion artifacts that way.
Looks pretty good in Firefox on the Mac though: https://imgur.com/a/rqPgQFZ
Edit: Here it is in video form https://imgur.com/a/B0UOigX
Works well on Firefox 78.0.2 on Ubuntu 20.04
While the output of this algorithm (just like my traces) isn't as faithful to the source material as, say, H.264 is, the result looks great and has an amazing style.
This might be a great target for mobile-first webtoons.
The funny thing is that Flash animations where the more straight-forward way. You create vector animations and export them as such.
Obviously, this is more universal. You can take any input as opposed to something that was vectorized to begin with.
Companies that ignore this and patent the “system and method” for implementing algorithms are being jerks.
This might have a future as part of a regular video codec, being used when there's mostly vector graphics on screen (or just for those areas that are vector graphics).
The problem with CAD is that whilst a conversion to a rough sketch will be possible, 'CAD' implies very high accuracy or exactness, and the amount of work required to render a converted sketch accurate will likely trump that by just doing it manually in the first place.
I suppose in future some ML-aided setup that takes a drawing in AND focuses on dimensional accuracy might fare better than this video-centric solution.