H.266/Versatile Video Coding (VVC)
newsletter.fraunhofer.de
newsletter.fraunhofer.de
Is this continued improvement related to the improvement of technology? Or just coincidental?
Like, why couldn't have H.266 been invented 30 years ago? Is it because the computers back in the day wouldn't have been fast enough to realistically use it?
Do we have algorithms today that can compress way better but would be too slow to encode/decode?
Still, I've often thought it would be nice if text were a more first-class citizen within video codecs. I think it's more a toolchain/workflow problem than a shortcoming in video compression technology as such. Whoever is mastering a Blu-Ray or prepping a Hollywood film for Netflix is usually not the same person cutting and assembling the original content. For innumerable reasons (access to raw sources, low return on time spent, chicken-egg playback compatibility), it just doesn't make sense to (for instance) extract the burned-in stylized subtitles and bake them into the codec as text+font data, as opposed to just merging them into the film as pixels and calling it a day.
Fun fact: nearly every Pixar Blu-Ray is split into multiple forking playback paths for different languages, such that if you watch it in French, any scenes with diegetic text (newspapers, signs on buildings) are re-rendered in French. Obviously that's hugely inefficient; yet at 50GB, there's storage to spare, so why not? The end result is a nice touch and a seamless experience.
This is all besides complexity like video/audio content synced text or handling multiple simultaneous speakers. Even that is besides workflow/tooling issues that you mentioned.
The MPEG-4 spec kind of punted on text and supports fairly basic timed text subtitles. Text essentially has timestamp where it appears and a duration. There's minimal ability to style the text and there's limits on the availability of fonts though it does allow for Unicode so most languages are covered. It's possible to do tricks where you style words at time stamps to give a karaoke effect or identify speakers but that's all on the creation side and is very tricky.
The Matroska spec has a lot more robust support for text but it's more of just preserving the original subtitle/text encoding in the file and letting the player software figure out what to do with that particular format and then displaying it as an overlay on the video.
It's unfortunate text doesn't get more first class love from multimedia specs. There's a lot that could be done, titles and credits as you mention, but also better integration of descriptive or reference text or hyperlink-able anchors.
https://www.semanticscholar.org/paper/A-Means-for-Achieving-...
The way I understand it is that one way to compress a document would be to store a computer program and, at the decompression stage, interpret the program so that running it outputs the original data.
So suppose you have a program of some size that produces the correct output, and you want to know if a smaller-sized program can also. You examine one of the possible smaller-sized programs, and you observe that it is running a long time. Is it going to halt, or is it going to produce the desired output? To answer that (generally), you have to solve the halting problem.
(This applies to lossless compression, but maybe the idea could be extended to lossy as well.)
If you are looking at Kolmogorov complexity you are right, we can't ever know. But Kolmogorov complexity is about single points in the space of possible outputs. It basically says "there might be possible outputs that do look random, but are actually produced by a very short encoding". One example would be the digits of pi.
But if you look at the overall statistics of possible output streams, and at their averages, there is a lower bound for compression on average. As soon as the bitlength in the compressed stream matches the entropy in the uncompressed stream in bits, you reached maximum compression. There will be some streams that don't conform to those statistics, but their averages will.
However, we are somewhat far away from matched entropy equilibrium for video compression. And even then, improvements can be made, not in compression ratio but in time, ops and energy needed for de/encoding.
You could actually use ML for all of the video decoding, but that research is still in it's early stages. It has been done rather well with still images [1], so I'm sure it'll eventually be done with video too.
Those ML techniques are still a little slow and require large networks (the one in [1] decodes to PNG at 0.7 megapixels/s and its network is 726MB) so more optimizations will be needed before they can see any real-world use.
[1] https://hific.github.io/ HN thread: https://news.ycombinator.com/item?id=23652753
Right now neural networks allow for higher compression for tailored content, so you need to ship a decoder with the video, or have several categories of decoders. The future is untold and it might end up not being done this way.
[0] https://en.wikipedia.org/wiki/Deep_learning_super_sampling
Once you get to an AI that has full comprehension of what humans perceive to be reality, you can just give them a rough outline of a story, add some information on casting, writers, and Spielberg's mood during production, and they'll fill in the (rather large) blanks.
That's a bit exaggerated, but I remember reading about one such algorithm a few days ago (by Netflix, maybe?). It was image compression that had internal representations such as "there is an oak tree on the left".
It would then run the "decompression", find the differences to the original, and add further hints where neccessary.
> Because H.266/VVC was developed with ultra-high-resolution video content in mind, the new standard is particularly beneficial when streaming 4K or 8K videos on a flat screen TV.
Compressing video is very different from gzipping a file. It's more about human perception than algorithms, really. The question is "what data can we delete without people noticing?", and it makes sense that answer is different for an 8k video than a 480p video.
This new techniques can also be used for 1080p video (for example) but with lower gain. Also the "old" algorithm/system are generally still used but they may be improved/extend.
if you film(still film no movement) the macbook pro top to bottom in h264/MP4 at 1024 p resolution and again you take a picture from your camera.
the results will be shocking,
the video of 5-10 seconds will have lower storage size than the size of single Image. but when you inspect video carefully you will see the tiny details are missing, like the edges are not as sharp for metal body, the gloss of metal is bit different, the tiny speaker holes on top of keyboard are clear in Images and can be individually examined while in video they are fuzzy pixelated and so on.
so, The end results: A 5 second video with 10s of frames per second is smaller in size than single Image taken from same camera.
I know that's not fair comparison, but imaging clever compressions were not invented and you would download terabytes of data to view a small movie.
https://sidbala.com/h-264-is-magic/
GREAT article
We don't know yet. There are no public technical details (that I know of) for H266 yet, but if I recall H265 gave the same 50% reduction in bandwidth claims, and for years people stuck with H264, because it was higher quality due to dropping off less subtle parts of the video you really want to see. Only in the last couple of years has H265 really started to become embraced and used by piracy groups. Frankly, I don't know what changed. I wouldn't be surprised if there was some sort of H265 feature addition that improved the codec.
Encoder maturity is also a big factor: when H.265 first came out, people were comparing the very mature tools like x264 to the first software encoders which might have had bugs which could affect quality and were definitely less polished. It especially takes time for people to polish those tools and develop good settings for various content types.
this. our TV (a few years old, "smart") can play videos from a network drive, but doesn't support H.265. reencoding a season's worth of episodes to H.264 takes a while...
That didn’t stop, of course, but I generally don’t notice the improvements if I’m not looking for them. Someone with a 4K or better home theater will 100% benefit from newer codecs’ many improvements on all those extra pixels but if you’re the other 95% of people watching on a phone or tablet, lower-end TV with the underpowered SoC the manufacturer could get for $15, etc. you probably won’t notice much difference and convenience will win out for years longer.
One area where this can be important to remember is when comparing codecs: a fair number of people will make the mistake where they'll take a relatively heavily compressed video, recompress it with something else, and get a size reduction which is a lot more dramatic than what you'd get if you compared both codecs starting from a source video which has most of the original information.
Even though x264 quickly started seeing better results compared to DivX and XVid, you didn't see pirate encodes switch to x264 for years.
So it was kind of a magical codec upgrade and now it seems more incremental to me.
Money quote :
"H.264 is a newer video codec. The standard first came out in 2003, but continues to evolve. An automatically generated patent expiration list is available at H.264 Patent List based on the MPEG-LA patent list. The last expiration is US 7826532 on 29 nov 2027 ( note that 7835443 is divisional, but the automated program missed that). US 7826532 was first filed in 05 sep 2003 and has an impressive 1546 day extension. It will be a while before H.264 is patent free."
(emphasis mine)
[1] https://www.osnews.com/story/24954/us-patent-expiration-for-...
Is it just basically the same mechanism that leads to so much drug development happening in the US despite how backwards its medical system is, because those regressive institutions create profit incentives not available elsewhere (to develop drugs or video codecs for profit) and thus the US already has capitalists throwing money at what could be profitable whereas everyone else would look at it as an investment cost for basically research infrastructure.
https://en.wikipedia.org/wiki/Moving_Picture_Experts_Group
The article we are all commenting on is by a German research organization that has been a major contributor to video coding standards.
Perhaps you're confused by the patent issue? European companies are happy to file for US patents and collect the money.
Compared to 30 years ago, we now have better knowledge and statistics about what low level primitives are useful in a codec.
E.g. jpeg operates on fixed 8×8 blocks independently, which makes it less efficient for very large images than a codec with variable block size. But variable block size adds overhead for very small images.
An other reason can be common hardware. As hardware evolves, different hardware accelerated encoding/decoding techniques become feasible that gets folded into the new standards.
Depending on the client or the number of images on the site the huge JPEG could be a crippling performance issue, or even a “site doesn’t work at all” issue.
What kind of overhead?
The extra code shouldn't make a big difference if nothing is running it.
And space-wise, it should only cost a few bits to say that the entire image is using the smallest block size.
Is the worry just about complicating hardware decoders?
For me, this is another pointless advance in video technology. 720p or 1080p is fantastic video resolution, especially on a mobile phone. Less than 1% of the population cares or wants higher resolution.
What new technologies are doing now is re-setting the patent expiration clock. As long as new video tech comes out every 5-10 years, HW manufacturers get to sell new chips, phone manufacturers get to sell new phones, TV manufacturers get to sell new TVs, rinse, repeat.
720p is far from fantastic. It's noticeably blurry, even on mobile.
1080p is minimally acceptable, and is now over 10 years old.
> Less than 1% of the population cares or wants higher resolution.
That's a very bold claim. Have any studies or polls to back that up?
From 1:
> In Japan, only 14 percent of households will own a 4K TV in 2019 because most households already have relatively new TVs, IHS said.
Let's dissect the reasoning. Most households already have a relatively new TV, thus a low adoption rate. Implying that needing a new TV generally, not the desire to upgrade resolution, is the primary motivating factor in purchasing a TV. In fact, most 4K TVs are already as cheap as the rest of the market.
I truly believe that almost everyone does not care about 4K whatsoever. In fact, even if they do 'care' it's not because they know what they're talking about. Most of the enhancements that 4K TVs bring are an artifact of better display technology, rather than increased resolution. See 2.
Streaming 4K+ video is a waste of resources with no tangible benefit to anyone other than marketing purposes. Netflix streams '4K' because everyone has a '4K' TV now and they demand it.
[1]: https://www.twice.com/research/us-4k-tv-penetration-hit-35-2...
[2]: https://www.cnet.com/news/why-ultra-hd-4k-tvs-are-still-stup...
> In addition, “with the Japanese consumer preference for smaller TV screens, it will be more difficult for 4K TV to expand its household penetration in the country"
With a smaller TV screen, yeah, you don't need 4K.
But in the USA, larger screens are desirable. And that's seen in the expected 34% 4K adoption rate in the USA your article describes.
I still use a 1080p TV, but it's also only 46" and is 10 years old. I'll probably be buying a 4K 70-75" OLED later this year.
My computer monitor is 1440p. I could have bought 4K when I upgraded, but I'm primarily a gamer and I wanted 144 hz, and 4K 144 hz monitors didn't exist yet.
> Streaming 4K+ video is a waste of resources with no tangible benefit to anyone other than marketing purposes. Netflix streams '4K' because everyone has a '4K' TV now and they demand it.
It's not inherently wrong, just practically so. The difference is that Netflix (and often other streaming services) max out at a much higher bitrate for their 4k streaming, and are using a better codec (H.265) as well. By comparison, Netflix's bitrate for 1080p is severely limited and so if you compare the two, even watching at the 1080p level of detail, streaming in 4k will often be a vastly superior experience.
So it's not inherent (not a result of the resolution), but still, streaming 4k is not pointless at present.
Not OP. I would say these are far-fetched claims that need defending. Most blurriness of mobile video comes from the low bitrate it's encoded at. Basically nobody is watching Bluray quality 720p or 1080p materials on a phone - and that's the problem, not the resolution.
My guess is that a typical middle or even middle-upper class family is going to have a TV that is less than 70 inches, and is 10+ feet from most viewers. Even with 20-20 vision, the full quality provided by 1080p is not even visible at that distance! (You'd need to go all the way up to 78 inches at 10 feet, or sit 7 feet from your 55 inch set to even get the full benefit from a 1080p set.) See this very helpful chart: http://s3.carltonbale.com/resolution_chart.html
Most of the benefit in 4k video comes from recent advances in HDR presentation and better codecs, not from the resolution. Sure, if you're a real stickler for quality, you might be sitting 6 feet from your 80 inch OLED set, and 4k is definitely for you in that case, but it's really not that important to the average person. In my case, I can barely distinguish between 720p and 1080p on my set even with glasses on.
Now granted, it's great to have laptops and tablets at higher resolutions, because your face is smashed up against them and you're often trying to read fine text. But that's not the video case that's being talked about here.
On mobile, 720p starts to get shoddy once your screen hits 5 inches across.
So while 4k is situational, and encoding quality is more important than the extra resolution most of the time, 1080 vs. 720 is pretty clear-cut; 1080 should be considered the minimum for most content.
> On mobile, 720p starts to get shoddy once your screen hits 5 inches across.
Even assuming you're right abut this, I guess I really just have a hard time caring. Anything you watch on your phone is at best something you don't give a shit about, artistically speaking, and it's hard to imagine 1080p vs 720p making any kind of difference to the experience. (I suppose I might be biased since my screen is "only" 5.2 inches diagonally - I get the smallest one I can whenever I buy a new phone.)
And for what it's worth, using the same math as for my previous comment, you can't even get the full benefit of 720p on a 5 inch diag screen unless you're holding it less than a foot from your face. Granted, you can get the full benefit of 1080p at a little under 8 inches, but I'm suffering even imagining trying to watch a video this way. Even at this distance, I would dispute using "shoddy" to describe how 720p will look.
The math is actually pretty simple: for a 720p screen, there are sqrt(12802 + 7202) pixels on the 5 inch diagonal, so a distance of 5 inches / sqrt(12802 + 7202) per pixel. 20/20 visual acuity can resolve roughly 1 arc minute, or pi/10800 radians. By the arc length formula, the distance we calculated subtends that angle at (5 inches / sqrt(19202+10802)) * (10800/pi), or 11.7 inches.
> And in my experience there's a lot of screens closer than 10 feet to couches.
Note that I addressed this point. If you have a pretty typical 50 inch TV, you've got to have it closer than 7 feet from your couch for 4k to make any difference at all.
There seems to be disconnect about how people are consuming media. For me holding a 2280x1080 6.3" phone ~8" from my eyeballs is a natural viewing distance and I can see the full resolution without difficulty. And at least from my point of view a 65" TV is also a pretty typical size.
Yes, so am I.
> therefore that 1080p is usually better than just "minimally acceptable", since most of the time we don't even get the full benefit of it.
If most of your users would be limited by 720p, and 1080p is standard and easy to do, then I'm comfortable calling 1080p the minimum.
> I'm suffering even imagining trying to watch a video this way.
The official "Retina" numbers have a phone 10 inches from your face. And that's about where I hold it when I have my glasses on.
> 20/20 visual acuity can resolve roughly 1 arc minute, or pi/10800 radians.
That's a good baseline number, but a lot of people can beat it by a significant fraction.
Would you? AV1 was only officially released 2 years ago, h.265 7, h.264 14, …
It is all a matter of trade offs and engineering.
For MPEG / H.26x codec, the committees start the project by asking or defining the encoding and decoding complexities. And if you only read Reddit or HN, most of the comment's world view would be Video codec are only for Internet Video and completely disregard other video delivering platform. Which all have their own trade off and limitations. There is also cost in decoding silicon die size and power usage. If more video are being consumed on Mobile and battery is a limitation, can you expect hardware decoding energy usage to be within the previous codec? Does it scale with adding more transistors, are there Amdahl's law somewhere. etc It is easy to just say adding more transistor, but ultimately there is a cost to hardware vendors.
Vast majority of the Internet seems to think most people working on MPEG Video Codec are patents trolls and idiots and paid little to no respect to its engineering. When as a matter of fact Video Codec are thousands of small tools within the spec, and pretty much insane amount of trial and error. It may not be as complicated as 3GPP / 5G level of complexity, but it is still lot of work. Getting something to compress better while doing it efficiently is hard. And as Moore's Law is slowing down. No one can continue to just throwing transistors at the problem.
One of MPEG-1's design goals was to get VHS-quality video at a bitrate that could stream over T1/E1 lines or 1x CD-ROMs. The limit on bitrate led to increased algorithmic complexity. It was well into the Pentium/PowerPC era until desktop systems could play back VCD quality MPEG-1 video in software.
Later MPEG codecs increased their algorithmic complexity to squeeze better quality video into low bit rates. A lot of those features existed on paper 20-30 years ago but weren't practical on hardware of the time, even custom ASICs. Even within a spec features are bound to profiles so a file/stream can be handled by less capable decoders/hardware.
There's plenty of video codecs or settings for them that can choke modern hardware. It also depends on what you mean by "modern hardware". There's codecs/configurations a Threadripper with 64GB of RAM in a mains powered jet engine sounding desktop could handle in software that would kill a Snapdragon with 6GB of RAM in a phone. There's also codecs/configurations the Snapdragon in the phone could play using hardware acceleration that would choke a low powered Celeron or Atom decoding in software.
In MPEG-1's heyday there would haven't been a lot of point in encoding presets producing content common hardware decoders couldn't handle.
There were several other video codecs in the same era that didn't have hardware decode requirements. Cinepak was widely used and could be readily played on a 68030, 486, and even CD-ROM game consoles. As I recall Cinepak encoders had more knobs and dials since the output didn't need to hit a hardware decoder limitation.
Decoders are relatively simple book keepers/transformers. Encoders are complex systems with tons of heuristics.
This is also why hardware decoders tend to be in everything and are relatively cheap with equal quality to software counterparts. On the flip side, hardware encoders are almost always worse than their software counterparts when it comes to the quality of the output (while being significantly faster).
That's what I meant by my first sentence.
And I'll throw out there that the vast majority of 'hardware codecs' are in fact software codecs running on a pretty general purpose DSP. You could absolutely reach the same quality as a high quality encoder given the right organizational impetus of the manufacturer; they simply are focused on reaching a specific real time bitrate for resolution rather than overall quality. By the time they've hit that, there's a new SoC with it's own DSPs and it's own Jira cards that needs attention. If these cores were more open, I'm sure you'd see less real time focused encoder software targeting them as well.
This has been an issue with AV1, it's got relatively high decode complexity and there's not a lot of hardware acceleration available. The encode complexity is fantastic though and is very slow even on very powerful hardware, less than 1fps so ~30 hours to encode a one hour video. Even Intel's highly optimized AV1 encoder can't break 10fps (three hours to encode an hour of video) while their h.265 encoder can hit 300fps on the same hardware.
Also, in many applications, it’s suitable to exchange time for memory / compute. You can spend an hour of compute time optimally encoding a 20-minute YouTube video, with no real downside.
Neither of these approaches are suitable for things like video conferencing, where there is a small number of receivers for each encoded stream and latency is critical. At 60fps, you have less than 17ms to encode each frame.
Interestingly, for a while, real-time encoders were going in a massively parallel direction, in which an ASIC chopped up a frame and encoded different regions in parallel. This was a useful optimization for a while, but now, common GPUs can handle encoding an entire 1080p frame (and sometimes even 4K) within that 17ms budget. Encoding the whole frame at once is way simpler from an engineering standpoint, and you can get better compression and / or fewer artifacts since the algorithm can take into account all the frame data rather than just chopped up bits.
TL;DR: it's partly because we're using higher video resolutions. A non-negligible part of the improvement stems from adapting existing algorithms to the now-doubled-resolution.
Almost all video compression standards split the input frame into fixed-size square blocks, aka "macroblocks". To put it simply, the macroblock is the coarsest granularity level at which compression happens.
- H.264 and MPEG-2 Video use 16x16 macroblocks (ignoring MBAFF).
- H.265 use configurable quad-tree-like macroblocks, with a frame-level configurable size up to 64x64.
- AV1 makes this block-size configurable up to 128x128.
Which means:
Compression to H.264 a SD video (720x576, used by DVDs) results in 1620 macroblocks/frame.
Compressing to H.265 a HD video (1920x1080) results in at least 506 macroblocks/frame.
Compressing to AV1 a 4K video (3840x2160) results in at least 506 macroblocks/frame.
But compressing to H.264 a 4K video (3840x2160) will result in 32400 macroblocks/frame.
The problem is, there are constant bitcosts per-macroblock ((mostly) regardless of the input picture). So using H.264 to compress 4K video will be inefficient.
When you take an old compression standard to encode recent-resolution content, you're using the compression standard outside of the resolution domain for which it was optimized.
> Is this continued improvement related to the improvement of technology? Or just coincidental?
Of course, there also "real" improvements (in the sense "qualitative improvements that would have benefited to compression old video resolutions, if only we had invented them sooner").
For example:
- the context-adaptive arithmetic coding from H.264, which is a net improvement over classic variable-length huffman coding used by MPEG-2 (and H.264 baseline profile).
- the entropy coding used by AV1, which is a net improvement over H.264's CABAC.
- integer DCT (introduced by H.264), which allow bit-accuracy checking and way lot easier and smaller hardware implementations (compared to floating point DCT that is used by MPEG2).
- loop filters: H.264 pioneered the idea of a normative post-processing step, whose output could be used to predict next frames. H.264 had 1 loop filter ("deblocking"). HEVC had 2 loop filters: "deblocking" and "SAO". AV1 has 4 loop filters.
All of these a real improvements, brought to us by time, and extremely clever and dedicated people. However, the compression gains of these improvements are nowhere near the "50% less bitrate" that is used to sell each new advanced-high-efficiency-versatile-nextgen video codec. Without increasing - a lot - the frame resolution, selling a new video compression standard will be a lot harder.
Besides, now that the resolutions seems to have settled up around 4K/8K (and that "high definition" has become the lowest resolution we might have to deal with :D), things are going to get interesting ... provided that we don't start playing the same game with framerates!
In the mid '90s, PCs often weren't fast enough to decode DVDs, which were typically 720x480 24FPS MPEG2. DVD drives were often shipped with accelerator cards that decoded MPEG2 in hardware. I had one. My netbook is many orders of magnitude faster than my old Pentium Pro. But it's not fast enough to decode 1080p 30fps H.265 or VP9 in software. It must decode VP9/H.265 on the GPU or not at all. MPEG2 is trivial to decode by comparison. I would expect a typical desktop PC of the mid '90s to take seconds to decode a frame of H.265, if it even had enough RAM to be able to do it at all.
It's an engineering tradeoff between compression efficiency of the codec and the price of the hardware which is required to execute it. If a chip which is capable of decoding the old standard costs $8, and a chip which is capable of decoding the new standard costs $9, sure, the new standard will get lots of adoption. But if a chip which is capable of decoding the new standard costs $90, lots of vendors will balk.
By the end of that year the G4 Power Macs were just barely fast enough to play DVD's with software decoding and assistance from the PCI or later AGP video card. And after a while (perhaps ~ 2002?), even the Blue G3's could do it in software even if you got a different video card, as long as you also upgraded to a G4 CPU (they were all in ZIF sockets).
It was very taxing on computers at y2k!
Later autumn 2000 G3 iMacs could also play DVD's but I think they needed special help from a video co-processor.
https://www.theregister.com/2013/08/06/xerox_copier_flaw_mea...
Here's something to consider:
In 1995, a typical video stream was 640 x 480 x 24fps. That's 7,372,800 pixels per second.
In 2020, we have some content that's 7680 x 4320 x 120fps. That's 3,981,312,000 pixels per second, or a 540 fold increase in 25 years.
The massive increase in image size actually makes it easier to use high compression ratios. I found this out the hard way recently, when I was trying to compress and email a powerpoint presentation that a coworker had presented on video. In a nutshell, the powerpoint doc with it's sharp edges and it's low resolution made it difficult to compress.
Increase framerate plays a factor too; due to decades of research on motion interpolation, algorithms have become quite good at guessing what content can be eliminated from a moving stream.
Would you? Video compression is one of the few things that we will work on for the next 1000 years and still be nowhere near finished. The best video compression would be to know the state of the universe at the big bang, have a timestamp of the beginning and end of your clip and spatial coordinates defining your viewport. Then some futuristic quantum computer would just simulate the content of your clip...
So yeah, sure we are done with video compression :). This is of course an extreme example of constant time compression that may or may not be ever feasible (if we live in a computer simulation of an alien race, then it is already happening).
But the gist is the same. Video compression is mostly about inferring the world and computing movement not by storing the content of the image.
For instance by taking a snapshot of the world, decomposing it into geometric shapes (pretty much the opposite of 3D rendering) and then computing the next frames by morphing these shapes + some diff data that snaps these approximations back in line with the actual data.
We are all but in the very infancy of video compression. What should surprise you is why it takes us so long to get anywhere.
Early pc's had separate and very expensive mpeg decode boards just to decode dvd, creative sold a set, the cpu simply couldn't even handle mpeg 2. I know its hard to believe but there was a time when playing back an mp3 was a big ask, all these algorithms could be made long ago, but they would have been impractical fantasy. Only now are we seeing real partial cheated resolution ray tracing in modern high end gaming hardware now which is a good comparison, ray tracing has been with us for a long time, only hardware advancement over decades has made it viable.
It amused me that they claimed 4k uhd h265 is now 10GB for a movie, that's garbage bitrate, they always ask too much of these codecs.
can confirm. audio playback would stutter on my 486dx if one dared to multitask.
http://www.videolan.org/developers/x264.html
But if you use an H.264 encoder or decoder in a country that recognizes software patents then you need to buy a patent license if your usage comes under the terms of the license:
If you build x264 or other open source implementation of h.264/h.265, and embed it for example in commercial video conferencing software/appliance, you have to pay patent licensing fees for that product.
It's also why Firefox downloads a blob from Cisco to handle MPEG-4 video - Cisco covers the licensing for distribution et al.
Cisco has some products which use compressed video in a browser setting. It would be useful if all browsers supported a good codec. Individually downloaded codec plugins suck, because installing is iffy. Therefore, give something away which doesn't cost licensing money to make your existing licensed products more usable.
And get some good feels on the interwebs.
That wasn't enough, and WebRTC requires both VP8 and H.264 as MTI codecs.
However, patent clerks have also from time to time registered algorithmic patents.
I would still take freedom over patented software
> […] but h265 is now significantly more efficient than av1.
What did you mean?
Also this kind of performance claims are 100% hot air. Real world benchmarks talk.
H266 will be adopted by broadcasting and archival and will make mpeg tankers of money. Whatever the next generation of physical home media is after blu-ray will use it, the player for it will read it, and your TV cable box will take h266 signals in to decode. The costs of paying mpeg will be in the cost of the discs, the cost of the cable package, etc.
The real win we should... hope? For is that h266 never sees a personal computer hardware decoder from Qualcomm / Samsung / Intel /AMD / Nvidia / etc. If online video is exclusively distributed with AV1 then none of these companies need touch the festering MPEG patent hell and consumers avoid that parasite leeching money out of their computer purchases.
Because the cable box and physical media player are dying. You can generally opt out of them and avoid filling the MPEG coffers with software patent money. And the big web companies that have the power to dictate what computers are using for the next decade and beyond are all way favoring AV1 with the exception of Apple.
AV1 will probably win in most circumstances (big tech) but is unlikely to win where there are big gains to be had by reducing file size (broadcasters with gigantic libraries).
Broadcasters are also used to paying a lot and not getting much.
This is a market that is voluntarily paying for perceived value.
$0.20 for MPEG-LA, $0.40 for HEVC Advance, and "call us" for Technicolor and Velos Media. $0.99 doesn't sound far off.
Very few Win10 users would want a CPU-targeted HEVC codec.
Intel, nVidia and AMD have that codec in their hardware. They are probably paying for a license to use these patents, they ship Win10 drivers for their hardware, and Microsoft publishes that drivers on Windows update.
Nearly 30 years after MP3, the only audio codec that could rivals mp3 at the standard rate of 128kbps at a significant lower bitrate was Opus at 96Kbps.
And MP3 is still by far the most popular codec due to compatibility reason.
This is similar to JPEG, although things are about to change.
AAC and Vorbis were doing this for years before Opus was on the scene. Opus is a further improvement on audio codecs, but not an unprecedented one.
Is it? Because Google/YouTube, Amazon/Twitch, Netflix, Microsoft, Apple, Samsung, Facebook, Intel, AMD, ARM, Nvidia, Cisco, etc, are all part of AO Media:
* https://aomedia.org/membership/members/
The main major tech player I don't see is Qualcomm.
And most of those companies are also part of MPEG as well.
They're part of MPEG because of legacy reasons in having to deal with H.264.
You were probably thinking about hardware decoding though.
They're complementary options rather than competing. Each does something well, and it sounds like you want the thing that software encoding does.
AV1 is unlikely to ever be practical for "muggle" encode use, at least in this decade. It will only be worth committing that much compute workload to making a smaller file if the recipients will number in at least what, millions?
I'd be really curious what a hardware realtime AV1 encoder would even look like. How much silicon would that take? That kind of chip would have be be colossal even if it sacrifices huge amounts of efficiency to spit out frames at reasonable time (in the same way hardware hevc and vp9 encoders kind of suck).
Can we agree not to work on such projects? I feel that the lack of a good open-source encoder/decoder would spell the death of most codecs nowadays. That would also teach Fraunhoffer about it.
Of course, everyone is free to scratch their itch. And the bigger the void, the more itchy it gets. Luckily, we still have AV1.
I appreciate that it costs money and time to develop these algorithms, but when you're backed by multi-billion dollar "partners from industry including Apple, Ericsson, Intel, Huawei, Microsoft, Qualcomm, and Sony" perhaps they could swallow the costs? It is 2020 after all.
This is also why you see articles from time to time highlight a stupid patent as if it's an outrage that the patent office allowed it. It's not the patent offices job to enforce patents. You can literally go ahead and patent swinging on a swing (1) and the patent office would approve it if the paper work is in order. The media would then likely pick up on this with outrage as if that's an enforceable patent. The truth is that it's simply not the patent offices job. Patents are meant to be enforced by courts.
Actually, it is (at least in the US). USPTO can deny patents on the basis on nonpatentability, and its general refusal to do so post-State St. decision is often cited as one of the problems of the modern patent system.
Broadly speaking, however, if the argument is that software patents are invalid in Europe because they'll be found so by the courts, it should be noted that SCOTUS is actually pretty likely to rule software unpatentable were it to hear a software patent case. A little background is in order:
In Parker v Flook (1978), SCOTUS said that mathematical algorithms (i.e., basically software) is unpatentable. In Diamond v Diehr (1981), they said that part of the patent being software doesn't make the entire thing invalid. The big decision is State St (1998), which is a CAFC decision holding that anything was patentable so long as it produced a "useful, concrete, tangible" result and basically broke the patent office. When SCOTUS decided Bilski v Kappos (2008), they emphatically (and unanimously!) called out State St as wrong, but declined to endorse any guidelines as to what the limits of patentability should be. Later Mayo (2012) and Alice (2014) decisions again unanimously and unambiguously laid out what wasn't patentable: natural processes, and "do it on a computer" steps.
A few years ago, we had a patent attorney at work tell us (paraphrasing somewhat) that Alice made it really hard to figure out how to write a software patent that wouldn't be invalidated. Their continued existence (and pretense to their enforceability) is less because it's secure and more because no one wants to spend the money to litigate it to the highest level (see also the Google v Oracle case, which is exactly the sort of thing a software patent case history would entail).
At least the patent licenses usually used with MPEG mean that private use of open source implementations is free.
Why is that?
Now compound this by the fact that a) trying to make an exhaustive patent search to get a verifiable claim that you don't infringe on any patent is very problematic b) known patent pools like MPEG-LA are known not to cover everything.
So you can make a reasonable bet that you avoid infringing patents by avoiding patents from MPEG-LA and few other better known groups, but you can't actually guarantee that you're not infringing on any patents.
This resolves, sort of, into a game of chicken and depends heavily on whether a lesser known patent holder decides it's worth it to bother executing against you... but even if they don't, unless they come out with a royalty-free license, the possibility of patent is a Damocles' Sword hanging over your codec.
Is there even such a thing?
Isn't the problem that one has to actually go to court to get the answer to this question?
In fact, I heard more than once that current advice is to explicitly avoid searching :/
The way to get around this is to exist in the EU and avoid providing anything to the US.
Care to try your luck?
I don't think they'll be successful. The Alliance for Open Media was careful to avoid potential patent problems during AV1 development. So, unless AOMedia seriously failed in that effort, AV1 will be alright.
The patent system was not designed to have every tiny little technique patented, and this is its failure mode.
In the USA. Wilful and unwilful infringement of the patent costs the same in Europe.
They can be reasonably sure they do not infringe known patents from certain Patent Pools and patents declared as part of MPEG-LA bundles. They can't provide reasonable data that they do not infringe on any submarine patent, something that killed 3 attempts by MPEG-LA to provide a royalty-free codec for the web - all that was required to kill it was a note from a company that they "might" have patents covering things they tried to release, or that they decided not to allow royalty-free license for their known patent.
Meanwhile patent search is complex enough that it's unreasonable to impossible to make a statement that you definitely don't infringe any unless you keep clear of anything invented within last ~20 years.
That's true of every new video codec. It didn't stop the use of H.264 or VP9 or even HEVC.
Multiple companies have now rolled out AV1 into production. We'll see what happens.
What's wrong is claiming that AV1, VP9, VP8 are "patent-free". They are not.
The existence of Theora, VP8, VP9, and now AV1 seems to contradict that theory.
You could argue that they infringe on some unknown patents, but that is also arguably true of patent cabals like MPEG (you just hope that the cabal is big enough that there aren't any patentholders lurking outside). The only difference is that with a patent cabal you have the fun of having to obey the restrictions of everyone who showed up with a possibly-related-in-some-way patent and joined the cabal.
Not to mention that it isn't necessary for a patent pool to be a cabal. AOMedia has a similar structure to a patent cabal except it doesn't act like a cabal (its patent pool is royalty-free in the style of the W3C). So even if the argument is that a patent pool is a good idea (and video codecs cannot be developed without them), there isn't a justification behind turning the patent pool into a cabal.
> At least the patent licenses usually used with MPEG mean that private use of open source implementations is free.
You say that, but there's a reason why some distributions (openSUSE for one) still can't ship x264 (even though the code itself is free software). Not to mention the need for Cisco's OpenH264 in Firefox (which you cannot recompile or modify otherwise you lose the patent rights to use it). The existence of the MPEG patent cabal isn't a good thing, and any minor concessions you get from them do not justify their actions.
And retaliation clauses are present in basically every free software license that has clauses dealing with patents (including Apache-2.0).
I'm more and more fond of calling it all a big game of chicken, a MAD without nukes,
The video patents aren't just "patent troll" patents, either. They are highly enforceable, and were registered by corporations like Ampex.
I have been trying to write a simple app to stream RTSP (security cameras), and that has been a pain.
I need to basically use either proprietary (paid) or GPL software to do it.
Video software is not for the faint of heart. Much as I grouse about the licensing, I am not about to develop my own codec.
I did write this one app, which is an ffmpeg wrapper, to convert RTSP to HLS (Which is not -currently- suitable for realtime streaming): https://github.com/RiftValleySoftware/RVS_MediaServer
It's GPL, because I need to use the GPL ffmpeg H.264 codec.
edit: Your project reminds me, https://github.com/arut/nginx-rtmp-module is super worth checking out and might be helpful to you.
One of the frustrating things about implementing video software, even licensing it, is pretty much everything out there is an expression of ffmpeg, which is a really, really good system, but does have some baggage.
> It's GPL, because I need to use the GPL ffmpeg H.264 codec.
I don't understand the problem?
I can see the problem with people patenting things and preventing you from writing your own implementation, but it seems you just want other people to do the hard work of implementing it so you can wrap a skin around it and do what?
It was a statement of fact.
I've been writing open-source software for well over 20 years (actually, well over 30 years –Where does the time go?).
I think I understand the issues involved.
Apparently, they aren't. There are at least 2 other patent pools that claim patents for HEVC, and I think I saw 3 in some other article before:
https://streaminglearningcenter.com/codecs/hevc-ip-mess-wors...
Ultimately, it's a question of how much you're gonna risk to get where you want, and how much power/influence/wealth you can bring to squash a possible lawsuit.
To completely avoid risk your only choice is to to use old technology where all the patents have expired (20 years in US), like MPEG-2. The next lowest risk is to use H.264 and VP9 which have been out for a while and whose the patent pools have stabilized over the years (and the original parts of the standard will have their patents expire soon - but not some of the newer profiles). After that I would argue that AV1 is less risky that H.265 and H.266, as a lot of work was put into intentionally avoiding patented technology that was not part of the pool, and no one outside the pool has yet made patent claims against it.
EVC baseline is basically that, only using technology based on H.264 that are already expired or soon to be expired and patented techniques from companies that are giving it away to the standard.
In some way EVC is even more exciting than VVC.
"1.3. Defensive Termination. If any Licensee, its Affiliates, or its agents initiates patent litigation or files, maintains, or voluntarily participates in a lawsuit against another entity or any person asserting that any Implementation infringes Necessary Claims, any patent licenses granted under this License directly to the Licensee are immediately terminated as of the date of the initiation of action unless 1) that suit was in response to a corresponding suit regarding an Implementation first brought against an initiating entity, or 2) that suit was brought to enforce the terms of this License (including intervention in a third-party action by a Licensee)."
This makes it much harder for practicing entities or their licensees to assert claims against other practicing entities over the formats.
Edited because I didn't know that some European countries accept software patents.
For example, many countries in EU do not allow patents on software, but that's not something you can claim to be true for all of them - at least before Brexit, since iirc UK was pretty happy to provide software patents.
Then there's a case where if you're really willing you can, as far as I understand, force a patent dispute through WTO, with possibility of patent valid in USA being executed for example in Poland, despite the fact that the patent is invalid in Poland (it doesn't matter if your software is part of physical solution in Poland, algorithms of any kind are not patentable).
The question is, does a "practicing entity" holding the patent cares enough to go through the hardest route to get the patent executed using WTO as a forum? One needs to compare costs and benefits. It's why patent trolling involved pretty much few counties in Texas, because that's where the costs were lowest compared to benefits.
https://en.wikipedia.org/wiki/Software_patents_under_the_Eur...
Not a private organization.
Look at all those European patents in the MPEG-LA license pool!
Wanna take the risk? You might end-up winning the lawsuit, but at this moment, there's a good chance for you to be already out-of-business.
- in Europe: http://www.bailii.org/ew/cases/EWCA/Civ/2002/1702.html
- in US: https://web.archive.org/web/20061205050434/http://eolas.com/...
But as I said, if your startup is being sued by Dolby, whether the enforcement is successful or not is actually irrelevant. Showing that your work doesn't infringe a patent, or that Dolby's patent is invalid, is a money and time-consuming process (unsurprisingly, patents are not generally written to facilitate re-implementation or defense).
(Moreover, in the US, in some cases, the patent owner might even get a preliminary injunction ( https://www.tms.org/pubs/journals/jom/matters/matters-9712.h... ), which might seriously and immediately harm your business. I don't know if such a thing exists in Europe).
Big tech companies like Dolby and IBM use a preventive racket-looking technique ; it involves trying to sell to potential infringers a "protective" subscription, but there's no preliminary analysis of whether there actually is any patent being infringed.
During broadcasting tech events like IBC or NAB, Dolby actually sends people to other company's booths for this ; and there's a famous story about IBM against small-at-this-time SUN : https://www.forbes.com/asap/2002/0624/044.html , whose gist is:
> "OK," [the IBM lawyer] said, "maybe you don't infringe these seven patents. But we have 10,000 U.S. patents.
> Do you really want us to go back to Armonk [IBM headquarters in New York] and find seven patents you do infringe?
> Or do you want to make this easy and just pay us $20 million?"
This is a quirk of some UK courts, where you can literally just start a case to ask a question on some detail of the law and get an answer.
The question was:
> "Is it a defence to the claim under s.60(2) of the Patents Act 1977, if otherwise good, that the host computer claimed in the patent in suit is not present in the UK, but is connected to the rest of the apparatus claimed in the patent."
From Wikipedia:
> Questions of validity were never considered by the court.
> Edited because I didn't know that some European countries accept software patents.
European patents are granted at the European Patent Office (individual european countries also have their own patent offices, whose patents can only be enforced in their home country).
You can, but fraunhofer certainly isn't trying!
> At least the patent licenses usually used with MPEG mean that private use of open source implementations is free.
0_o that is not at all the truth.
So yes you are correct, but in effect it might as well be patent free as far as a 3rd party end user is concerned as they have been provided what is in effect a safe harbour.
https://aomedia.org/membership/members/
AV1 decoding has already been in Chrome and Firefox since at least a year ago. We're just waiting for hardware decoding and encoding support now, which should start appearing this year.
The next version of Chrome will also support the AV1-based AVIF image format this month:
https://chromium-review.googlesource.com/c/chromium/src/+/22...
YouTube, Netflix, and Amazon/Twitch are also likely to not support VVC (and some of them don't support h.265 either) for their streaming services.
FWIW, Firefox supports it already (behind the image.avif.enabled knob).
They should swallow so people again come back and say they are doing it to kill competition? Just received bill from doctor's visit for hundreds of dollar, I guess things are not gonna be free in 2020 after all.
In the early days of MP3, all MP3 rippers and players were built off of their implementation.
Hardware and software companies had to license in order to play MP3 files. As such there was not native support for MP3's for quite some time.
In the late 90's right around the explosion of MP3's on the internet, Fraunhofer was going after companies for doing so.
In my humble opinion, that license mess set back innovation in the portable audio space by a good 5 years.
Seeing all this, I'm convinced that copyright in general and patent system in particular does more harm than good by slowing down the technical progress of the humanity as a whole for the sake of some already rich people becoming a bit richer.
The initial idea behind patent system was sensible, but the way it's abused now... I mean, it could work in today's world as intended if patents lasted a year or two, not what is effectively eternity.
There are plenty of societies that don't respect intellectual property and copyright. And those societies don't innovate at the rate as those who do.
There are certainly abuses in the copyright, trademark, and patent systems. But throwing out the baby with the bathwater is not the answer. Identifying the abuses and improving the system is the answer.
Perhaps that progress won't happen at the rate you'd prefer, but it's significantly better than burning the whole thing to the ground.
Changing patent length (pc's suggestion) is hardly seems like burning the whole thing to the ground.
In the US, if I understand correctly, there are ways to defend against patent abuse, though it typically involves costly legal fees. Fees many cannot justify in spending.
Addressing patent challenges may be a way to protect innovators, both those who should gain reward for their innovations, and those who seek to grain reward by building upon innovations.
That's the key: In the U.S. Good luck stopping someone infringing on your patient in a good number of other countries.
What it does do though is allow someone to start from zero, and catch up to the rest of the competition very fast and cheap. They can then offer their "product" cheaper in order to gain market share. As long as they are getting/keeping customers, there's no need to innovate. You can have a viable business without spending tons of cash on R&D, and if you're making money that way, who cares?
*I am no way endorsing this kind of business model, but it exists and does well.
For one, this claim suffers from a correlation/causation issue. But also, do you have an actual citation for research which shows this is true?
Comparable examples being ?
That's not necessarily so - for example, during the industrial revolution, Germany largely overtook the British in mechanical engineering skill during a period where the British had copyright but before the Germans got it: https://www.wired.com/2010/08/copyright-germany-britain/
Also, it's quite possible the causation goes the other way: A society might not succeed at innovation because of IP laws - it's just as possible that because a society has succeeded at innovation, it passes IP legislation, in order to 'kick away the ladder'. But just like regulatory capture shows in general, legislation that helps yesterday's winners seek rent is not necessarily the same (and is often in fact the opposite) of legislation to help tomorrow's winners see the light of day.
That sounds nice and reasonable, in theory.
In practice, all the money is behind expanding the copyright and patent systems. When is the last time the duration of copyright terms was shortened? When is the last time the patent system was adapted to be less draconian and less protective of those poor, poor multinational corporations that somehow end up holding all those patents?
Spoiler alert - that has happened exactly never. Instead, all we get is 'harmonization' which always means extending terms and giving those laws more teeth, to match the strictest law implemented anywhere in the world.
Steamboat Willy's copyright expires January 1st, 2024. How much do you want to bet that Disney will be pushing for another copyright term extension before then?
> Perhaps that progress won't happen at the rate you'd prefer, but it's significantly better than burning the whole thing to the ground.
The current systems only ever get stricter. Where are the much shorter terms? Where is the recognition that cooperation and remixing fosters innovation and progress and as such should be encouraged, not punished? Where is the PTO following the actual law that says math and logic (i.e. software) are not patentable? Etc, etc, etc.
The abuses have been very well documented over the past several decades. There has been zero progress on incremental improvements. Tell me again how you propose we improve the system without a drastic overhaul?
I'm not 100% sure about this. Just look at how hard China laughs at intellectual property and copyright in general and tell me if you still think the same.
Not being burdened by expensive license fees is a competitive advantage.
But then again, also not having to care about workers right is...
This needs to vary based on field. Some areas, like drugs, take forever to get to market and have exorbitant development costs to recover. (Though, there, other abuses need to get fixed, like renewing patent lifetime with slightly different applications or formulation.
Unfortunately, most companies require that any patent awarded to an engineer in their employ is automatically assigned to the company.
Get rid of that loophole and employees will be able to license their patents as they see fit. Of course this is fraught with practical difficulties, but some kind of compromise could be reached.
Never happen, obviously.
Or did it push it forward? If not for the licensing at the time, would it have been developed? Would it be allowed to be used by anyone just paying for the technology?
I think a more subtle "open-sourcing" of this IP could still be possible, though. Maybe one that still requires that large corporate players that are going to sell their derivative products, acquire a license the traditional way (this is, after all, what the contributors to the codec's patent-pool and R&D efforts based their relative-R&D-labor-contribution negotiations around: that each contributor would end up paying for the devices of theirs that run the codec.)
Maybe there could be a foundation created under the stewardship of the patent-pool itself, which nominally pays the same per-seat/per-device licensing costs as every other member, but where this money doesn't come from revenue but rather is donated by those other members; and where this foundation then grants open-source projects an automatic but non-transferrable license to use the technology.
So, for example, a directly open-source project (e.g. ffmpeg) would be granted an automatic license (for its direct individual users); but that license wouldn't transfer to software that embeds it. Instead, other open-source software (e.g. Handbrake, youtube-dl, etc.) that embeds ffmpeg would acquire its own automatic license (and thus be its own line-item under the foundation); while closed-source software that embeds ffmpeg would be stuck needing a commercial license.
Is there already a scheme that works like this?
Fraunhofer gets roughly 30% of its funding from public sources, the remainder is raised on a per-project basis. It's a fair assumption that those industry partners provided some funds towards the development here. Maybe they even covered all the payroll costs for the involved scientists for the duration of the project.
And yet more income means more money for other research projects. Maybe ones that are not as commercially interesting, or for which a partner decides to terminate a contract rather unexpectedly. While I am also a fan of OSS and would love for work like this to either have no patents or a liberal patent grant, I can also appreciate the desire to fund your research institute.
Why would the taxpayer fund anything that isn’t open and free. Crazy.
The argument that anything funded by tax money must be open is a very fair stance. Though the line gets very blurry when you mix various sources of funding like this. To the best of my recollection I have yet to be paid from any public funding (rather than project specific funding raised from the industry, for example).
Personally, I have no qualms with the funding model, but other points of view are presumably equally valid.
Does it mean money, or not? Because if not, then Fraunhofer does not exist.
But there is absolutely a problem here, because said 'mega businesses' actually should have a strategic imperative to want to make internet technologies more widespread.
Why on earth would MS want to limit their main line of business for a tiny big of IP related revenue?
It would seem to me, that G, MS, Huawei and all of the various patent holders should be trying their best to remove any and all barriers to adoption. There are enough bureaucratic hurdles in the way to worry about, let alone legal concerns.
Even if MS or whoever had to buy out some laggard IP owners who didn't want to play ball, it would probably still make sense for them.
Fraunhofer or anyone else are not in that situation, but the behemoths running vast surpluses are, it just seems shortsighted for them to hamstring any of this.
Fraunhofer has a budget of over 2 billion Euros, and 30% of their money comes from public funding. They run over 70 institutes, so they do much more than this.
The root problem here is that nowadays public research is funded with industry money, which means there has to be a return on investment, hence the patents. In fact, this has metastasized into universities being graded by their patent portfolio volume. So I would expect there to be patents even if 100% of the funding came from the tax payer.
It would have been possible to do the whole process just with public money and zero patents. In fact, I would love it if some research team collated all the patent tax payments across the population of Germany and compare the bottom line cost for the country.
I wager it would have been cheaper without patents, too.
Both h264 and h265 have these implementations, I think FFMPEG library has both under the terms of GPLv2.
The decoders are almost completely useless. The video codec, at least the decoder, needs to be in the hardware, not in software.
Mobile devices just don’t have the resources to run the decoders on CPU. The code works on PC but consumes too much electricity and thermal budget. Even GPGPUs are not good enough for the job, couple generations ago AMD tried to use shader cores for video codecs, didn’t work well enough and they switched to dedicated silicon like the rest of them.
You'd be surprised at how often these are used.
Even Raspberry Pi has a hardware decoder: https://github.com/Const-me/Vrmac/tree/master/VrmacVideo
Linux Kernel support for HEVC is WIP but I’m pretty sure they’ll integrate eventually.
is this because of licensing/copyright?
But the very first Pi already had hardware h264 decoder (and even encoder!) which didn’t need any extra keys to work. No idea how they did it, maybe the license was included in the $25 price. Pi 1 was launched in 2013, at that time h264 has been already widespread, while mpeg-2 use was declining.
I think that’s why they did not include the license. It increases price for all users but only useful for very few of them, who connected a USB DVD or BluRay drives to their Pi-s.
A GPL implementation doesn't guarantee a patent grant. Even if you wrote the code yourself, even just for your personal use, your own work could still be illegal to use due to lacking a patent license from the original patent holders.
Be careful about using implementations of H.26x codecs, because in countries that recognize software patents the code may be illegal to use, regardless whether you've got a license for the code or not. Even when a FLOSS license says something about patent grants, it's still meaningless if the code author didn't own the patents to grant.
I'm sure it does better, but I'm equally sure it'll turn out to be an incremental benefit in practice.
> I wonder how many more CPU instructions are required per second of decoded video compared to HEVC.
CPU cycles are cheap. The real cost is the addition of yet another ?!@$!!#@ video codec block on every consumer SoC shipped over the coming decade.
Opinionated bile: video encoding is a Solved Problem in the modern world, no matter how much the experts want it to be exciting. The low hanging fruit has been picked, and we should just pick something and move on. JPEG-2000 and WebP failed too, but at least there it was only some extra forgotten software. Continuing to bang on the video problem is wasting an absolutely obscene amount of silicon.
No. In the end you transmitted a file that has a certain size. It doesn't really matter if you save that file or just use a volatile buffer.
As long as someone sees the benefit, people will keep pursuing it. Compression (as like all tech) is a moving target with the platforms regularly improving on many axes.
This is part of the problem. What is an "objective" measure of perceptual fidelity to the original?
https://openaccess.thecvf.com/content_cvpr_2018/papers/Zhang...
For comparison, HEVC claimed 50 to 60% compared to AVC. You can compare with reality...
Objective video quality is a tricky thing to do and you can easily come out with a codec that’s great at fine details in foreground people but terrible at high speed panning of water and trees. Depending on the material depends where you want the quality to go.
For lower res you will see lower percentage. You see the same claims from other codec as well.
https://en.wikipedia.org/wiki/Essential_Video_Coding
There are 3 video coding formats expected out of (former) MPEG this year:
https://www.streamingmedia.com/Articles/Editorial/Featured-A...
So this isn't necessarily the successor to HEVC (except that it is, in terms of development and licensing methods).
Maybe. On the other hand, maybe not. Leoanardo Chiariglione, founder and chairman of MPEG, thinks MPEG has for all practical purposes ceased to be:
https://blog.chiariglione.org/a-future-without-mpeg/
The disorganised and fractured licensing around HEVC contributed to that. And, so far, VVC's licensing looks like it's headed down the same path as HEVC.
Maybe AV1's simple, royalty-free licensing will motivate them to get their act together with VVC licensing.
NVIDIA's DLSS 2.0 supersampling is already moving into that direction.
https://papers.nips.cc/paper/9127-deep-generative-video-comp...
How do you know this?
TV hardware is on par with browsers. Anything is a program.
So I guess it comes down to 266 hw support. Or powerful CPUs that can push sw decoding?
MPEG, WiFi, GSM…
IMHO, intentional standards must be implementable without any patent fees, or they are very bad standards.
"ECMA for instance has made all the standards for DVD and optical disks. There were 5 recording formats. So there you are a little bit uneasy, of course. And again after a few beers I can ask the people in the room. Why do you want to have 5 formats? Do you still call that standardization? The answer is always the same: You are well paid. Shut up"
https://youtu.be/wITyO71Et6g?t=226This is really just a press release, what's actually new? Can it be implemented efficiently in hardware?
It seems like the major trade-off being taken right now is along lines of using more memory to buffer additional frames. This can help you in certain scenarios, but in the general case, you cannot ever hope that a prior frame of video has any bearing on future frames of video. It is just exceedingly likely that most frames of video look much like prior frames. So, you can certainly play this game to a point, but you will quickly find yourself on the other end of the bell curve.
You can also play games with ML, but I argue that you are going even further from the fundamental "truth" of your source data with this kind of technique, even if it appears to be a better aesthetic result in isolation of any other concern.
There are also lots of one-off edge cases that have always been impossible to address with any interframe video compression scheme. Just look at the slowmo guys on youtube dump confetti on a 4K camera. No algorithm except for the dumbest intraframe techniques (i.e. JPEG) can faithfully reproduce scenes with information this dense, and usually at the expense of dramatic bandwidth increases.
Bandwidth is cheap and ubiquitous. I say we just use the algorithms that are the fastest and most efficient for our devices. We aren't in 2010 sucking 3G or edge through a straw anymore. Most people can get 20+mbps in their smartphones in decently-populated areas.
As for "information-dense scenes": Pathologic cases such as the HBO intro screen are encoded into modern codecs as noise, and regenerated client-side, because there's no actual information there. These scenes are either engineered or pure noise.
webrtc based video chats are all still using h264, did they not adopt 265 yet for technical or licensing reasons? what is the likelihood of broad browser support for h266 anytime soon?
Is that with x265 built into both browsers ? I build it into mine but I don't think it is the default for ffmpeg.
This makes it great for a company like Netflix or YouTube, but less good for one-to-one and/or battery sensitive use cases like video calls. However, specialized chips help, and some mobile devices can record in HEVC in real time (mine from 2019 can). I believe current smartphones have HEVC encoding hardware, but I'm struggling to find a source for that right now.
I haven't seen the details of this new codec yet, but it's quite possible it also has a large encoding cost which will make it better suited to particular use cases, as opposed to a blanket upgrade.
[1] http://www.praim.com/en/news/advanced-protocols-h264-vs-h265...
1: https://support.apple.com/en-gb/HT207022
2: https://www.qualcomm.com/snapdragon/processors/comparison
The latter isn't practical, I'll eat the couple hundred MB in order to save a lot of time.
"Sequel to Dark Night starring Joaquin Phoenix"
Of course we will not have to film movies in the first place then. We will just put a description into a compressor start watching.
I'm not sure if 265 is worth spending efforts on now when 266 is about to crash the party, and will be equally adopted at least "equally poorly"
It's just a slow percolation throughout the ecosystem as people buy new hardware and video servers selectively send the next-generation streams to those users.
The effort on h.265 has already been spent. Now it looks like h.266 is the next generation. It's going to be years before chips for it will be in devices. That's just how each new generation works.
AV1 is supposed to be 30% better than HEVC and they claim H.266 is 50% better than HECV. This would mean that H.266 is roughly 30% better than AV1. By better I'm always referring to the bandwidth/space needed.
But take this with more than a grain of salt since bandwidth/space are only one of many things that matter and also these comparisons are dependent on so many things like resolution, material (animatic/real), etc. etc.
Regardless, these early performance claims are most likely complete bullshit.
Given that model sizes for decoding seem like they'll be on the order of many gigabytes, it will be impossible to run AI decompression in software, but will need chips, and chips that are a lot more complex (expensive?) than today's.
I think AI compression has a good chance of coming eventually, but in 10 years it will still be in research labs. There is absolutely no way it will have made it into consumer chips by then.
Are normal 1080p videos going to see this fabled 50% savings over h.265? Or is the 50% only for 4K/8K, while 1080p gets maybe only 10-20% savings?
The press release unfortunately seems rather ambiguous about this.
"for equal perceptual quality"
Put a different way: We can fool your eyes/brain into thinking you are looking at the same images.
For most consumer use cases where the objective is to view images --rather than process them-- this is fine. The human vision system (HVS, eyes + brain processing) is tolerant of and can handle lots of missing or distorted data. However, the minute you get into having to process the images in hardware or software things can change radically.
Take, as an example, color sub-sampling. You start with a camera with three distinct sensors. Each sensor has a full frame color filter. They are optically coupled to see the same image through a prism. This means you sample the red, green and blue portions of the visible spectrum at full spatial resolution. If we are talking about a 1K x 1K image, you are capturing one million pixels of each, red, green and blue.
BTW, I am using "1K" to mean one thousand, not 1024.
Such a camera is very expensive and impractical for consumer applications. Enter the Bayer filter [0].
You can now use a single sensor to capture all three color components. However, instead of having one million samples for each components you have 250K red, 500K green and 250K blue. Still a million samples total (that's the resolution of the sensor) yet you've sliced it up into three components.
This can be reconstructed into full one million samples per color components through various techniques, one of them being the use of polyphase FIR (Finite Impulse Response) filters looking across a range of samples. Generally speaking, the wider the filter the better the results, however, you'll always have issues around the edges of the image. There are also more sophisticated solutions that apply FIR filters diagonally as well as temporally (use multiple frames).
You are essentially trying to reconstruct the original image by guessing or calculating the missing samples. By doing so you introduce spatial (and even temporal) frequency domain issues that would not have been present in the case of a fully sampled (3 sensor) capture system.
In a typical transmission chain the reconstructed RGB data is eventually encoded into the YCbCr color space [1]. I think of this as the first step in the perceptual "let's see what we can get away with" encoding process. YCbCr is about what the HVS sees. "Y" is the "luma", or intensity component. "Cb" and "Cr" are color difference samples for blue and red.
However, it doesn't stop there. The next step is to, again, subsample some of it in order to reduce data for encoding, compression, storage and transmission. This is where you get into the concept of chroma subsampling [2] and terminology such as 4:4:4, 4:2:2, etc.
Here, again, we reduce data by throwing away (not quite) color information. It turns out your brain can deal with irregularities in color far more so than in the luma, or intensity, portion of an image. And so, "4:4:4" means we take every sample of the YCbCr encoded image, while "4:2:2" means we cut down Cb and Cr in half.
There's an additional step which encodes the image in a nonlinear fashion, which, again, is a perceptual trick. This introduces Y' (Y prime) as "luminance" rather than "luma". It turns out that your HVS is far more sensitive to minute detail in the low-lights (the darker portions of the image, say, from 50% down to black) than in the highlights. You can have massive errors in the highlights and your HVS just won't see them, particularly if things are blended through wide FIR filters during display. [3]
Throughout this chain of optical and mathematical wrangling you are highly dependent on the accuracy of each step in the process. How much distortion is introduced depends on a range of factors, not the least of which is the way math is done in software or chips that touch every single sample's data. With so much math in the processing chain you have to be extremely careful about not introducing errors by truncation or rounding.
We then introduce compression algorithms. In the case of motion video they will typically compress a reference frame as a still and then encode the difference with respect to that frame for subsequent frames. They divide an image into blocks of pixels and then spatially process these blocks to develop a dictionary of blocks to store, transmit, etc.
The key technology in compression is the Discrete Cosine Transform (DCT) [4]. This bit of math transforms the image from the spatial domain to the frequency domain. Once again, we are trying to trick the eye. Reduce information the HVS might not perceive. We are not as sensitive to detail, which means it's safe to remove some detail. That's what DCT is about.
So, we started with a 3 sensor full-sampling camera, reduced it to a single sensor and three away 75% of red samples, 50% of green samples and 75% of blue samples. We then reconstruct the full RGB data mathematically, perceptually encode it to YCbCr, apply gamma encoding if necessary, apply DCT to reduce high frequency information based on agreed-upon perceptual thresholds and then store and transmit the final result. For display on an RGB display we reverse the process. Errors are introduced every step of the way, the hope and objective being to trick the HVS into seeing an acceptable image.
All of this is great for watching a movie or a TikTok video. However, when you work in machine vision or any domain that requires high quality image data, the issues with the processing chain presented above can introduce problems with consequences ranging from the introduction of errors (Was that a truck in front of our self driving car or something else?) to making it impossible to make valid use of the images (Is that a tumor or healthy tissue?).
While H.266 sounds fantastic for TikTok or Netflix, I fear that the constant effort to find creative ways to trick the HVS might introduce issues in machine vision, machine learning and AI that most in the field will not realize. Unless someone has a reasonable depth of expertise in imaging they might very well assume the technology they are using is perfectly adequate for the task. Imagine developing a training data set consisting of millions of images without understanding the images have "processing damage" because of the way they were acquired and processed before they even saw their first learning algorithm.
Having worked in this field for quite some time --not many people take a 20x magnifying lens to pixels on a display to see what the processing is doing to the image-- I am concerned about the divergence between HVS trickery, which, again, is fine for TikTok and Netflix and MV/ML/AI. A while ago there was a discussion on HN about ML misclassification of people of color. While I haven't looked into this in detail, I am convinced, based on experience, that the numerical HVS trickery I describe above has something to do with this problem. If you train models with distorted data you have to expect errors in classification. As they say, garbage-in, garbage-out.
Nothing wrong with H.266, it sounds fantastic. However, I think MV/ML/AI practitioners need to be deeply aware of what data they are working with and how it got to their neural network. It is for this reason that we've avoided using off-the-shelf image processing chips to the extent possible. When you use an FPGA to process images with your own processing chain you are in control of what happens to every single pixel's data and, more importantly, you can qualify and quantify any errors that might be introduced in the chain.
[0] https://en.wikipedia.org/wiki/Bayer_filter
[1] https://en.wikipedia.org/wiki/YCbCr
[2] https://en.wikipedia.org/wiki/Chroma_subsampling
I have books on AI that are thirty years old. I think I can say they cover somewhere between 80% and 90% (if not more) of what AI is today. The difference is computing that is thousands, millions, of times faster, massive amounts of storage, etc. In other words, one could very well argue we haven't done much in 30 years other than build faster computers.
If we take the example of a puppy, it seems to generalize pretty well using something like one-shot learning, but is it? I cannot confirm for sure how much data a puppy has already digested before being able to do what we could call "one shot learning". So maybe, the exposure to data is already there, waiting for a specialization toward a particular task.
Giving the ability to a network to be probabilistic enable it to do inference using uncertainty, which is clearly a neat feature when you are gravitating toward AGI for scene understanding.
In the case of video compression, scene understanding may introduce more artifacts IMO: Even if the scene is captured with high end cameras, on a pixel level basis, the edges will never be perfectly neat. I think this will decrease the ability of any network to "understand" which object is at the edges, this results in low classification rates on them, resulting in bad compression/decompression quality (?) for features that are important to the human eye.
All in all, I'm not sure that NN are the right tool for this kind of problems. But we are diverging from the main subject VVC, Thanks for the very interesting comments :)
My internet connection speeds and hard drive space have increased much faster than my CPU speeds (internet being basically a free upgrade).
So I don't appreciate new codecs coming out and obsoleting my hardware to save companies a few cents on bandwidth. H.264 got a good run in, but there isn't a "universal" replacement for it where I can buy hardware with decoding support that will work for at least 5-10 years.
This isn't about saving companies a few cents on bandwidth. It's about halving internet traffic, about doubling the number of videos you can store on your phone. That's pretty huge. You can still get h.264 video in 1080p on YouTube so your computer is still meeting the expectations it was manufactured for.
The bigger problem that I didn't mention is with videoconferencing: FaceTime is hardware accelerated and has no issues with 720p, but anything WebRTC seems to prefer VP8 or VP9 codecs, which fails on my Mini and strains my 2015 MBP. Feels like a waste of perfectly good hardware to me.
https://github.com/erkserkserks/h264ify
Or is there no longer an H.264 4k version?
You could say something similar about any other technological advance.
I don't mind replacing a machine after 8 years of service, but h.265 still isn't supported by Google/YouTube, and Apple refuses to add hardware decoding for VP8/VP9, so there's no universal codec that will work as efficiently as h.264 did on my Mini and multitude of MacBooks and iPhones all this time.
Can't do that anymore without either buying a new computer or youtube-dl and recompress the bigger version (which takes hours for minutes of video on my poor machine).
But FWIW, H.266 isn't going to be in any sort of wide use for a few years. Buy something that supports H.265 and you'll probably be good for at least 5 years.
And it's not like any actual 4K content (besides porn, real, nature, or otherwise) actually exists. Broadcast and movie media is done in 2K then extrapolated and scaled to "4K" for streaming services.
What is 2K? I've never even heard of a "2K" camera. Where did you get the idea things are being filmed in "2K" and being scaled to 4K?
Genuinely curious where you're getting this information from. Or are you confused because 1080p refers to the vertical resolution while 4K refers to the horizontal resolution?
edit: here's another https://old.reddit.com/r/cordcutters/comments/9x3v4e/just_le...
The top link in the reddit thread disproves what you're saying though:
https://4kmedia.org/real-or-fake-4k/
Somewhere between a third and a half of films are listed as "real 4K".
So there is actually tons of real 4K content. (And the list is just films -- there are plenty of streaming TV shows in real 4K too, like Mrs Maisel.)
There might be another reason for the misperception -- it's true that film editing is generally done in something lower-quality like compressed 1080p, but that's just for speed/space while you work. All the clips "point" to the 4K originals, so when the final master is produced, it's still produced out of that "real" 4K.
Largely due to the CPU over head of H265, though I am not sure why more people do not use GPU encoding over CPU Encoding, I have never been able to notice the difference visually