2,065 karma · joined July 20, 2008
Ask me anything.
A proper implementation will make sure the worst-case latency is accounted for and not cherry-pick the best case.
It's left as an exercise to the reader which methods of lying are being used in this case.
We cannot tell volunteers to not do something (i.e help Microsoft) in the same way we cannot tell volunteers to do something.
To clarify, Microsoft didn't offer a one-time payment to fix any bug. A volunteer explained for free that a command line option had changed between FFmpeg versions.
Microsoft offered a one-off payment instead of a support contract.
/s
v210_planar_pack_8_ssse3: 402.5
v210_planar_pack_8_avx: 413.0
v210_planar_pack_8_avx2: 206.0
v210_planar_pack_8_avx512: 193.0
v210_planar_pack_8_avx512icl: 100.0
23x speedup. The compiler isn't going to come up with some of the trickery to make this function 23x faster.
800% is nothing.
How is this "impossible" if the video clock and audio clock are on different physical devices?
You cannot iterate quickly based on the glacial embedded industry. I am amazed by the way we are quoted 3 month wait for a shipment (even pre-covid) without anyone batting an eyelid.
You're a bit like the satellite industry looking at Starlink and saying "well it's not true satellite, doesn't have X, Y and Z" and pointing to a few specialist applications where legacy satellite is highly suited and ignoring the masses of business applications that have moved to Starlink.
The state of the embedded industry continues to amaze me.
The lightweight macro layer in ffmpeg takes care of v prefixes.
In FFmpeg, x264 and dav1d there are many different examples of code that couldn't be written in intrinsics or other abstraction layer.
https://twitter.com/FFmpeg/status/1705543447245988245?t=Ul9e...
There have been many SIMD abstraction layers created in the past but none of them will beat the raw speed of handwritten assembly. Try and implement something like vpternlogd in one of these abstraction layers.