FFmpeg lands CLI multi-threading as its "most complex refactoring" in decades
phoronix.com
phoronix.com
It wasn't that long ago that reading, processing, and rendering the contents of a single image took a noticeable amount of time. But both hardware and software techniques have gotten significantly faster. What may have made sense many years ago (lots of workers on a frame) may not matter today when a single worker can process a frame or a group of frames more efficiently than the overhead of spinning up a bunch of workers to do the same task.
But where to move that split now? Ultra-low-end CPUs now ship with multiple cores and you can get over 100 easily on high-end systems, system RAM is faster than ever, interconnect moves almost a TB/sec on consumer hardware, GPUs are in everything, and SSDs are now faster than the RAM I grew up with (at least on continuous transfer). Basically the systems of today are entirely different beasts to the ones commonly on the market when FFmpeg was created.
This is tremendous work that requires lots of rethinking about how the workload needs to be defined, scheduled, distributed, tracked, and merged back into a final output. Kudos to the team for being willing to take it on. FFmpeg is one of those "pinnacle of open source" infrastructure components that civilizations are built from.
Using either processes or threads should bear no difference in terms of die space. In a die there are cores, each core with dual streams of computation, that are interleaved to compensate for memory latency, so while one waits the other computes and vice versa. It doesn't have anything to do with processes or (virtual) threads, which are software constructs.
Update: Found here! https://www.youtube.com/watch?v=Z4DS3jiZhfo&t=1221s
IIUC all current encoders that support parallelism work by multiple threads working on the same frame at the same time. Often times the frame is split into regions and each thread focuses on a specific region of the frame. This approach can have a (usually small) quality/efficiency cost and requires per-encoder logic to assemble those regions into a single frame.
What if instead/additionally different keyframe segments are processed independently? So if keyframes are every 60 frames ffmpeg will read 60 frames pass that to the first thread, the next 60 to the next thread, ... then assemble the results basically by concatenating them. It seems like this could be used to parallelize any codec in a fairly generic way and it should be more efficient as there is no thread-communication overhead or splitting of the frame into regions which harms cross-region compression.
Off the top of my head I can only think of two issues:
1. Requires loading N*keyframe period frames into memory as well as the overhead memory for encoding N frames.
2. Variable keyframe support would require special support as the keyframe splits will need to be identified before passing the video to the encoding threads. This may require extra work to be performed upfront.
But both of these seem like they won't be an issue in many cases. Lots of the time I'd be happy to use tons of RAM and output with a fixed keyframe interval.
Probably I would combine this with intra-frame parallelization such as process every frame with 4 threads and then run 8 keyframe segments in parallel. This way I can get really good parallelism but only minor quality loss from 4 regions rather than splitting the video into 32 regions which would harm quality more.
Plus, computers aren't quad cores any more, people with powerful streaming rigs probably have 8 or 16 cores; and key frames aren't every second. Suddenly you're in this hellish world where you have to balance latency, CPU utilization and encoding efficiency. 16 cores at a not-so-great 8 seconds of extra latency means terrible efficiency with a key frame every 0.5 second. 16 cores at good efficiency (say, 4 seconds between key frames) means terrible 64 second of extra latency.
That's a good point. In the general case of reading from a pipe you need to buffer it somewhere. But for file-based inputs the buffering concerns aren't relevant, just the working memory.
It's one of several reasons why live streams of this type are often 10-30 seconds behind live.
* Of course it also depends on where in the pipeline they hook in - some take the feed directly, in which case every frame is essentially a key frame.
What make you think that? I very much care about encoding performance (for a fixed quality level) for offline use.
Also to be able to seek anywhere in the steam without decoding all previous frames.
It also has a lossless codec ffv1 where the entropy coder doesn't reset, so it truly can't be multithreaded.
I'd assume if each thread is working on its own key frame, it would be difficult to make b-frames work? Live content also probably makes it hard.
Big win too! This is going to really speed things up!
Agree with the general premise, of course, if I've got 10 different videos encoding at once then I don't need additional efficiency because the CPU's already maxed out.
You could also imagine they might apply some kind of heuristic to decide to re-encode something based on some condition... Like fine tune encoder settings when a title becomes popular. No idea if they do that, just using some imagination.
Also, their devs likely want fast feedback on changes - I imagine they might have CI running changes on some standard movies, checking various stats (like SNR) for regressions. Everybody loves if their CI finishes fast, so you might want to compress even a single movie in multiple threads.
(Plus a couple more "legacy" encodes with PIFF instead of CENC for ancient devices, probably.)
New tech advances, sure, they probably do re-encode everything sometimes - even knocking a few MB off the size of a movie saves a measurable amount of $$ at that scale. But are there frequent enough tech advances to do that more than a couple of times a year..? The amount of difficult testing (every TV model group from the past 10 years, or something) required for an encode change is horrible. I'm sure they have better automation than anyone else, but I'm guessing it's still somewhat of a nightmare.
Youtube, OTOH, I really can imagine having thousands of concurrent ffmpeg processes.
Their tech blog and tech presentations discuss many of the requirements and steps involved for encoding source media to stream to all the devices that Netflix supports.
The Netflix tech blog: https://netflixtechblog.com/ or https://netflixtechblog.medium.com/
Netflix seems to use AWS CPU+GPU for encoding, whereas YouTube has gone to the expense of producing an ASIC to do much of their encoding.
2015 blog entry about their video encoding pipeline: https://netflixtechblog.com/high-quality-video-encoding-at-s...
2021 presentation of their media encoding pipeline: https://www.infoq.com/presentations/video-encoding-netflix/
An example of their FFmpeg usage - a neural-net video frame downscaler: https://netflixtechblog.com/for-your-eyes-only-improving-net...
Their dynamic optimization encoding framework - allocating more bits for complex scenes and fewer bits for simpler, quieter scenes: https://netflixtechblog.com/dynamic-optimizer-a-perceptual-v... and https://netflixtechblog.com/optimized-shot-based-encodes-now...
Netflix developed an algorithm for determining video quality - VMAF, which helps determine their encoding decisions: https://netflixtechblog.com/toward-a-practical-perceptual-vi..., https://netflixtechblog.com/vmaf-the-journey-continues-44b51..., https://netflixtechblog.com/toward-a-better-quality-metric-f...
This is overrated - of course that's how you do it, what else would you do?
> Mean-squared-error (MSE), typically used for encoder decisions, is a number that doesn’t always correlate very nicely with human perception.
Academics, the reference MPEG encoder, and old proprietary encoder vendors like On2 VP9 did make decisions this way because their customers didn't know what they wanted. But people who care about quality, i.e. anime and movie pirate college students with a lot of free time, didn't.
It looks like they've run x264 in an unnatural mode to get an improvement here, because the default "constant ratefactor" and "psy-rd" always behaved like this.
That's not what has been done previously for adaptive streaming. I guess you are referring to what encoding modes like CRF do for an individual, entire file? Or where else has this kind of approach been shown before?
In the early days of streaming you would've done constant bitrate for MPEG-TS, even adding zero bytes to pad "easy" scenes. Later you'd have selected 2-pass ABR with some VBV bitrate constraints to not mess up the decoding buffer. At the time, YouTube did something where they tried to predict the CRF they'd need to achieve a certain (average) bitrate target (can't find the reference anymore). With per-title encoding (which was also popularized by Netflix) you could change the target bitrates for an entire title based on a previous complexity analysis. It took quite some time for other players in the field to also hop on the per-title encoding train.
Going to a per-scene/per-shot level is the novely here, and exhaustively finding the best possible combination of QP/resolution pairs for an entire encoding ladder that also optimizes subjective quality – and not just MSE.
This is unnecessary if the encoder is well-written. It's like how some people used to run multipass encoders 3 or 4 times just in case the result got better. You only need one analysis pass to find the optimal quality at a bitrate.
I think it should be possible to decide in one shot in the codec though. My memory is that codecs (image and video) have tried implementing scalable resolutions before, but it didn't catch on simply because dropping resolution is almost never better than dropping bitrate.
Netflix tries to optimize the encoding parameters per shot/scene.
from the dynamic optimization article:
- A long video sequence is split in shots ("Shots are portions of video with a relatively short duration, coming from the same camera under fairly constant lighting and environment conditions.")
- Each shot is encoded multiple times with different encoding parameters, such as resolutions and qualities (QPs)
- Each encode is evaluated using VMAF, which together with its bitrate produces an (R,D) point. One can convert VMAF quality to distortion using different mappings; we tested against the following two, linearly and inversely proportional mappings, which give rise to different temporal aggregation strategies, discussed in the subsequent section
- The convex hull of (R,D) points for each shot is calculated. In the following example figures, distortion is inverse of (VMAF+1)
- Points from the convex hull, one from each shot, are combined to create an encode for the entire video sequence by following the constant-slope principle and building end-to-end paths in a Trellis
- One produces as many aggregate encodes (final operating points) by varying the slope parameter of the R-D curve as necessary in order to cover a desired bitrate/quality range
- Final result is a complete R-D or rate-quality (R-Q) curve for the entire video sequence
That's the problem - if the encoding parameters need to be varied per scene, it means you've defined the wrong parameters. Using a fixed H264 QP is not on the rate-distortion frontier, so don't encode at constant QP then. That's why x264 has a different fixed quality setting called "ratefactor".
Possibly Netflix statistics are way better.
And that was years ago, I wouldn't be surprised to learn it's a bigger number now.
* I would be very surprised if Netflix even uses vanilla ffmpeg
I believe the vast majority of ffmpeg usages are web services, or one off encodings.
Subjectively, me compressing my holiday video is much more important than Netflix re-compressing a million of them.
If you have a queue of 100k videos to process and a cluster of 100 cores, assigning a video to each core as it becomes available is the most efficient way to process them, because your skipping the thread joining time.
Anytime there is a queue of jobs, assigning the next job in the queue to the next free core is always going to be faster than assigning the next job to multiple cores.
They use custom hardware just for encoding.
fyi they have to transcode over 500h of videos per minute. So multiple that by all the formats they support.
They operate at an insane scale, Netflix looks like a garage project for comparison.
For encoding, recently, they've built their own ASIC to deal with H264 and VP9 encoding (for 7-33x faster encoding compared to CPU-only): https://arstechnica.com/gadgets/2021/04/youtube-is-now-build...
There is a noticeable delay booting up this pipeline for each tool invoke right now. We are working on putting in some optimizations but improvements in FFmpeg will definitely help. https://github.com/trypromptly/LLMStack is the project repo for the curious.
The presentation says it's 700 commits. Was that a separate branch? Or was it slowly merged back to the project?
Well I can look at github I guess
as long as it works for them...
Not that this isn't great. Its fantastic. But TBH its not really going to change my workflow of VapourSynth preprocessing + av1an encoding for "quality" video encodes.
ffmpeg -v quiet -f data -i file.txt -map 0:0 -c text -f data -This is probably 80% of my experience with ffmpeg, to be honest, but the other 20% is invaluable enough anyway.
Later on though I've realized the quality of tesseract's OCR on arbitrary media is often quite bad. Google translates detection and replacement is so much ahead my current image system I'd think I would just somehow reutilize that for my app, either thru public API or browser emulation ...
(ps: and no, it's not Rick Astley/Never Gonna Give You Up)
$ ffmpeg -v quiet -codecs | egrep -i 'gif|ascii'
D.V.L. ansi ASCII/ANSI art
DEV..S gif CompuServe GIF (Graphics Interchange Format)
(“D” and “E” in the first field indicate support for decoding and encoding) dd if=./file.txt
Can you also format your drive with ffmpeg? I'm looking for a more versatile dd replacement.. ffmpeg -f data -i /dev/zero -map 0:0 -c copy -f data - > /dev/sda
is roughly equivalent to to dd status=progress if=/dev/zero of=/dev/sda ffmpeg -v quiet -f data -i file.txt -map 0:0 -c copy -f data -
does.[1] “Encoder 'text' specified, but only '-codec copy' supported for data streams”
How come it only has 4 measly entries in HN, and none got any traction. I've posted a new entry, just for the curiosity of others.
[edit]
found the docs; it's available on Linux[1]. I'm definitely looking into it tonight because it can't be worse than writing ffmpeg CLI filtergraphs!
1: http://www.vapoursynth.com/doc/installation.html#linux-insta...
This is a pretty good (but not comprehensive) db of the filters: https://vsdb.top/
The frameserver is one thing, but an ecosystem of (trustable! open source!) plugins is harder to replicate.
At least we don't need deinterlacers so badly any more, though.
are people just itching for reasons to dive into show & tell or to wax poetic about how they have solved the problem for years? I really don't understand people at all, because I don't understand why people do this. and I'm sure I've done it, too.
Hype for FEAT is beyond sensibility. People with similar FEAT are bristled by this and wish that their projects received even a fraction of FEAT's hype.
I think it's normal.
language is for communicating. don't impede that communication by using unnecessary terms.
It's been threaded since its inception, so it seems somewhat topical.
Now it can read the next chunk of file while decoding a chunk while applying filters while running multi-threaded encoding while writing the previous chunk.
I expect this will be a big help for complicated setups (applying multiple filters, rendering to multiple outputs with different codecs), but probably not much change for a simple 1-input 1-output transcode.
I'm using FFMPEG to encode MP3 with LAME for an audio hosting service and it would be great to improve encode times for long files.
That said, I don't think we still get input buffering (for HLS).
Talk with ChatGPT about it and see if you can do it.
Just Google it.
I wonder why did it took so long for FFmpeg?
BTW, MS Media foundation is a functional equivalent of FFmpeg. It was released as a part of Windows Vista in 2006, and is heavily multithreaded by design.
Couldn't do the same with larger codebases because the context is not enough for the input code and output refactoring.
I have been refactoring code using gpt-4 for some months now and the limiting factor have been the context size.
GPT-4 turbo now have 128k context and I can provide it with larger portions of the code base for the refactors.
When we have millions of tokens of context, based on what I'm experiencing now, I can see that a refactoring like the one made in ffmpeg would be possible. Or not? What am I missing here?
I guess you could use AI to help create a coccinelle semantic patch.
I most certainly have not. At work, I do greenfield development in a specialized problem domain, and I would not trust a model (or, for that matter, a junior developer) to do any kind of refactor in an acceptable manner. (That aside, there's no way I'm goingto expose company code to any sort of outside party without the approval of upper management).
At home, I program for fun and self-improvement, and a big part of both is thinking hard about problems. Why would I want to wreck that with asking a model to do it for me?
What if you have an idea and you tell the computer to implement it and then thoroughly check the code?
About the self-improvement part, I suppose you don't operate your graphics card yourself, but delegate it to your driver.
The LLM is just another tool.
Interesting.
Here's a sample question I just asked the AI:
We are working on a React mapping application. I am going to give you two components from this application. Our goal is to extract functionality to deduplicate code
We can make a single base component, extract code to hooks, or any other strategy which you think is appropriate
Here is travel.tsx:
<snip>
Here is heatmap.tsx:
<snip>
And here is the response:
To deduplicate code and improve maintainability between the `TravelMap` and `Heatmap` components in your React mapping application, we can identify common functionalities and extract them into reusable hooks or components. Here's a strategy considering your code:
1. *Extract Common Hooks*: Both components use similar hooks for managing markers, popups, and map references. We can create custom hooks that encapsulate these functionalities.
2. *Create Shared Utility Functions*: Functions like `initialViewState`, which are identical in both components, can be moved to a separate utility file and imported where needed.
3. *Refactor Map Layers and Sources*: The way map layers and sources are handled is quite similar in both components. A custom hook or component to manage map layers and sources would be beneficial.
4. *Generalize Popup Handling*: The popup handling in both components can be abstracted into a custom hook or a higher-order component.
5. *Combine Common Styling*: Shared styles, such as `MAP_CONTAINER_STYLE`, can be moved to a common CSS module or a styled-component that can be reused.
6. *Refactor Event Handlers*: Functions like `onMapClick`, `onMouseMove`, and `onMouseLeave` have similar structures. These can be part of a custom hook or a utility function with configurable parameters.
7. *Legend and Live Region Components*: These components are used in a similar way in both `TravelMap` and `Heatmap`. Ensure they are designed to be reusable with different props.
8. *Refactor Map Initialization*: The way the map is initialized and configured is quite similar. This can be abstracted into a custom component or hook.
### Example Implementations:
#### 1. Custom Hook for Map Markers:
```typescript // useMapMarkers.ts import { useState, useCallback } from 'react';
export const useMapMarkers = (initialLat, initialLng, world) => { // Similar marker logic from your components // ...
return {
marker,
onMarkerDrag,
onMarkerDragEnd,
updateMarkerPosition,
// other marker related states and functions
};
};
```#### 2. Shared Utility Function:
```typescript // utils/mapUtils.ts export const initialViewState = (map, mapType) => { // Implementation }; ```
#### 3. Custom Hook for Map Layers and Sources:
```typescript // useMapLayers.ts import { useState, useEffect } from 'react';
export const useMapLayers = (props) => { // Logic to handle layers and sources // ...
return {
dataLayer,
patternLayer,
// other layer related states and functions
};
};
```#### 4. Generalized Popup Handling:
```typescript // useMapPopup.ts import { useState } from 'react';
export const useMapPopup = () => { // Popup logic from your components // ...
return {
popupState,
openPopup,
closePopup,
setPopupData,
// other popup related states and functions
};
};
```### Integration:
Once these hooks and utilities are created, integrate them into your `TravelMap` and `Heatmap` components. Replace the duplicated logic with calls to these shared resources. This will significantly reduce code duplication and make future maintenance easier.
### Testing and Refinement:
After refactoring, thoroughly test both components to ensure that all functionalities work as expected. Watch out for any edge cases or specific behaviors that might be affected by the refactoring.
For those suggestions, I might use five out of eight of them, and probably do one or two things differently. But you cannot, with a straight face, say the model did not understand. It clearly did. It suggested reasonable refactors. If being able to refactor means understanding, I guess we have understanding!
I could continue with this conversation, ask it to produce the full code for the hooks (I have in my custom prompt to provide outlines) and once the hooks are complete, ask it to rewrite the components using the shared code.
Have you ever used one of these models?
Cleaning up code also follows some well established patterns, performance work is much less pattern-y.
Codebases like FFMPEG are one of the kind. I bet you need 10 or 100 times more understanding than the react thing you mentioned above.
One day maybe AI can do it, but it probably won't be LLM. It would be something which can understand symbols and math.
> Because refactoring requires understanding, which LLMs completely lack.
<demonstration that an LLM can refactor code>
> Cleaning up code also follows some well established patterns, performance work is much less pattern-y.
Just as writing shitty react apps follow patterns, low-level performance and concurrency work also follow patterns. See [0] for a sample.
> I bet you need 10 or 100 times more understanding
Okay, so a 10 or 100 times larger model? Sounds like something we'll have next year, and certainly within a decade.
> One day maybe AI can do it, but it probably won't be LLM. It would be something which can understand symbols and math.
You do understand that the reason some of the earlier GPTs had trouble with symbols and math was the tokenization scheme, completely separate from how they work in general, right?
[0]: C++ Concurrency in Action: Practical Multithreading 1st Edition https://www.amazon.com/C-Concurrency-Action-Practical-Multit...
It's obvious from context here that the refactoring that was mentioned was specifically around concurrency, not simply cleaning up code.
https://chat.openai.com/share/7c41f59a-c21c-4abd-876c-c95647...
What you've shown is actually a great example of the what folks mean that LLMs lack any sort of understanding. They're fundamentally predict-the-next-token machines; they regurgitate and mix parts of their training data in order to satisfy the token prediction loss function they were trained with.
In the linked example you provided, *you* are the one that needs to provide the understanding. It's a rather lengthly back-and-forth to get that code into a somewhat useable state. Importantly, if you didn't tell it to fix things (sqlite connections over threads, etc.), it would have failed.
And while it's concurrent, it's using threads, so it's not going to be doing any work in parallel. The example you have mixes some IO and compute-bound looking operations.
So, if your need was to refactor your original code to _actually be fast_, ChatGPT demonstrated it doesn't understand nearly enough to actually make this happen. This thread conversation got started around correcting the misnomer that an LLM would actually ever be able to possess enough knowledge to do actually valuable, complex refactoring and programming.
While I believe that LLMs can be good tools for a variety of usecases, they have to be used in short bursts. Since their output is fundamentally unreliable, someone always has to read -- then comprehend -- its output. Giving it too much context and then prompting it in such a way to align its next token prediction with a complex outcome is a highly variable and unstable process. If it outputs millions of tokens, how is someone going to actually review all of this?
In my experience using ChatGPT, GPT4, and a few other LLMs, I've found that it's pretty good at coming up with little bits to jog one's own thinking and problem solving. But doing an actual complex task with lots of nuance and semantics-to-be-understood outright? The technology is not quite there yet.
Let's see how your smooth talking LLM is going to do with things that are not web development or leetcode medium, for which so much stuff has been written. All the best.
alwayshasbeen.png
> The AI effect occurs when onlookers discount the behavior of an artificial intelligence program by arguing that it is not "real" intelligence.[1] > Author Pamela McCorduck writes: "It's part of the history of the field of artificial intelligence that every time somebody figured out how to make a computer do something—play good checkers, solve simple but relatively informal problems—there was a chorus of critics to say, 'that's not thinking'."[2] Researcher Rodney Brooks complains: "Every time we figure out a piece of it, it stops being magical; we say, 'Oh, that's just a computation.'"[3]
> "AI is whatever hasn't been done yet."
> —Larry Tesler
That's very different from code which (perhaps surprisingly) isn't well formalized. The goals are often vague and it's difficult to figure out what is intentional and what incidental behavior (esp. with imperative code).
As someone who was deeply involved in the Go scene since the early 2000s let me emphatically assure you it was not at all clear. Indeed it was a major point of pride among Go enthusiasts that computers could not play it well, for various reasons (some, like the branching factor, one could potentially grant that advances in hardware and software could solve for eventually. Others, like the inherent difficulty in constructing an evaluation function, seemed intractable).
Betting markets at the time of the AlphaGo match still had favorable odds for Sedol, even with the knowledge that Google was super-confident baked in.
It is extreme hindsight-bias of exactly the type the grandparent was talking about to suggest that obviously everybody knew all along that Go was very beatable by "non-real AI".
It's typical for enthusiast of anything to believe their area is special. To be able to judge how feasible besting humans is a question for computer scientists / engineers anyway. You need some knowledge of Go for sure, but don't actually have to be a master Go player.
> It is extreme hindsight-bias of exactly the type the grandparent was talking about to suggest that obviously everybody knew all along that Go was very beatable by "non-real AI".
Maybe it was a surprise for you, but I was familiar with the pre-AlphaGo state of the art, and it was always clear it's just a technical problem. There's no fundamental difference between a game like Go and Chess, it's only the parameters which differ, where Go's space is exponentially larger.
> Others, like the inherent difficulty in constructing an evaluation function, seemed intractable
The branching factor is what makes the evaluation function difficult in Go and is therefore again a technical problem.
> Betting markets at the time of the AlphaGo match still had favorable odds for Sedol, even with the knowledge that Google was super-confident baked in.
This is a non-argument, betting markets are not rational.
Actually it doesn't, as demonstrated by the fact that people have made fairly decent chess implementations in 1K lines of code and things like that. Chess is comparatively easy because it has well-defined rules and well-defined concepts of "good" and "bad". Refactoring something to be multi-threaded is incomparably more complex and any comparison to this is just pointless.
The trick to this is continuous iteration and feedback. It's remarkable how far I've gotten with GPT using these simple primitives and I know I'm not the only one.
I do renames in a big window with my IDE.