FFMPEG from Zero to Hero
ffmpegfromzerotohero.com
ffmpegfromzerotohero.com
ffmpeg -i part0.mp4 -i part1.mp4 -filter_complex "[1:v]scale=1280:720:force_original_aspect_ratio=decrease,pad=1280:720:(ow-iw)/2:(oh-ih)/2[v1]; [0:v] [0:a] [v1] [1:a] concat=n=2:v=1:a=1 [v] [a]" -map "[v]" -map "[a]" out.mp4
to concatenate two videos with different sizes today.My relation with it is almost identical to my relation with any bash scripting more than a one-liner: I have to relearn it every time I want to engage with it.
eg copying without conversion between different formats isn’t really possible when it looks like it should be, because there’s a lot of incorrect handling of timestamps both in libavformat and in files themselves.
In ffmpeg's case, I would be looking for a GUI beyond just "here's all the command line parameters, only in a form"... the hardest thing is taking something I already know what I want to do, and encoding it into the command line, because the way to encode a graph into a command line is not obvious. It is easy in principle but there are so many degrees of freedom in how it is done that I can never remember it all. This is a perfect time for a visual language.
IMO the best option is the best of both worlds. A GUI that outputs the CLI command along with good documentation would be ideal.
For example English orthography is a horrible mess. The inventory of squiggles needed to write English isn't too bad, but the correspondence between the words you know and the correct sequence of squiggles to write them is unnecessarily complicated. No benefit accrues to us from this, and in some other written languages it's much easier.
I'd place ffmpeg somewhere in the not-bad but not-great part of the continuum. As with English orthography of course the problem is if you change things with the intent to make them better you actually introduce a cost for existing users which may be impossible to sustain.
Kinda like running into an old friend you haven’t seen in decades.
Doesn't apply to well-designed GUI programs. In good GUIs features are discoverable, you don't need prior knowledge or documentation to use software.
Experience speeds things up, e.g. keyboard shortcuts are faster than menus or buttons. That's optional though, one can still use the software without these shortcuts.
#!/bin/bash
size=${1:-"1920x1080"}
offset=${2:-"0,0"}
name=${3:-"video"}
ffmpeg -video_size $size -framerate 25 -f x11grab -i :0.0+$offset -c:v libx264 -crf 0 -preset ultrafast "$name.mkv"
ffmpeg -i "$name.mkv" -movflags faststart -pix_fmt yuv420p "$name.mp4" echo ‘!!’ > ~/bin/name-of-script
where !! is the previous command alias sl='fc -ln -1 | sed "s/^\s*//" >> ~/.saved_commands.txt'
alias slg='< ~/.saved_commands.txt grep'
You can also add a comment to the command before saving. For ex: $ foo_cmd args #this command does this awesome thing
$ slTheir man page even had a like three line long example for basic stuff. ffmpeg is at least a bit more consistent with syntax though only barely!
Their example from the readme:
ffmpeg -i input.mp4 -i overlay.png -filter_complex "[0]trim=start_frame=10:end_frame=20[v0];\
[0]trim=start_frame=30:end_frame=40[v1];[v0][v1]concat=n=2[v2];[1]hflip[v3];\
[v2][v3]overlay=eof_action=repeat[v4];[v4]drawbox=50:50:120:120:red:t=5[v5]"\
-map [v5] output.mp4
Which turns into import ffmpeg
in_file = ffmpeg.input('input.mp4')
overlay_file = ffmpeg.input('overlay.png')
(
ffmpeg
.concat(
in_file.trim(start_frame=10, end_frame=20),
in_file.trim(start_frame=30, end_frame=40),
)
.overlay(overlay_file.hflip())
.drawbox(50, 50, 120, 120, color='red', thickness=5)
.output('out.mp4')
.run()
)
[0] https://github.com/kkroening/ffmpeg-python[1]: https://zfsonlinux.org/manpages/0.8.3/man8/zfs-program.8.htm...
The entire script is executed atomically, with no other administrative operations taking effect concurrently.
Which would be somewhat hard to ensure with an external language. If a fatal error is returned, the channel program may have not executed at all, may have partially executed, or may have fully executed but failed to pass a return value back to userland.
If an external tool is allowed to acquire this kind of exclusive lock, I don't see the difference.That's the whole point, as I understand it. The channel programs are executed in the kernel, in a way that cannot be done by external programs. Some more details here[1].
edit: I also think you should have included the note, which says
Note: ZFS API functions do not generate Fatal Errors when correctly invoked, they return an error code and the channel program continues executing.
So while it's not quite ACID-level, its not as bad as it sounds without that note.
This is how I feel about jq, every time I want to parse some JSON I have to re-read their documentation. Their API is not very intuitive (at least to me).
That's my relation with regex.
1) take your raw uncompressed y4m file and write it out to a directory of PNG files, one png file per frame. this is your static image reference baseline for subjective eyeballs.
2) take your raw uncompressed y4m file and encode it to x265 or whatever codec you're testing, at various different bitrates and encoder settings
3) take your various encoded x265 files and also write those out to PNG files in separate directories
3) pick exactly the same frame number filename from your 'master' PNGs and your encoder-output PNGs, copy them and put them in the same directory so you can quickly flip back and forth between them in a image slideshow application.
If you want to do this with 2160p24,p25 or p30 videos at several minute lengths, be prepared to have 150-160GB of scratch disk per y4m and per PNG-dump-directory.
For example, x264 will “allocate” more data to parts of the image that change less, as those remain on screen longer so artifacts on them are more visible. While fast moving parts can be compressed into a blurry blocky haze as they appear for maybe 20 frames so few people will really notice artifacts.
(Of course the actual implementation is 100x more sophisticated and complex)
You are doing the opposite with PNG’s: focusing on quality of 1 frame instead of on the perceived quality of the frames when viewed as video.
For objective comparisons there is of course VMAF, which is an essential tool for doing automated comparisons of videos vs their uncompressed original. One of the reasons why VMAF was created is that subjective eyeball evaluation of codecs (whether still or moving) will vary between person to person, and is very labor intensive.
But still image PNGs of frames also serve a useful purpose when dealing with lower bitrate videos, for your personal opinion comparison of blockiness, color banding, blobs of color in areas that are mostly the same color, etc.
1. https://www.rtl-sdr.com/rtl-sdr-tutorial-receiving-noaa-weat...
I wish there was a scriptable / command-line interface for HandBrake (which is already based on FFMPEG) [1], where the user just provides the high-level commands: 99% of the time I want to specify high-level commands: extract clip from this timestamp, including SRT subtitles, crop to this geometry, and shift the audio by 3 seconds.
There should be an quick command-line utility to concatenate multiple video files according to exactly the timestamps the user has provided. It's such a common operation.
There's no reason that the tool can't simply do a streaming decode of multiple different file formats and concatenate the video and sub-second precision. If input video resolutions are different, scaling the smaller video to the largest resolution is what the user almost always wants.
I get that FFMPEG is a "plumbing" CLI tool, but a "porcelain" wrapper would be amazing!
Due to how keyframes work, cutting on keyframe boundaries is a lot faster and easier and doesn't require re-encoding in many cases. This is the default for the segment muxer.
Cutting between keyframes is a fair bit more effort, and requires re-encoding, which is why I guess it's not the default.
Even if your two files were encoded with x265 at exactly the same bitrate. It's a much more complicated problem than it appears at first glance, once you really dig into the command line options and encoding parameters of codecs like x264, x265 and vp9.
It's not as simple as concatenating two files together. You can also select down to per-frame precision using kdenlive and loading different x264,x265,vp8,vp9 files into it and cutting/editing them together. You will then need to re-encode the resulting output. kdenlive is ultimately a nice GUI front end on top of this:
Handbrake does have a CLI. [2] I haven't used it and I'm not sure what advantage it might have over ffmpeg. I personally use mkvmerge or ffmpeg for my muxing/cutting and VapourSynth for encoding.
[1] https://trac.ffmpeg.org/wiki/Seeking
[2] https://handbrake.fr/docs/en/latest/cli/cli-options.html
What I don't understand is, how can professional video editing tools trim accurately (and very quickly)? What are they doing differently to ffmpeg?
If do things the "fast way" with ffmpeg, the exported video has random black frames which I think is related to the keyframe issue you mention. If I do things the "slow way" (e.g. accurately) with ffmpeg, it takes a huge amount of time (at least with large 4k videos). But I don't understand how I can drop that same 4k video into Screenflow, trim 1 second out of it and export it in a matter of seconds.
Regardless, I don't think these software are "matching" anything. TMPGEnc for example has settings to choose what quality you want for these re-encoded frames.
There are however, plenty of great software that can but at any frame and only re-encode the frames that are outside the whole GOP. Most of them are commercial though, I haven't find one that is free and good.
----
Also, seeking in FFMPEG in practice, is actually more complicated than the guide [1] you linked. Below is a note I keep for own reference for keyframe-copy. Hope someone will find it useful.
How to keyframe-cut video properly with FFMPEG
FFMPEG supports "input seeking" and "output seeking". The output seeking is very slow (it needs to decode the whole video until the timestamp of your -ss) so you want to avoid it if unnecessary.
However, while -ss (seek start) works fine with input seeking, "-to/-t" (seek ending) is somehow vastly inaccurate in input seeking for FFMPEG. It could be off by a few seconds, or sometimes straight up does not work (for some mepeg-ts files recorded from TV).
The best of the two worlds is to use input seeking for -ss and then output seeking for -to. However, this way, the timestamp will restart from 0 in output seeking. So instead of using -to, you should calculate -t (duration) yourself by subtracting -ss from -to, and use `-t duration` instead. Below is a quick Python script to do so.
https://gist.github.com/fireattack/9a100c5a200154937babd1823...
(You can also try to use -copyts to keep timestamp, but not recommended because it doesn't work if the video file has non-zero start time.)
Could you clarify what you mean with "use output seeking for -to"?
From your Python script it seems that you're just using input seeking and then specifying the duration in seconds with `-t`, which is actually the same as using `-to` when doing input seeking.
Also, input seeking should be inaccurate when doing stream copy, so I'm not sure your script actually works as expected?
(And unless I'm missing something, it seems that all of this is well-explained in the ffmpeg guide linked above.)
Thanks!
In my Python script, I did input seeking for -ss (start point) part, and then output seeking for -t part (end point). As you can see, the -ss part is before -i {inputfile}, and -t is after.
-ss 1:00 -i file -t 5
is NOT the same as
-ss 1:00 -t 5 -i file.
The latter has a bug that happens frequently when I'm trimming MPEG-TS files recorded from HDTV. It literally doesn't stop at the -t/-to timestamp for reason I don't know. And it only happens when stream copy.
Below is a quick showcase: t.ts is the source, and the filenames show how I generate them with FFMPEG (for example, ss_t_i means input seeking -ss first, then -i t.ts, then -t 1:00).
https://i.imgur.com/lLUSzEM.png
As you can see, if I use -t/-to before -i, it doesn't cut the file properly.
>Also, input seeking should be inaccurate when doing stream copy
Yeah, it's not frame accurate, can only cut at keyframes, but enough for my application. By the way the same inaccuracy exists for output seeking if you're doing stream copy.
I think it might be differences between how the the container describes the video and the video itself, and which one is chosen as the truth during operations.
- Inaccurate time handling (https://github.com/yermak/AudioBookConverter/issues/21#issue...)
- Incorrect handling of mp3 chapters https://github.com/sandreas/m4b-tool/issues/71#issuecomment-...
If anyone who is interested in using ffmpeg in a docker container (without the dependencies / compiling stuff), this alias is pretty useful (with relative paths ;-):
alias ffmpeg='docker run --rm -u $(id -u):$(id -g) -v "$PWD:$PWD" -w "$PWD" mwader/static-ffmpeg:4.3.2'But that chapter shouldn't simply be the first n pages covering "what is ffmpeg", "where do I download ffmpeg" and "how do I install ffmpeg".
Show us a practical example what cool things I can do with ffmpeg. Your TOC is enticing, pick one of those things (OPUS or audio from waveform or whatever) and show that! Including the command line and your explanation up to it.
So far I have no idea whether you explain anything, or whether you're just listing "100 cool one-liners".
(The latter may also be fine, but I don't know what I'm buying here)
If the author is here, I suggest looking into RawTherapee. It offers a GUI for authoring and applying raw processing profiles, as well as a CLI tool that can be used for batch application of a given profile to raw images. If you consider raw image processing within the scope of the book, you might as well cover it properly.
[0] Setting aside the insistence of the author to spell “RAW” as an abbreviation, which it isn’t. “PodCast” is another instance of weird, excessive and sometimes inconsistent capitalization style used in the book. Not to say this detracts from the substance, but I would suggest that an editor or a proofreader could help your books be more professional.
Thanks for the reply !
Just like the rest of the internet, there's probably very little that you could do right now, that someone else has not already done or struggled with as well. Most of the time, it's reading about what people have tried that did not work before which is still incredibly useful in and of itself. It sucks while your under pressure of a deadline, but you eventually you just sort of "get it".
The trick is, you have to use it frequently, and not just every now and then when some task comes up. But that's no different than any other tool. Yes, the commands can get harry and scary looking, but so can SQL statements.
Lots of power to be had in mastering them.
roll face on keyboard.
hit enter.
repeat until results look good!
A tangentially related question: I've always found FFMPEG-the-cli-program an amazing piece of software, incredibly powerful, versatile and well-made. I therefore expected the same when I had to interact with its library interface (i.e. libavcodec, libavformat) recently. How disappointing, and very frustrating an experience! The docs felt extremely thin and full of "ah yeah don't use the foo function afterall, it's since been replaced with foo_2 and foo_2really3forreal, but the docs don't mention it", and the API conventions seemed very random and inconsistent. Is it just me?
This is not a complaint; thank you to the people who spend their free time developing a free multimedia suite for me to use! I was just surprised about the perceived quality differences between FFMPEG-the-cli-program and FFMPEG-the-library.
If you are using ffmpeg with hardware-accelerated codecs like H.264, remember to take the free 10x speed boost!
The speed boost isn't free, hardware accelerated encoders often can't compress as well as software encoders.
A video codec isn't a fixed rule book that all the encoders follow to get the same result, it's more akin to CSS: there are a bunch of tools you can use but there's no one way to combine them to get a particular result. If you want to make a chess board with HTML and CSS, just think of how many ways there are to do it, it's similar when it comes to video encoding.
So different encoders have different results but why would a hardware encoder compress worse than a software encoder? The answer is simple: a hardware encoder needs to be implemented in hardware. That comes with a ton of constraints that CPU encoders don't have to deal with, such as physical space on silicon, power, heat, limited memory and so on and so forth. This means that the encoder has to make compromises that a software encoder doesn't.
So how big is the difference? It depends a bit how you measure it but Moscow State University has a ton [0] of data about different encoders and just last year, they evaluated a bunch of hardware encoders [1] and compared them to software. You can see in their results [2] that relative to x264 (software, which happens to be ffmpeg's default), NVENC (NVIDIA's encoder, present on their GPUs) took 21.3% more bytes to produce the same subjective result.
[0]: https://www.compression.ru/video/codec_comparison/index_en.h...
[1]: https://www.compression.ru/video/codec_comparison/hevc_2020/...
There are a myriad of StackOverflow questions, mailing lists, forums... but never a well structured and comprehensive analysis of this topic. And FFmpeg itself, while being a feat of a software project, has such a lackluster and incomplete doc.
I'd like to see a discussion about recovering H.264 and VP8 timestamps, with minimal processing (i.e. not transcoding, if possible, otherwise of course it becomes an easy thing), which covered the whys and whens of using these FFmpeg options:
• -fflags +genpts
• -fflags +igndts
• -vsync
• -copyts
• -use_wallclock_as_timestamps
Does this book cover these options and the topic of timestamp reconstruction?
FFMPEG stands for “Fast-Forward-Moving-Picture-Experts-Group”.
Are we sure FF doesnt stand for "Free Frame" ? I think in the 90s there were a lot of projects with this ff prefix-c copy - It doesn't re-encode the file. What does re-encoding mean?
-async 1 -which is deprecated in favour of aresample. If the correct syntax is aresample=async=1 then would it just cut off the audio timestamp and match it with video timestamp.
It's something you may want done or not whenever a decoding step is happening by necessity anyway. The "possible benefit" would be if the original encoding was partially broken somewhere (which players/decoders can handle usually but they complain to stderr) or otherwise sub-par. Not sure about video but audio files floating about out there from the last 25 years are a wild mess, with "mp3"s really being layer 2 or 1 or not even mpeg at all inside etc.
Typical video-file-related reasons for using the "copy" codec would include just switching between container formats: MP4 vs. Matroska vs. AVI vs. whatever else. As long as the new container supports the types of audio and video you have, just copying them will save both time and quality. Or maybe you want to alter the video but keep the audio untouched, or vice-versa. Or you just want to add a track to something and not disturb the existing ones.
Even doing "copy" for all streams and using the same container can be useful, even though nothing would seem to change in that situation: "remuxing" like this is often all that's necessary to fix a wonky file.
I use the h264_videotoolbox hardware encoder on macOS all the time! But in my experience, CPU based encoding is always better and offers more fine tune control. It's just slow. And there's all sorts of filters and transformations that aren't possible in hardware.
FFmpeg's Nvidia hardware support seems pretty good too. There's an official support doc on using ffmpeg: https://docs.nvidia.com/video-technologies/video-codec-sdk/f...
Sometimes a streamer wants to overlay his stream with a tournament channel, but because of copyright they can't. It'd be nice to have an easy way to set that up, if twitch won't do it on their end the user can do it themselves
I think that a process of clicking a link (meshstreams.io/streamer1,streamer2) would let you download a script that would install FFMPEG and open both streams in VLC
https://github.com/whyboris/Video-Hub-App
Extract screenshots from video, join them together in a horizontal filmstrip (letterboxing as needed) for super-fast preview of each video :)
Handbrake worked too, sometimes easier for things like 4:3 deletterbox or ad removal or b&w restoration on colourised screening.
I still use mpeg_streamclip for some things. I used to use an Italian PhD students Java tool to fix TS flags without explicitly reencoding. ffmpeg can probably do all the things I am doing, if I look harder.
I have no idea what most of the options mean beyond bullshitting, and I rely on the wisdom of others "try adding this option"
It's digital alchemy. Do not offend Yagoth, incanting for Morgoth. The candles are optional but the goat skull is not: i tried removing the -a: option but something broke.
https://github.com/yuanqing/vdx - An Intuitive CLI for processing video, powered by FFmpeg.
I'm looking for some easy and beautiful way to generate audio visualization; the books' TOC doesn't seem show much.
Can someone please tell me a bit more what I may get from that chapter?
Thanks
I got to know about FFMPEG when I wanted to generate gifs using videos.
I’ll definitely buy this book
https://alfg.github.io/ffmpeg-commander/
Hoping to cover more options soon.
I'm always surprised at how often macOS is misprinted, especially in technical publications. I suppose Apple hasn't made it easy to keep up (Mac OS, Mac OS X, OS X[1]), but macOS has been the official name since Sierra[2]. Perhaps it shouldn't, but it does make me wonder about the accuracy of other details a publication might offer.