FFmpeg Drawtext Filter for Overlays, Scrolling Text, Timestamps on Videos
ottverse.com
ottverse.com
ffmpeg -ss 66 -i crab.mp4 -t 30 -crf 27 -preset veryfast -vf "drawtext=fontfile=/usr/share/fonts/truetype/liberation/LiberationSans-Bold.ttf:text='YOUR TEXT HERE':fontcolor=white:fontsize=240:shadowx=5:shadowy=5:box=1:boxcolor=black@0.0:boxborderw=5:x=(w-text_w)/2:y=(h-text_h)/2:enable='between(t,9,30)',scale=iw/3:ih/3" -c:v libx264 -filter:a "volume=0.2" -c:a libmp3lame output.mp4
You just need to get the original crab rave video from youtube: https://www.youtube.com/watch?v=LDU_Txk06tMYou might also need to supply your own font path I'm not sure how universal they are.
(and suffer a few weeks learning it! I don't even remember C++ being so painful)
ffmpeg -ss 66 -i crab.mp4 -t 30 -crf 27 -preset veryfast \
-vf "drawtext=fontfile=~/Library/Fonts/SF-Compact-Display-Bold.otf:text='YOUR TEXT HERE':fontcolor=white:fontsize=68:shadowx=5:shadowy=5:box=1:boxcolor=black@0.0:boxborderw=5:x=(w-text_w)/2:y=(h-text_h)/2:enable='between(t,9,30)'" \
-c:v libx264 -filter:a "volume=0.2" output.mp4
Thank you very much, crazy how powerful but hard is FFmpeg.It is unbelievably powerful though.
Some sort of deprecation or salience decay needs to be added.
Except for typos, syntax..etc, I don't edit someone else's answer; I leave a comment and let the original respondent edit.
And the worse part is everything is turned into command line arguments
It would be easier if you could specify it in something more human readable like (sigh) yaml
It was: `split[a][b];[b]lut3d=${path}[c];[a][c]blend=all_mode=overlay:all_opacity=${opacity}`
You can't use some expresions on the command line like `x=(w-text_w)/2:y=(h-text_h)` but it can be workaround with some scripting using python bindings for example.
Some (simple) text overlay can be acomplished with:
gst-launch filesrc location=sample.mp4 \
! decodebin \
! videoconvert \
! textoverlay text="Hello" font-desc="Sans Bold 150px" y-absolute=0.5 \
! autovideosink
1. http://manpages.org/gst-launchUnrelated but wow, BBC has quite a bit of interesting and relevant open source projects. Simorgh, their react SSR framework, caught my eye “used on some of our biggest websites”. Encouraging for those looking to build out performance react/amp platforms https://github.com/bbc/simorgh
As far as I know in gst-launch you can't access properties from one element (video stream width) in other elements.
ffmpeg of 15 years ago was confusing, but now it's so simple - to compile, to add your own filters, and to use, they've done a great job.
I still get confused by the terminology.
You are looking at multiple hours (>2) for a 2-hour movie from my experience. One way to reduce this is to spin up a DigitalOcean droplet (a $80 per month droplet) that'll give you a lot of processing power, finish your encoding in 2-3 hours,and shut it down. It'll save you a lot of headache and time.
Not sure I answered your question, but, hope that helped!
In principle, would it be possible to only re-encode the macroblocks affected by the subtitles, leaving the rest of the frame bit-for-bit identical?
It's been a long time since i read a video encoding specification, so i don't know if things still work this way.
Can possibly work, but depending on movie content might not save all that much encoding.
Much faster to encode. Can be toggled. But players need to support it.
What's also impressive is it can stream your media at a lower quality, say 4k to 720p to your phone, also in near real time, on a very old laptop that I repurposed into a NAS.
See https://support.plex.tv/articles/200250377-transcoding-media...
You're going to have to reencode it, which is a lossy and CPU-intensive process (unless you manage to get hardware encoding working with your video card, which is possible but very fiddly IME).
> Is it a super cpu intensive process that takes a long time for let's say a typical 2 hour movie?
There are a lot of compute/filesize tradeoffs that you can make. If you don't mind a file that's 2-3x larger than your original then you can do it fairly quickly (say 1/3 of realtime). If you don't mind an effectively uncompressed video (so tens or hundreds of gigabytes) then you can do it as fast as your disks will write. If you want something similar to the original without sacrificing too much quality then it'll probably be slower than realtime.
It can do 300fps+ (for simply transcoding. Burning in subtitles probably will slow it down a little ibt) already on my very dated GPU, likely would be even faster on newer ones.
Edit: Obviously it's more complicated than this. But I think this is one of the reasons behind the need for things like Netflix's VMAF
In any case, they massively lag behind software encoding.
But I guess you can call it "firmware" or "software" dependant too, because they don't really use/support the same version of NVENC to begin with.
And I doubt the result is significantly different, if any, for software encoders like x264 - of course assuming you use the exactly same version and parameter.
The above line will hardcode the subtitles onto the video and stop after five minutes. This should give you an idea of how slow re-encoding is going to be on your machine, and what the quality of the result is going to look like.
What you ideally want is that your rendered video has approximately the same bitrate as the original, and the same visual quality. You can play with the -crf parameter to get the quality the same, and the -preset parameter to get the bitrate the same.
For reference, when I'm rendering 720p video I'm getting about 4x speed on my 10th gen Core i5. 1080p should take about twice the time.
I've played around a bit with using the gpu-accelerated encoders in ffmpeg, but I could never get the same quality or bitrate as the cpu encoder, it seems to me that they're tuned for quickly turning raw capture into something that can realistically be streamed and handled.
At my job we have an API that lets you do this to "burn in" subtitles onto a video.
A demo for this: https://transloadit.com/demos/video-encoding/add-subtitle-to...
See eg: https://mutsinzi.com/add-srt-subtitles-to-quicktime/
TL;DR:
ffmpeg -i yourVideo.mp4 -i yourSubtitle.srt -c:v copy -c:a copy -c:s mov_text -metadata:s:s:0 language=eng yourOutputVideo.mp4
I'm uncertain if there are some metadata you could toggle to hint at default subtitle language and/or default to subtitles on (or off). I generally use vlc, so this isn't much of a problem for me.
To make the first subtitle track default:
-disposition:s:0 default
Or you can make it forced: -disposition:s:0 forced
And to clear the disposition of the second subtitle track (if it was default in the source stream) -disposition:s:1
The problem is that video players are inconsistent in term of how they honour these parameters.https://trac.ffmpeg.org/wiki/Hardware/QuickSync
Even the most low-end processors like the Celeron, if they have embeded GPUs, have this encoding acceleration built in and it makes a huge difference versus using the general purpose CPU portion. The generation of CPU will determine the encoders/decoders available.
I learned about this from the below post, and have been using the functionality where ever I can as it applies many of the popular transcoding sofwares out there, for example handbrake also.
https://forums.serverbuilds.net/t/guide-hardware-transcoding...
My startup is using ffmpeg filters to create video courses so the GUI is aimed towards that use case rather than being flexible: https://blog.modernlearner.org/product-update-split-content-...
The results aren't bad: https://www.youtube.com/watch?v=L-NjLrwTyxs
Great results depend on fine-tuning of the ffmpeg filters and having great quality video to work with.
The ffmpeg filter mini-language is easy to work with when you're targeting a particular use-case.
I'm using it for creating short videos: https://www.youtube.com/watch?v=L-NjLrwTyxs
And I'm scratching the surface since ffmpeg has so many amazing filters: https://ffmpeg.org/ffmpeg-filters.html
-c copyI keep a document with arguments I use for each of my use case.
I've been running basically this, which is a simple conversion to ogv, mp4, and VP9 webm, with a tiny bit of audio cleaning:
ffmpeg -y -nostdin -i $f -threads 8 -codec:v libtheora -qscale:v 4 -codec:a libvorbis -af afftdn -qscale:a 0 $f.ogv -codec:v libvpx-vp9 -qscale:v 4 -codec:a libvorbis -af afftdn -qscale:a 0 $f.webm -codec:v mpeg4 -codec:a aac -af afftdn -qscale:a 0 -qscale:v 4 $f.mp4
Across a huge dataset from a huge number of backgrounds, usually incredibly crap, for about 18 months, now. Without downtime. And with decent quality output, or at least pretty close to equal the input (did I mention it is often incredibly bad? Including near-broken files, and even some files with less than 20,000 pixels, total).I'm not sure exactly what "never worked" means without more detail, but I'd say it's a workhorse.