StreamPot: Run FFmpeg as an API with fluent-FFmpeg compatibility, queues and S3
github.com
github.com
I never understood why it hasn't grown a DSL to specify the input/output/transformation. Could be JSON-based for what I care. Then we could have nice things like easy programmatic generation of these scripts and schema verification (for each particular ffmpeg version).
Also, when I looked at how LosslessCut (an Electron app) does the timeline preview thumbnails, it just calls the ffmpeg CLI app however many times it needs (once per thumbnail). With the number of heavy users of ffmpeg, I'm wondering how nobody has yet started a project to run ffmpeg in-memory as a server and avoid the process startup cost as I'm sure LosslessCut isn't the only one who does this.
There are also a few things that I had to patch on top of ffmpeg (for example it doesn't do well with generating singular HLS packets on demand). Would be nice to have a pluggable architecture that can link multiple things together like that.
Might implement this myself if I end up having the time for that.
It's not quite JSON and I have to look up the syntax every time I use it, but it's well-defined.
The general syntax is:
[input1][input2]... filterchain [output1];
[input3] filterchain [output2];
...
to specify chains of filters that link input streams to output streams.For example, the following:
ffmpeg -i INPUT -vf "
split [main][tmp];
[tmp] crop=iw:ih/2:0:0, vflip [flip];
[main][flip] overlay=0:H/2
" OUTPUT
is doing something like this (python pseudocode): main,tmp = input.split(2)
flip = (tmp
.crop(iw, ih/2, 0, 0)
.vflip())
out = overlay([main, flip], 0, H/2)
Here's a more complicated example (visualization on https://ffmpeg.guide/graph/demo ): ffmpeg -i ./long-video.mp4
-i ./background-music.mp3
-filter_complex "[0:v]trim=duration=30[out0];
[0:a]atrim=duration=15[out1];
[out1][1:a]
amix=inputs=2:duration=longest:
dropout_transition=2:weights=1 1:
normalize=1
[out2]"
-map "[out0]" -map "[out2]"
./shorter-video.mp4
Pseudocode: # input files (-i options)
in0 = read("./long_video.mp4")
in1 = read("./background_music.mp3")
# filter graph
out0 = in0.video.trim(duration=30)
out1 = in0.audio.trim(duration=15)
out2 = amix([out1, in1.audio],
duration="longest", dropout_transition=2,
weights=[1,1], normalize=1)
# output stream mapping
out_stream[0] = out0
out_stream[1] = out2I was watching a streamer trying to use ffmpeg in a streaming configuration and it didn't seem to work very well in that case. However, if someone was belligerent enough, I would guess it would be possible to use libavformat/libavcodec directly to parse out frames from video. For example, the streamer "sphaerophoria" posted a short series where he built a video editor from scratch [1]. I believe the code he wrote could be used as the basis for a custom-built editor to extract thumbnails from videos.
1. https://www.youtube.com/watch?v=l_hD99zpPBY&ab_channel=sphae...
[1] https://github.com/jjcm/nonio-video-cdn/blob/master/route/en...
Ultimately the only element that isn’t polling in a classical machine architecture is the interrupt handler in the top half of the kernel (or system equivalent).
Not necessarily. The server may internally use interprocess signalling or sockets or something to find out when the process is ready instead of busy-waiting. For example, if ffmpeg is run as a child process, the parent process can sleep a thread waiting for the child to terminate.
Just about any sort of busy-wait polling behaviour is a code smell in my book. Stop wasting cycles. Let the thing you're waiting on tell you when its ready. If it can't, fix it until it can. This is all opensource code.
Seriously, unless you know we’re delivering an interrupt or a signal, assume it’s polling all the way down. And I’m not so sure about the signals.
Maybe there's cases in the kernel where polling is unavoidable. But polling should happen as little as possible - for the sake of both latency and efficiency. I don't spend much time thinking about the kernel. In userland, polling is almost never necessary. I'd love it if just about all software that manages data that changes over time was written such that downstream consumers can be notified when that data changes. This should be the case for filesystems, databases, message queues, web servers, device lists over USB and so on.
Bonus if all of Termux can be fixed for secondary Android profiles,
You'd figure anyone with a project with enough throughput for their "most popular" tier (5TB/month, 1000 requests/minute) and a budget of $165/month for video processing, they could wire up something on their own servers with not much effort. Even barring the cost of the service, I imagine it'd be cheaper to run that processing yourself staying within the server than it would be to upload it to another service and then download it again for thousands of videos.
Am I missing something?
Streampot lets you send 300 requests per minute for the cost of the cheapest dedicated server instance on hetzner.