Running FFmpeg on AWS Lambda for 1.9% the Cost of AWS Elastic Transcoder
intoli.com
intoli.com
[1] demo in a talk: https://www.youtube.com/watch?v=O9qqSZAny3I&t=55m15s (the actual run (sans uploading) is at https://www.youtube.com/watch?v=O9qqSZAny3I&t=1h2m58s ); code: https://github.com/StanfordSNR/gg ; some slides (page 24): http://www.serverlesscomputing.org/wosc2/presentations/s2-wo...
* FFmpeg supports http/https as input protocols if compiled with the options enabled. See `ffmpeg -protocols`
* You can parallelize or chunk FFmpeg to enable longer inputs, e.g. I found https://github.com/nergdron/dve/blob/9f1ca516b18f50d1d99d15e...
* Try with larger memory sizes. Larger memory = more CPU for Lambda which may result in shorter transcodes. You might even pay the same amount if the transcodes are CPU bound and finish in roughly linear time wrt CPUs
You definitely misunderstand cloud pricing.
Is there any effort to on-prem lambda stuff yet? I know it's a moving target, but I wouldn't recommend getting into cloud stuff you can't migrate out of.
That will be when the ubiquitous cloud truly arrives -- run on whatever provider in the sky, and as long as they run kubernetes you can run your workloads there.
Edit: OpenFaaS can run on Nomad as well apparently: https://www.hashicorp.com/blog/functions-as-a-service-with-n...
The much harder challenge would be provisioning thousands of on-prem servers to handle the load, but I wouldn't necessarily qualify a dependency on cloud-like autoscaling as lock-in.
I guess the lock-in might come as you unwittingly couple yourself to the intricacies of the rest of AWS's offerings as your app architecture grows more complex.
"Painless relocation of Linux binaries–and all of their dependencies–without containers."
Also, if you found exodus interesting, you may find the following interesting too.
Elastic Transcoder Audio is $0.00450 per minute [1], this article says that with Lambda it cost "$0.00008273 per minute of audio, a full factor of 54 times less than Elastic Transcoder.".
0.00450 / 0.00008273 = 54
This works just fine on AWS Lambda. The `ffmpeg` binary there weighs in at 46 MB. Unless you need something not bundled with that build, it seems like this is sufficient and is easier to set up.
This problem comes up a lot with storage blobs. The bigger they are the worse it is to serialize write/reads.
I'm going to stick with Elastic Transcoder for now though. I like that I have no upgrades to maintain and very little code. I feel like if I did this, it would take me years to recoup the cost even with a 99% savings.
But that is only because I only have a few videos a month. Roughly $1.00 on Elastic Transcoder. If I had thousands or even hundreds of videos this seems like a great and worthwhile project. Especially since this article appears to take a lot of the trial and error and proof of concept out of the mix.
I worked for a large Internet company that had a Netflix like product back in 2007. The transcoders were literally just plugged in underneath people's cubicles. Kept things nice and warm in the winter and I'm sure the costs were pretty low.
Rolling your own approach like this is certainly more complex to build/maintain than using Elastic Transcoder though.
ffmpeg -i "file.mp4" -ss 01:16 -to 02:16 -c:v libx264 -crf 32 "newFile.mp4"
ffmpeg -i "file.mp4" -ss 02:20 -to 02:45 -c:v libx264 -crf 32 "newFile1.mp4"
ffmpeg -i "newFile.mp4" -c copy -bsf:v h264_mp4toannexb -f mpegts temp.ts
ffmpeg -i "newFile1.mp4" -c copy -bsf:v h264_mp4toannexb -f mpegts temp1.ts
timeout /t 5 /nobreak
ffmpeg -i "concat:temp.ts|temp1.ts" -c copy -bsf:a aac_adtstoasc "Finished.mp4"
Looking forward to the next post in the series.
251 webm audio only DASH audio 143k , opus @160k, 78.96MiB
Then if I run youtube-dl -f 251 -g https://www.youtube.com/watch?v=r_fxB6yrDVo then I get this horrible URL: https://r1---sn-hxugvj5nu-cvnl.googlevideo.com/videoplayback...
However if I wget "[horrible_url]" -O audio, it still takes forever to download, so I guess rate-limiting might be the issue. But if download time is the problem, you could have one server that just downloads the data slowly to S3 and then kicks off the lambda job on the completed file.