LosslessCut: The Swiss army knife of lossless video/audio editing
github.com
github.com
(Unfortunately, VideoReDo was proprietary, produced by an indie developer, and that indie developer recently passed away.)
For those who don't really get what "lossless" video editing is all about, consider that most video editing software always involves these stages: importing video/audio "clips", storing the video/audio timelines in some sort of "app native" format (e.g. Pitivi, Premiere, Final Cut), and then exporting the completed video in one or more output formats (e.g. MP4, MOV), re-encoding the entire thing from scratch.
This means that if your only goal is to, say, cut 30-90 seconds out of a 1-hour video, you're still going to have to re-encode the entire 1-hour video. That also means if your re-encoding system isn't a match for however the original video was encoded, you'll make some changes you didn't intend via the re-encoding (e.g. video or audio quality changes).
With this "lossless" style of editor, however, it'll figure out a way to "snip out" the 30-90 seconds (you can think about this being "at the byte level") without re-encoding the entire thing.
White it has a few odd UI characteristice, it's free, it's actively developed, and I like it a lot. Not throwing shade at LosslessCut; it's fine. But I do prefer Avidemux.
Most video formats are a stream of delta-encoded frames that must be decoded and displayed in-order. If you want to go back even a single frame, you have to restart decoding from the beginning, and if one frame is damaged, you lose picture. So encoders have to regularly restart the video stream by inserting keyframes. And it's valid to insert keyframes basically anywhere into a foreign video stream, at least for the same codec and resolution.
Proper NLEs don't do this AFAIK - most of the cuts you do with them aren't approximatable with simple GOP[0] cutting tricks, and most people aren't editing video clips that are encoded to their final delivery format. Hell, at the professional level a lot of people record in ProRes, which is deliberately not delta-encoded[1], because delta encoding is actually really bad for nonlinear video editing.
[0] Group of Pictures - a keyframe plus all the delta-encoded frames up to the next keyframe.
[1] And has enormous size because of it
---
Basically, lossy compression codecs store groups of frames, or groups of samples. You just can't "cut" at any frame or sample that you want. If you do, you have to re-encode the whole thing.
If you want to understand why, the "delta encoded" thing means that only the difference between frames (or samples) is stored. It usually takes less bits to store the difference between frames, or samples, than it does to store each frame / sample separately.
The reason why "groups" exist is so that you can jump around in a video / audio file. The beginning of each group stores a complete frame / sample. (This also helps if the file is corrupted.)
As long as you cut a compressed audio / video file at the groups / jumps, you can cut it losslessly.
---
(Note: Techniques like companding can use less bits when each frame / sample are stored separately. Some people don't even call this lossless. The technique is basically a digital version of the old "Dolby B" button on tape decks from the 1980s and 1990s.)
If you had a video with keyframes 'd'K and delta frames 'd' like
0 1 |<-->| 2
KdddddddddddKdddddddddddKddddddddddd
and you wanted to cut out the indicated frames, could you create a new keyframe from the keyframe marked '1' + the two following delta frames, and then another new keyframe from K{1}dddddddd? Something like 0 1 |<-->| 2
KdddddddddddKdddddddddddKddddddddddd
Kdd = keyframe "L"
Kdddddddd = keyframe "M"
KdddddddddddKdLMddKddddddddddd
Or I guess if you're creating new frames that weren't in the original, could you just rewrite all the deltas in between K1 and K2 such that those 5 frames aren't there? 0 1 |<-->| 2
KdddddddddddKdddddddddddKddddddddddd
KdddddddddddKddeeeKddddddddddd
(where 'e' marks recomputed deltas)See https://news.ycombinator.com/item?id=40844633. Some container formats support metadata so you can cut without reencoding.
This allows cutting at completely arbitrary places (I'm 99% sure that this can also be at a higher time resolution than the original data), with the cost that the original (unplayed) data remains in the file.
However, some players may have incomplete support for them. For example I believe FFmpeg only supports edit lists used at the beginning of the video, which is commonly used to align video and audio properly.
Reminds me of the early MP3 days, when people would download 96kbps files, then reencode to 320mbps to 'improve the quality'.
I was using a Palm PDA, so I found a player that supported ogg. That allowed me to push to 48kps (or maybe even 32kps) with "acceptable" quality.
Though using the same logic described, just imagine the quality of 320mbps!
320 milli bits per second? 0.32bps. That's not that much.
This means that if your only goal is to, say, cut 30-90 seconds out of a 1-hour video, you're still going to have to re-encode the entire 1-hour video
If you’re cutting 1 minute then you are encoding one minute. You might need to decode more than one minute of the source video but it’s rare that it would be the full 1 hour to decodeThink of removing a commercial break from a broadcast clip or removing a "waiting for speaker" prelude from a 1 hour recorded lecture/speech.
It's no wonder that it uses FFMpeg to do the heavy-lifting, but I think it's worthwhile for the community to understand how this process ultimately works.
In a nutshell, every single modern video format you know about - mp4, mov, avi, ts, etc - is ultimately the extension of the container that could contain multiple video and audio tracks. The tracks are called Elementary Streams (ES) and they are separately encoded using appropriate codecs such as H264/AVC, H265/HEVC, AAC, etc. Then during the process called "muxing" they are put together in a container and each sample/frame is timestamped, so the ESes can be in sync.
Now, since the ES is encoded, you don't get frame-level accuracy when seeking for example, because the ES is compressed and the only fully decodable frame is an I-Frame. Then every subsequent frame (P, or B) is decoded based on the information from the IFrame. This sequence of IPPBPPB... is called GOP (Group of Pictures).
The cool part is that you could glean the type of the frame, even though it's encoded by looking into NAL units (Network Abstraction Layer), which have specific headers that identify each frame type or picture slice. For example for H264 IFrame the frame-type byte is like 0x07, while the header is 0x000001.
Putting all this together, you could look into the ES bitstream and detect GOP boundaries without decoding the stream. The challenge here is of course that you can't just cut in the middle of the GOP, but the solution for that is to either be ok with some <1sec accuracy, or just decode the entire GOP which is usually 30 frames and insert an IFrame (fully decoded frame can be turned into an IFrame) in the resulting output. That way all you do is literally super fast bit manipulation and copy from one container into another. That's why this is such an efficient process if all you care about is cutting the original video into segments.
LosslessCut: lossless video/audio editing - https://news.ycombinator.com/item?id=33969490 - Dec 2022 (153 comments)
Lossless-cut: The swiss army knife of lossless video/audio editing - https://news.ycombinator.com/item?id=24883030 - Oct 2020 (10 comments)
LosslessCut – Save space by quickly and losslessly trimming video files - https://news.ycombinator.com/item?id=22026412 - Jan 2020 (1 comment)
Show HN: LosslessCut – Cross-platform GUI tool for fast, lossless video cutting - https://news.ycombinator.com/item?id=12885585 - Nov 2016 (33 comments)
C=0;echo -n "-filter_complex '";
while read f;do
A=($f);echo -n "[0:v]trim=start=${A[0]}:end=${A[1]},setpts=PTS-STARTPTS[${C}v];[0:a]atrim=start=${A[0]}:end=${A[1]},asetpts=PTS-STARTPTS[${C}a];";C=$((C+1));
done <$1
for i in `seq 0 $((C-1))`;do echo -n "[${i}v][${i}a]";done
echo -n "concat=n=$((C)):v=1:a=1[outv][outa]'"
echo ' -map "[outv]" -map "[outa]"'It's basically just a GUI for ffmpeg.
I then use Permute if I need to recompress, or Davinci Resolve if I need to add effects or lossy edits.
What does it give you that something like Handbrake doesn’t (or even just directly calling ffmpeg)? Seems like a paid app. I don’t mind paying for software if it’s really better than other existing free options.
I also find that lots of tools artists are using generate multi-gig files that compress with the default settings in handbreak to 1/10th the size, and at least personally, I can't tell the difference at a glance. I sometimes think the size is being used as a proxy for "quality" so the artist and some of their patrons think larger file size = better
However while Handbrake supports this use case perfectly I don't think there's a way to re-encode the audio but keeping the video. I guess that's when this submission comes handy.
I use LosslessCut to trim and then Permute to speed up or crop. I almost always speed up my screen recording by 20% to make them feel higher energy and more engaging.
ffmpeg -ss 5:40 -i $(yt-dlp -f 313 -g https://www.youtube.com/watch?v=3SykjY3GdMs) -acodec copy -vcodec copy -t 20 waterdrop_mountainlake.mkvEDIT. Some decoders can be reinitialized by adding parameters in band. But this is technically illegal in some formats (like mp4) but it may "just work" in some playback environments
https://courses.cs.washington.edu/courses/csep590a/07au/lect...
Basically it will attempt a precise cut and if it doesn't coincide with a keyframe it will reencode both ends on each side of the cut and the nearest keyframes.
It is generally very small segments to reencode so it is pretty quick and the quality of these is nearly lossless too so the cut feels invisible.
The only issue is if your segments are really short (under a second) it will generally fail... but chances are keyframe cuts in this case don't work well either.
Not sure that I have a use for it, myself (although I could have used it back in my wageslave days).
I know exactly how difficult this type of thing is, so hats off to making it easy.
Best of luck with it!
I don't see anything like LosslessCut's "smart cut" feature, which re-encodes just the part between the cut point and next keyframe. Unfortunately, smart cut doesn't appear to work reliably. https://github.com/mifi/lossless-cut/issues/126
So, not in reply to you, but the parent. It then seems entirely accurate to say they are not comparable.
Ffmpeg as workhorse is common. But what handbrake and losslesscut are in terms of (mostly) UI, they solve different problems, of which there isn't much to compare.