We have to call free when the program shuts down because the program is available through a library interface, and so it can be called as a linked library in addition to being used as a standalone program. If it wasn't for that, we wouldn't even need free.
The program is structured so that based on the parameters it is given, we know exactly what needs to be allocated: how many video frames need to be malloced and so forth. This is all done in the initialization functions and after the program has started, it never needs to grab any memory ever again. This is not a case of "we're allocating the max we need and distributing it later"--it's a case of since we know how the program works, we know what exactly what it will need to do and how many frames it will do it to simultaneously.
There are only four things in the entire program that are malloced:
1. Video frame objects (big structs with tons of pointers to data that goes along with a video frame). A new_frame function exists to malloc and return a video frame. This makes it easy to check mallocs. We know exactly how many frames will be needed on startup. These frames make up the vast majority of memory usage of the program.
2. One data struct for each thread. Each thread has its own local struct that it passes around containing data currently being worked on.
3. A scratch buffer, per thread. This is a small malloced buffer used for a few calculations that require variable-size buffers, usually depending on video frame width (i.e. something known at initialization, but not known at compile-time).
4. Bitstream output buffers for each thread. In the extremely rare case that a bitstream exceeds the initial allocated value, this is realloced on demand.
Every single one of these can be free'd trivially at the end of runtime.