Each video is node. Figure out where to add edges to existing content, graph theory some shit and tune for similar content maximizing viewing time of short video clips. No ML necessary. Of course, given the magic ML sometimes seems to be, they could certainly be using it but the inputs and outputs are remarkably simple if you get the data structures right. To be clear, I'm not saying it's easy, just that ML's not necessary.