Ideally on the fly, with each box sold to the highest bidder, adsense-style.
I conceived it as video bloggers setting up a dynamic replacement volume in their scenes, and compositing a 3-D scene for the sponsored product into it as the stream goes out. One viewer might see an open Coke bottle. Another might see a half-eaten package of Oreos. Another might see framed photographs of adoptable rescue pets. All the vlogger does is avoid entering the ad-volume and the software does the rest.
Inserting into old movies would require a bit more sophistication, I think. You have to discern the geometry of the scene, model it, then discover the insertable volumes, match the camera movement in software, and finally use a shader to match the object textures to the film grain.
But let's not do this. Please.
Only if you're changing volume. We can start with swapping flat surfaces (say, cereal boxes).
I'm guessing this would dovetail nicely with up-converting 2-D movies to work with VR rigs. If you can reconstruct a scene well enough to fake some depth, it wouldn't be difficult to insert extra models into it.
The dystopic endgame would be to automatically generate the models from the 2-D video and randomly fill some of the volumes that don't interfere with the existing scene with advertising models.
It's a weird thing to watch an film turn out differently to how you rememebered it when you're not told it's a different cut. Feels like low-grade gaslighting.
(Another example is the translation of "天下" - Tianxia - in the film Hero, which is both critical to the film and translated differently in different versions)
These small clips can take several hours to render; my (unwarranted) assumption is that the machine being used likely has the equivalent (minimum) of a 1060 or 1070 (at least, that's what I'd use).
Now - if you have the resources to own or build a multi-GPU machine with scads of RAM and the very best CPU - it will still take hours, but you might be able to get it down to "less than a day's worth".
That's my best guess based on what I have seen so far and my own limited personal experience with deep learning and ML tech (nothing involving face swapping or such - more mundane things thru MOOCs I've taken in the past, using my 750ti SC as the GPU with tensorflow).