It only works if the background is continuous and fairly static - you can't easily synthesize a complex background if you have no reference material. You can of course extrapolate the visible background using other AI that has knowledge of real surfaces; however if the background has familiar content (like a sign) that is not visible in the reference frame(s), it's unlikely to fool anyone no matter how well the face is animated.