The de-interlacing block is where most of the delay comes from. Interlaced video (standard definition, 1080i, etc.) only transmits half the lines per field. This means that the processor has to synthesize the missing pixels in order to have full frames to display. A really cheap de-interlacer can be done with just a few lines of delay. You wouldn't want this as it would introduce bad imaging artifacts. Pretty much all consumer TVs use a method called "motion adaptive de-interlacing" (MADI). There are many implementations of MADI. In general terms you use data from frames before and after the one you want to process in order to detect pixels that have moved. Those that did not move can simply be replicated or averaged into the missing slots. Anything that moved requires different treatment. The key here is that you need to store a couple of frames worth of video before you can start with MADI. That's your two frames of delay (or more).
There might be other delays introduced by other subsystems such as the receiver.