We are big fans of Hedra. Do you know if they've publicly commented on their model architecture? As far as we know, our particular choice of an end-to-end diffusion + transformer is novel.
We don't know what Hedra is doing. It could be the approach EMO has taken (https://humanaigc.github.io/emote-portrait-alive/) or VASA (https://www.microsoft.com/en-us/research/project/vasa-1/) or Loopy Avatar (https://loopyavatar.github.io/) or something else.