After a quick scan, it seems to me that CUPED is supposed to be run at a per-unit level (ie. per user normalization), to reduce variance, which seems to be a bit different than the fallacy I describe here (computing lifts from the T and C group's "before" and "after" overall mean separately and subtracting the lifts, which is what I observed and triggered me to write this post).
(I wrote the post.)