So the only place you can do dedup is inline, as the data is being written the first time, not after-the-fact.
In addition, this requires you keep a huge indirection table (the DDT, or dedup table) that needs to be read for all writes, so either that has to be kept in memory or on fast storage, or you've just turned every write into one or more random reads, plus writes.
This also means that even if you turn off dedup after turning it on, the performance implications remain until the DDT no longer contains any blocks (e.g. you rewrote all the data after turning dedup off).
There are feature proposals to make the performance of dedup less pathological, but nobody's taken up implementing them so far. (Someone even did a proof of concept implementation of one of them, and it still hasn't been finished and integrated.)