Anyway, it's completely legal. You just have to scrub the data pretty thoroughly before you sell it.
The dataset spanning all of them is likely to be in the tens or hundreds of TB range, if not PB.
Now, I don't understand DP well enough and information theory/signal processing still seems a bit like "dragons be here" to me. But, I want to take a stab at trying to reason why he said that.
For example, take randomized response (the only DP technique I understand). That is vulnerable to a longitudinal attack: a person can query repeatedly to wash out the randomness. If you think about it, isn't it the almost the inverse of a repetition code (error correction)? There, you're trying to use redundancy (repetition) to remove noise.
If your signal processing professor was already taking that into account then I would be curious to know how that attack would work.