Of course, the major concern in publishing that data was precisely this - people's income is super sensitive. If there was some program at some university that only graduated one student in a particular year, you couldn't report the average wage outcomes for that program, because you'd essentially be putting that student's salary online for everybody to see. Instead of adding randomness to the data, as described here, we'd simply hide any information that was based on 3 or less individuals, or which could be entered into a mathematical formula to enable you to DERIVE information on 3 or less individuals.
For example, suppose you reported the mean salary of everybody who graduated with an art degree from community colleges in 2010, then reported the mean salary for everybody who graduated with an art degree from each individual community college in 2010. Suppose further that one of those individual community colleges only had a single art degree graduate and you hid the data. Under those conditions, somebody could still do some simple algebra to calculate the hidden value. So you'd have to hide the value from a second community college to prevent somebody from working out the unknown value at the first.
As you reported more and more data across a higher and higher number of dimensions, the problem grew more and more complex. It actually ended up being really cool to reason about, and we developed both a greedy heuristic and a binary-integer-programming approach to solving to near or true optimality, respectively.