When i was last using gentoo (which admittidly was a while ago) it was pretty much the bottle neck for most things if you wanted to compile in parallel, largely because intermediate steps are written out to the disk before linking and less so because of reading. That said it is possible to do it all in memory but then you end up memory bound which is also possible to work around by throwing money at your ram (or downloading more off the web).
Could you elaborate on that? I would assume that most compiler steps are linear traversals of graphs and the time comes from lots of small IOs.