Why is the problem inherently hard? To see why, let's imagine that the All-Knowing Fairy Godmother of Software Estimation descends from the heavens and lends to us her magic estimating function F that, applied to any software project S, will tell us exactly how much time and money our preferred software team will consume to implement S. Don't worry about how F works, just believe that it does. (Okay, okay. Let's just say that F peers into a parallel universe in which our team has already implemented a software system line-for-line identical to S, and it just observes how much time and money were actually consumed in that universe. Anyway...)
Now that we have the magic F, our problem is easy, right? Nope. The real problem is that we don't know what S is. All we actually know is that our S, whatever it ends up being, must satisfy some set of fuzzy constraints C (commonly called "software requirements").
In truth, there are a countless number of possible values for S. In other words, S is a random variable, a mapping that ascribes a probability to each of the possible software systems that we could build to satisfy C. Therefore, our best estimate, even with the perfect estimating function F, is itself a random variable. In other words, it's not an estimate but a distribution of estimates.
See how fun this is getting?
But, wait, it gets funner. That's because people in the business world don't want a distribution of possibilities. They want a budget, a date on the calendar. So we must squeeze point estimates out of the true distribution. (And, remember, we don't even know what the true distribution is!)
That's where the second reason kicks in. Let's just think about the distribution of possible budgets for our distribution of possible values of S. Because S can range from "the simplest thing that could possibly satisfy C" to "the most insanely complex thing that a frighteningly gifted salesperson for an enterprise consulting firm could get our CIO to throw money at," that distribution is going to be w-i-d-e. From X to 10X wide.
So if we're the folks tasked with coming up with those point estimates, we could plausibly estimate X on the low end, or 10X on the high end, or anything in between. Guess which estimates are going to get us the most push-back from the higher-ups?
Now, we're swell guys and all, and we want to do a good job with our estimates. No question. But, still... We're human. If we think that giving a higher-end estimate is going to make us unpopular with the higher-ups, maybe we'll estimate a little lower. And when we get push-back on that estimate, maybe we'll "take another look at the numbers" to see if we "missed any opportunities."
And that's the old one-two. We start with a hard, fuzzy problem and then add to it the pressure to deliver only feel-good solutions. The result: estimates that are often way low.