I think you're overstating the meaning of the empirical observation.
The observation as I understand it is that if you take the average over a large set of programs, then the hypothesis holds.
The observation is _not_ that if you take any program, then the hypothesis holds.
I'm hedging that it's the average over a large population of programs. I'm not making a claim about every program, only most programs, or rather - "most programs most of the time". I think that this is important because without a doubt, some programs are not generational at all.
Great example: go navigate from this page to another one. You want a full GC, not an eden GC, at this point. If your primary activity when browsing is navigating around and you're not spawning a new process for each navigation, then you are violating the generational hypothesis big time.
Note that this is somewhat orthogonal to the question of whether you should implement a generational GC. Generational GCs are great in part because they are rarely a regression on non-generational workloads.
Classic example: in the MMTk project they built this goofy "only for testing" non-generational collector that allocates in a bump nursery and then during GC it simultaneously evacuates the nursery and marks the "old" space. This non-generational GC, which has some traits of a generational GC, was slightly faster than their mark-sweep collector. It's often the case that generational GCs are really about a lot more than just exploiting the generational hypothesis.