While I find it good that the article explicitly addresses issues with trace-based compilation (usually this is not the case), a completely fair account needs to present the additional memory requirements for using the PyPy tool chain. Quite recently, somebody here has addressed this by mentioning that he does not really care for all the performance speedup he gets, if the memory requirements become outlandish at the same time.
It would also be very informative to know what the differences in automatic memory management techniques are (i.e., what did the previous implementation do?) Personally, I am also interested in interpreter optimization techniques, and it would therefore be interesting to me what--or if at all--the previous VM used for example threaded code or something along these lines.