We tried our best to incorporate these results in the cases where they reflect fundamental shortcomings not bugs (of which Aphyr found quite a few). But you are right, we could include a list of popular examples, where description and experimental findings diverge. Which cases did you have in mind? MongoDB being CP?