The Andy Pavlo CMU Crew has a great paper summarizing, comparing, and benchmarking different MVCC implementation techniques -- including garbage collection -- here: https://db.cs.cmu.edu/papers/2017/p781-wu.pdf
I wonder why they chose such an unrepresentative dataset size, ~84MB.
Yo. This is me. The point of this paper was to evaluate the core MVCC algorithms under different contention scenarios. So you strip out all the internal features of the system that can influence performance that are unrelated to the experiments. This ensures that you can make it a true apples-to-apples comparison. And since everything is in memory, you don't need a large data set.
IIRC, we ran ran the same experiments with 100m tuples instead of 10m and it did not change the results.