In your 2015 article, you criticized that the ArangoDB team restarted the instances after each test run. In this 2018 edition, they don't do this anymore.
"I think I've been in the top 5% of my age cohort all my life in understanding the power of incentives, and all my life I've underestimated it. Never a year passes that I don't get some surprise that pushes my limit a little farther." -- Charlie Munger
Unfortunately, there is no independent organization that believes in this and defines scenarios that are tested in different environments.
What's the best strategy here? I've defaulted to MySQL, pg, or sqlite, based on my preference of the moment. I haven't had enough good/bad outcomes to form a stronger opinion, but I'm now realizing I've never really thought this decision through (despite building a few data systems that I'm pretty proud of).