Very good points. Seems like this leaves us with "trust benchmarks that probably aren't generalizable to your use-case" or "run your own expensive benchmarks before you have the scale to have the data to simulate the scale".
What's the best strategy here? I've defaulted to MySQL, pg, or sqlite, based on my preference of the moment. I haven't had enough good/bad outcomes to form a stronger opinion, but I'm now realizing I've never really thought this decision through (despite building a few data systems that I'm pretty proud of).