Please let us know if you have any questions or comments about ToroDB. We will be happy to answer them :)
Enjoy!
Please let us know if you have any questions or comments about ToroDB. We will be happy to answer them :)
Enjoy!
Stampede looks nice, congrats!
Quick question about the 100x performance claim and the benchmarks in https://www.8kdata.com/blog/announcing-torodb-stampede-1-0-b...:
- I don't see any specs and/or methodology published for any of the benchmarks. I'd like to see some specs for the servers used (especially RAM and what HDD or storage type was used).
- For the 500GB dataset, I assume that it didn't fit in memory, but the 100GB could be "easily" fit in memory. I'd like to see if that's the case, and how does it compare when the dataset is all in memory (I'm anticipating that Mongodb still sucks big time, but it's nice to see a clear apples to apples comparison).
Anyway, kudos for the awesome work!
All benchmarks have been done on AWS i2.xlarge instances. They have 4 vCPUs, 30GBs of RAM and 800GB local SSDs.
It is true that we are not doing an apples to apples comparison on the 100GBs set. I think is just the opposite! We are helping MongoDB by giving it three machines (and therefore 90GB RAM!). This is specially true when the benchmark is using indexes because in this case the index can be always on RAM.
MongoDB is very fast when it has to retrieve a single document because, by design, it has an amazing spatial locality (the whole document is usually on the same page). But this feature is a weakness on aggregation queries, as they usually only care about a small subset of the document. ToroDB Stampede change that by storing your data on a relational way. Of course, as you said, MongoDB performance is horrible when it has to fetch documents from disk, but even if the documents are in memory, the same effect is expected (on aggregation queries) when data has to be move to the CPUs caches.
Anyway, my point stands. I'd use a single instance in which the dataset fits in memory, just for completeness. Then, you can compare one mongo with the full dataset to one stampede.
As you said, the aggregated data will still make Mongo suffer, but it will be a better comparison. I still like the 3-shard setup, though. It's also a good reference point.
If there's a larg-ish number of optional fields, but each document has only or a few of them, would it create a sparse table with lots of columns? Did you find any problem in these scenarios?
Sure, sparse tables are created. This is not a problem since nulls in PostgreSQL are quite cheap (they require no or a few bytes of storage per record).
Even if there is a high cardinality of optional fields, we have not seen in real cases that the number of columns goes beyond a few hundred. And that's perfectly manageable by PostgreSQL :)
There might be some pathological, degraded use cases. But we have found none of them on real datasets.