RethinkDB (YC S09) Raises $1.2 Million For Its Database For Solid-State Drives
techcrunch.com
techcrunch.com
It's not that the metadata is demanding; it's just that there's a lot of it. For each $/month Tarsnap takes in, I have about 200,000 table entries.
TechCrunch misses the point that Rethink is explicitly not doing this. The MySQL engine is below the SQL parsing layer, so as-is MySQL apps should be able to run against it.
- We garbage collect (see Mendel Rosenblum's Ph.D. thesis on log-structured systems)
- Our customers care about cost per IOPS, not cost per GB.
- The hot real-time stuff is usually handled by a different database and/or storage system than the older, less frequently accessed data anyway.
They clean segments to reduce fragmentation and allow for decent-sized extents to write new data into. Since your writes don't need to be long and contiguous, do you really need to empty the live data out of segments?
You DO need to identify which snapshots are stale, and consequently mark certain blocks as free. But I see no need for compaction.
The one we're implementing is really the one everyone wants - snapshot isolation. It can be implemented very efficiently, and is stronger than repeatable read, read committed, and read uncommitted (so you should never want these three). It's not as strong as serializable, but nobody can give you a scalable serializable isolation level.
Snapshot isolation also guarantees consistency, but requires all transactions to be idempotent (so they could be rerun in case of a conflict). It's the best of both worlds, in practice most other databases already behave this way anyway.
[sorry if this seems like an interrogation - it's just interesting stuff you're doing...]
Hash tables have a theoretical advantage over balanced trees, and an SSD would make a naive hash table implementation easier to implement. But if you are smart (like, say, BerkeleyDB), hash tables and balanced trees have almost the same real world performance.
RethinkDB might be better for write-heavy operations, but that's because SSDs are better for random writes.
True, but this is rarely the case for OLTP workloads. What happens when there is a credit card transaction with a user ID 100731, followed by a credit card transaction with a user ID 8762592? Even for range queries, what you're saying is true only if you're walking through the primary index. The second you start walking through the secondary indices, you're back to random read land (my Facebook friends, for example, are extremely unlikely to be stored in the user table sequentially).
SSDs are better for random writes
Random writes are very tricky on SSDs because of the slow erase operation. The FTL controllers are getting much better at this on micro benchmarks, but it's very difficult to measure random write performance profile over different timelines and different disk space utilization scenarios.
Obviously you need to be smart about how you do your locking (no giant lock!) but the mere fact of having locking is not automatically a problem.
(Disks vs. SSDs and transactions vs. "eventual consistency" are orthogonal.)
[1] http://gcc.gnu.org/onlinedocs/gcc-4.1.2/gcc/Atomic-Builtins....
(I'm sure Slava will correct me if I'm wrong here...)
You can get more isolation that this, and you need to to really keep your data consistent, but all DBs except Berkeley seem to have this off by default. So I am not too bothered by this, but I would be interested in seeing how well Rethink handles concurrent OLTP applications that actually care about data integrity. Caring about data integrity is slow, and Rethink might not speed this up all that much. Or it might :)