FoundationDB's Lesson: A Fast Key-Value Store Is Not Enough
voltdb.com
voltdb.com
Efficient expressiveness follows from the relationships directly preserved in the organization. Generally from least-to-most expressive you have:
- cardinality preserving (hash tables, most KV stores)
- order preserving (LSM, btree, skiplist, space-filling curves*)
- space preserving (space decomposition e.g. quad-tree, a zoo of exotics)
Competent implementation is progressively more complex, nuanced, and sophisticated as you go down the list, so most implementations reflect the comfort level of implementor. As you move down this stack, you can express things efficiently that will be very inefficient to express with a less expressive class of organization.
SQL was designed for databases built on order-preserving structures. While you can implement it on a KV store, it will never be as efficient as a database organized in the more expressive organization that SQL assumes.
KV stores are popular because they are relatively simple to design and implement, not because they are expressive. It is an architectural impedance mismatch to add query functionality that tacitly assumes a more expressive organization. It will never perform well against a database actually organized for the expressiveness of the query layer.
Any database can only do a few things really well. It is inherent in the tradeoffs. You can add mediocre support for a laundry list of other "good enough" capabilities but you never want to market and position your product around those mediocre capabilities.
In this case though FoundationDB's KV store was order preserving. The API supported (& efficiently implemented) ranged reads, not just individual get's.
Implementing layered architectures always looks good 'on paper' but the details often throw up performance issues that are hard to deal with without punching some holes in the abstractions.
In this case it seems that relatively few companies had a need for a massively scalable ordered KV store, and the whole SQL layer was an attempt to bridge the product to a wider audience. It would be fascinating to hear more of the story but I suspect that will never escape now.
I'm curious though - databases typically store data in B-Trees which are blocks of equal size which works great for block storage. So isn't "block storage" essentially a key value store, where the key is the block number and the value is the block itself? That I think is the proper way of using a key/value store as the database back-end. (And that's what I did in my SQLite/Redis experiment, BTW)
It would help if you could push down filter predicates to run locally inside Redis, but at that point you're already more than a key-value store. I wonder if you could do this using Lua?
But the core idea of pushing predicates to edges seems reasonable. At one point, I built this sql engine that coordinated queries and pushed down queries to the edges. It assumed that each edge store implemented an iterator over all its values, with optional filtering and sorting (if not implemented on the edge store, then the engine/client would filter/sort). It works great, but I haven't yet published it for other reasons.
Common operations like networking, transaction management and even index walks (except the key comparisons) are already compiled to native code, so you don't need to go all in. You just optimize the stuff that needs it.
Yes, it's critical about the "layers" model and about the SQL implementation in particular. Most of what I write isn't this critical, but I thought there was actually an interesting point here so I wrote it down. Take it for whatever it's worth.
Seems like the guy knows what he's talking about as well, b/c surprise, he's working on a DB.
I'm not sure why Apple would prefer FoundationDB to Cassandra for this usecase.
The only thing I can think of is Apple has an internal team that is building a Cassandra replacement and they need more qualified engineers with experience.
Now I don't think FDB (the product) is the answer, at least not in the short term. There are more problems scaling it to Apple's use case than there are working around Cassandra's lack of ACID.
So I'm convinced the value of FDB is the experience in the engineers' brains. Apple need brains to run Cassandra, but also to figure out if Cassandra is the right long term path. Build, buy, adapt? It takes veterans to make the right call.
My guess is that FoundationDB is replacing their Teradata installation. Better to buy the company and invest heavily in it then let it not met it's full potential as a small startup.
A more measured, intelligent consideration of pros/cons is needed.
Every April 1st :)
https://gigaom.com/2013/03/27/why-apple-ebay-and-walmart-hav...
Teradata is VERY expensive and my guess is reaching also scalability limits.
John makes some good technical arguments about why implementing transactional SQL on top of a distributed KV store (even a transactional one) is hard.
The points about metadata performance and consistency were actually new ideas to me. I already had beef with moving the data around to SQL nodes as that is an obvious waste of capacity.
But the importance of metadata in processing SQL queries in a distributed database never occurred to me even though I've lost a decent chunk of my life implementing that consistency. It's one of those things you take for granted if it's how you have always done it.
My primary issue is that the people best able to rebut John's well-argued points are no longer able to do so.
My secondary issue is that John could have easily made the same logical arguments _before_ FoundationDB was acquired, but chose not to do so for whatever reason. This would have led to a much more enlightening debate than the one we're able to have now.
It's actually precisely because FDB isn't a perceived competitor anymore that I can write this. As a vendor, I actually have much less of an agenda now. It's not like I'm worried about FDB stealing a customer. If I had posted this months ago, it would have been less credible and it would have felt tacky.
The other point is that I think there's actually something to say here. I'm not just dumping on a dead product, I'm trying to show there's a lesson to learn here about trying to bolt SQL onto things (a trend). To contrast, if NuoDB disappeared tomorrow, I could write a solid post on why they could never technically achieve what their marketing said they could, but there's no lesson there. "Don't make bad engineering choices" is too generic.
If I remember correctly they were doing/planing a lot of optimizations like:
1. delaying requests a little on purpose to take more advantage of batch requests
2. fancy techniques to improve join locality (based on Akibans previous work).
These two go hand-to-hand.
So in a good enough network (aka not any public cloud) it'd probably work reasonably well up to a point.
It seems like FDB-SQL was closer to 1x, with a much better replication story, but with huge limitations on the kinds of things you can do. (https://foundationdb.com/layers/sql/documentation/Concepts/k...)
So maybe you could push it to 2x or 3x with a few years of work, but other new systems with more SQL support and more customer traction are doing 10x and up today. It's a tough sell.
In which case it is a bit silly and you get the worst of both world.