Ex-Googlers Building Cloud Software That’s Almost Impossible to Take Down
wired.com
wired.com
EDIT: GitHub has more information. It's a scalable ordered key-value store (think a distributed version of Berkeley db.) Storage is based on RocksDB (a variant of LevelDB) and consensus is achieved using Raft. The database is written in Go. It's meant to "tolerate disk, machine, rack, and even datacenter failures with minimal latency disruption and no manual intervention."
What's not clear at all is where it sits CAP wise. It says it's available and strongly consistent. Which would be CA, which is not an option (especially for something claiming to be failure tolerant.) It must either sacrifice write availability or consistency in the event of a partition. No idea which way it leans.
Either you use a strongly consistent mode that can have poor performance under contention, or a strongly consistent mode that will have good performance under contention but lots of failed transactions. So you get to decide on performance vs availability.
To answer your question, it sacrifices availability, not consistency. It's an MVCC, after all.
If it's based on Raft, then it sacrifices availability if there aren't a quorum of nodes online.
And yeah, you kind of have to sacrifice availability if you want to stay consistent in the face of write skew...
And logical address space is still far in excess of what most disks or arrays can fit, right? 40 bits or so on linux?
EDIT: 47 bits, for 128TB -- http://stackoverflow.com/questions/2159456/whats-the-max-fil...
If Moore's law holds, a single SSD will outgrow the address space in around 7 years. In four years, an array of eight disks would outgrow the address space. This is just for a single server. If you want a linearly-scaling, robust solution for future requirements (like multi-petabyte and exabyte distributed datastores), there's no reason to lock yourself into technology that'll be obsolete in half a decade.
(edit: SanDisk says it may release 8TB SSDs next year, also adding "We see reaching the 4TB mark as really just the beginning and expect to continue doubling the capacity every year or two, far outpacing the growth for traditional HDDs")
As far as the address space and big SSDs thing.. I'd be willing to gamble on linux supporting mmap up to the biggest devices on the market, one way or another. Heck, there's only 16 more bits after that 47 before every FS under VFS has to be rewritten, right?
This makes sense for the current generation of storage sub-systems, though it would be misleading to say using memory map technology will be "obsolete in half a decade". The 48 bit limit is arbitrary. Manufacturers have 56 bit designs on the table right now, and there is nothing stopping them from implementing full 64 bit virtual address support.
The author of LMDB makes pretty bold performance claims and people are too eager to believe them.
You shouldn't propagate those claims unless you've done benchmarking to verify them.
The author of LMDB doesn't really make bold claims, he actually just included LMDB (and the venerable Berkeley DB) in LevelDB's published benchmarks. The benchmarks were developed by the LevelDB team.
it seems like it uses Megastore http://www.cidrdb.org/cidr2011/Papers/CIDR11_Paper32.pdf
If you want to know how spanner works, read the paper. If you want to use it yourself, you'll need to build it on top of your own infrastructure, just like Google did.
Github project: https://github.com/cockroachdb/cockroach/
Servers go down every day, and data centers go down every month. This project solves a real problem.
Maybe you'll come back with a completely-bogus, typical hackernews-ish response, deluding yourself, saying that yes you like it when advertisers target you based on the data they collect about you because of reasons X Y and Z, something something "better for me". Ads targeted at you do not benefit you in any way unless you work for an advertising firm.
And are you seriously arguing that NSA et al do not peek into cloud storage? Have you been living under a rock?
They mean I don't have to pay every website that generates content in order to have them...you know...exist.
Economics is not a hard subject. Stop being a reductivist. Or stop being a sneering jerk (and let's not pretend you aren't trying to pick a fight with your tone, 'kay?). Both not-reductivist and not-sneering-jerk would be nice, though.
You don't pay value. You pay a price, which provides value to the recipient. And the value received by the site being low does not mean the price is, since they're both subjective. Which is the problem with ads: for many, the price paid - the loss in privacy - is much greater than the pennies received by the advertiser. I'm glad you value your loss so low, but you shouldn't assume everyone does.
Besides, privacy is like vaccination - it also needs herd immunity. If everyone is exposed, the few "important" people who really need it will stand out like a sore thumb.
Nobody cares.
In any case, it's not every day people are given the choice. There's probably a handful of sites you can pay to remove ads, and the cost is usually much higher than the value the user would have provided in ad revenue.