Yelp rebuilds corrupted Cassandra cluster using its data streaming architecture
infoq.com
infoq.com
https://engineeringblog.yelp.com/2023/01/rebuilding-a-cassan...
> The investigation around the exception revealed that at-least one of the SSTable (Sorted String Table) rows was unordered, which caused the compaction operation to fail. SSTables are immutable files that are always sorted by the primary key
This was just... a bug in Cassandra? Is there anyone that can shed light on this? There seem to be plenty of people using Cassandra at scale -- is constantly repairing it normal practice?
https://www.youtube.com/watch?v=0QsLU9na2uE
But yes, you win some (mainly resilience, availability and disaster avoidance, possibly tunable consistency will help you) you lose some.
I remember reading about Discord switching from Cassandra to ScyllaDB I think.
Arguably Cassandra does sound like a weird choice, but we don't know the specifics of their setup. There's a lot of solutions presented on HN where SQLite and a Java application would have been a better choice and you can say for sure without knowing all the details, I feel like this is past that point.
I don't have the experience you do in this situation, but my first reaction to this was definitely "don't use Cassandra". But I also never really understood the use-case where Cassandra shines as a solution either (seems like only companies with a lot of data really seem to get wins from it?)
The specs called for a Kubernetes cluster and a Cassandra database cluster for production and the same for testing. Everything about this project was presented as being highly complex. Digging into it with the client and the developers we reduced it to: One Docker container running in Azure websites and CosmosDB on the backend, which could basically run on the free tier.
The whole thing hold less than 3GB of data and had maybe a few hundred requests per day. The client just loved the idea that they where special, dealing with massive amount of data and required scalability to keep costs under control. The developers more or less just ran with it and wanted to do Kubernetes and Cassandra sounded interesting and now they had a client that would pay for it. Technically I suppose that both Kubernetes and Cassandra where reasonable choices, had they had 1000x the load, but given their market they where never going to grow beyond 10x on this particular solution.
We didn't get the contract. Our bid was insanely low (not really worth the cost of bringing in a new contractor), delivered something different that asked for (fair enough).
I see a small hosted k8s setup with a hosted sql server as a fine app for many things.
In this case you had a potential zero management environment, truly Cloud as it was meant to be, vs. managing Kubernetes, plus a database. You could go with managed Kubernetes (AKS) and a managed SQLServer, but why take on that cost?
Edit: Even AKS isn't truly zero management, you need to do at least some of the work for the upgrades, so instantly more management, something you need to do, something that adds to the operational cost.
I have personally never build an app that was a single docker image running. Usually I am at a min of 2. One server going down shouldn't take down prod once you even have 1 paying customer.
There were half a million restaurants back then or thereabout, and they wanted to store ratings and comments.
That is not something you'd be able to put on one single database in 2004.
Why didn't they simplify afterward I can't imagine, but I can see how a business directory at that age looked at the numbers and went yeah not in a database.
Back in 2004 I was running larger MySQL databases than yelp has now:
https://www.enterpriseappstoday.com/stats/yelp-statistics.ht...
> That is not something you'd be able to put on one single database in 2004.
Sorry, but it was, plenty of companies had much larger databases back then running plain old master slave replication.
But even if it wasn't: just don't run it all on one db? Nothing says you need to be able to join on the restaurant table and the comments table. Put them on different servers. That's all you're doing with Casandra anyway.
I used to work at Yelp, and at least at that time, a big use case of Cassandra was basically for what were essentially materialized views created from log data. At that point in time, MySQL or logs were the "sources of truth", but there were enough transactions going on that it made sense to have things like Cassandra around too for some of the other use cases.
Also, whether you need a highly available database for a company like yelp is an implementation detail.
Oh no, the database is down, how will Jonny decide where to go for lunch now?!
> Also, whether you need a highly available database for a company like yelp is an implementation detail.
You surely know better than them what they need, that's impressive.
Whether they really needed cassandra or whether it was just using fancy complicated tech for the sake of it, I couldn't say.
1. In one case, developers hate DBAs and MongoDB is outside the scope of DBAs.
2. A developer starts using CouchDB to include in the curriculum.
Example of fairly standard Cassandra bug (don't know if present on latest release, certainly was a year or two ago): When you add a new node to the cluster, it 'bootstraps', where it copies ~1/n the data from other nodes. When you are done bootstrapping, it's copied a bunch of data from other nodes, but the other nodes still contain that data. You then run 'cleanups' on the other nodes to remove the (now stale and unusable) data so as to get your disk space back.
If you accidentally run a cleanup on the new node as it is being bootstrapped, it will succeed, you will delete all the data that's been copied over so far, and Cassandra will _not_ terminate the bootstrap. Everything will be green, but your new node will suddenly be using 0 disk space. When the bootstrap finishes, possibly days later, your cluster will be immediately corrupted due to violated replication guarantees - but only on data that hasn't been read or written over that period, because if it was written it'll be re-replicated, and if it was read Cassandra will silently repair at this time. Repairs resolve the issue, but if you've made this mistake due to scripting, if you get unlucky it's possible to just delete all replicas of some data between repairs.
Example of other Cassandra bug (again, might be outdated): Cassandra nodes identify themselves on startups with IPs, and the owned token ranges are not persisted, they're streamed from other nodes in the cluster. If you've deployed your Cassandra in K8s and you reboot multiple nodes in one go and they swap IPs upon reboot, you may now find yourself in a split brain situation in which nodes magically forget they own certain data ranges and think they own each others data (or maybe it's that the nodes still think they own the right ranges but other nodes think they own the wrong ranges). Wasn't close enough to fully debug that one.
It's a mess. Would seek to avoid problem spaces where I might need to use it again, though if by chance ended up in a space where it made sense, probably wouldn't avoid the tech.
> Example of fairly standard Cassandra bug (don't know if present on latest release, certainly was a year or two ago): When you add a new node to the cluster, it 'bootstraps', where it copies ~1/n the data from other nodes. When you are done bootstrapping, it's copied a bunch of data from other nodes, but the other nodes still contain that data. You then run 'cleanups' on the other nodes to remove the (now stale and unusable) data so as to get your disk space back.
Interesting, seems like there is a bunch of little knowledge like this needed to run a service properly... Managed Cassandra has more added value to provide I guess.
> If you accidentally run a cleanup on the new node as it is being bootstrapped, it will succeed, you will delete all the data that's been copied over so far, and Cassandra will _not_ terminate the bootstrap. Everything will be green, but your new node will suddenly be using 0 disk space. When the bootstrap finishes, possibly days later, your cluster will be immediately corrupted due to violated replication guarantees - but only on data that hasn't been read or written over that period, because if it was written it'll be re-replicated, and if it was read Cassandra will silently repair at this time. Repairs resolve the issue, but if you've made this mistake due to scripting, if you get unlucky it's possible to just delete all replicas of some data between repairs.
This seems... really bad -- I don't think I have the skill to run a Cassandra cluster (and not enough use cases to run it as a hobby to find these edges)...
This sounds like the space for a consultancy to make a tidy killing though.
To any Yelp data engineers who might happen to read - good work, and it's a good testament to the platform you provide.
https://www.yahoo.com/lifestyle/can-you-trust-yelp-crowd-fun...
https://thehustle.co/botto-bistro-1-star-yelp/
https://nypost.com/2014/10/13/restaurant-fights-yelps-allege...
https://www.wired.com/2010/02/yelp-sued-for-alleged-extortio...