ScyllaDB Closes $16M in Series B Funding
scylladb.com
scylladb.com
I gotta say though, going to each of their websites, and Scylla was very straight forward "here's what we provide" and I understand it pretty quickly. Datastax? I...still don't entirely understand what they offer. It's all loaded with enterprise buzz words.
But last I heard they bought in the titandb guy to build their own enterprise graph database on top of cassandra.
So they got that going for them.
Enterprise companies care about performance, sure. But far, far less than they care about being on a supported platform. DataStax has done the hard yards over the years to prove themselves capable of supporting Cassandra. SycllaDB has no street cred at all.
So SycllaDB might make some inroads in performance critical startups and web companies. But that is likely to be about it.
This funding is part of the story to make sure Scylla can support customers but they have a much better foundation to build upon with this database tech.
And the most important part of a database is confidence. You need to be able to trust that when you have a Production outage IT is able to talk to someone to fix it. DataStax has proven capable of meeting that task. It's very difficult for a second tier player to do the same especially with Cassandra being such a niche product.
Scylla is an upstart... but I would not call them a second tier player. DataStax is already trying and failing [0, 1] to incorporate ideas from Scylla in to their product.
Yes, support and SLAs matter. And this round of funding will go a long way toward helping Scylla build out those parts of their business.
Overall company budget for a technology group is besides the point---individual departments are budget constrained and being able to say I saved 30% last quarter, and whilst handing 200% more operations is something any manager would love.
But showed Scylla to my co worker. And he was like: holly cow! We need to install that!
And I was like: bag, we are at a bank. Forget it :)
Few orgs do use it in production. It's really cool but along with it the cloud needs to change. For instance, we wanted to have a 1-minute-granularity price scheme from the cloud vendors so it will resemble a container (just with hardware-based security..). This any many needs to happen for it.
I would love to move over from Apache Cassandra to Scylla but honestly I'm a bit afraid to do that. I have no doubt that it's much faster but I haven't seen hard numbers about consistency and availability. Apache Cassandra is a much older project with many installations and is battle tested (to a degree) how can I be sure that Scylla will behave as stable as Cassandra in that regard?
They dont have 100% parity yet with all cassandra features so waiting for that to move more work over to scylla, but performance per machine and lack of tuning hassle is very nice.
Absolutely 100% recommend.
Anyways, congrats on the funding guys, certainly not trying to cast shade.
I switched from Cassandra to citus + pipelinedb (and I'm JVM guy). Postgres is such an awesome platform. I'm planning on looking into some logical decoding.
It also has a way to go with just scale-up performance before it even gets to scale-out.
I agree on scale-out but scale-up performance I will say Cassandra is not even close to cost to performance ratio of Postgres w/ extensions for a real time analytics.
I have legitimate 6months experience that I wasted on Driud/Cassandra and could not match Postgres in terms of performance.
I don't want to hear CAP this and that when I can do all sorts of stream processing a priori. Besides real SQL is easy to understand and hack with then many proprietary query languages.
I only tell people so they don't was time like I did.
This doesn't mean anything.
The only thing that matters is what (and how much) data you have and what queries you want to run. If a relational database can do that for you then cassandra/scylla isn't a good choice.
And neither does your practically trolling comment. You could say that about anything.
> The only thing that matters is what (and how much) data you have and what queries you want to run. If a relational database can do that for you then cassandra/scylla isn't a good choice.
Theory aside what matters is what I can get to work... so prior knowledge is a big deal.
All I said is it would be interesting if Postgres had a better story for elasticity and that it might be a good fit for many. I think many are in my camp and a relational db would fit but are all to often pushed towards to NoSQL.
Again I don't want to hear about CAP this and that. You can totally take either system and make it have many of the properties that our touted for each one. People take relational databases and turn them into schemaless or columnar stores all the time with eventually consistency (often with mysql). Particularly with Postgres as it is a platform (it has a powerful extension model).
Regardless Cassandra is touted for analytics. The marketing really pushes it for that. Eventually consistency and columnar store is a good fit for real time analytics. Lots of data points, lots of aggregation, get to play with various queries in realtime, elasticity... etc.
We will probably hit a wall with Postgres.
All and all I think Cassandra is pretty good for what it does. Other than Redis, and maybe RethinkDB I think its my third NoSQL favorite (and yes I know each of those guys has a sweet spot for what they are good at requirement wise).
Cassandra is wide-column (which is just another buzzword for key/value), not columnar (as in storing data in a column-oriented format) so it's actually not great with aggregations and barely supports queries like that. It is good for range scans across data in a single partition and for spreading load around the cluster if your data is also spread evenly into these partitions, but ultimately my point was that in a thread about cassandra/scylla, it doesnt make much sense to bring up a relational db because it's completely different in every way.
If it does work better for you, that's great - and it means is that cassandra/scylla was never a good fit to begin with. The multi-master global replication is a key feature that will likely never be reached by postgres (which is just starting to get scale-up and some logical replication features now) and even mysql only just released the group-replication for multi-master which still only supports the concept of a single total cluster.
There is a general consensus what realtime analytics is ( memsql.com apparently tries to define it). It certainly is less nebulous than "big data".
Our biggest problem was continuous aggregates. Continuous aggregates are tough for databases (particularly for Postgres since it is MVCC). So it isn't the relational model that is the problem but the algorithms needed for consistency that conflict with constant read and write speed.
I did goof by saying Cassandra was column oriented (that is a loaded and confusing term) but people do use it all the time for aggregates (see Druid). Druid by the way is apparently column oriented (going back to my point how you can most data stores into something else).
Saying Cassandra is completely different than Postgres isn't really saying something terribly useful. I bring up Postgres not because it is a relational database but because it has some nice features and extensions that seem to be cost effective (compared to just loading everything in memory ala redis which is not cheap).
Cassandra certainly does try to offer familiar things to old school SQL guys like myself (namely CQL and various options for consistency)... again it isn't completely different.
> If it does work better for you, that's great - and it means is that cassandra/scylla was never a good fit to begin with.
You are also assuming some stuff like that we didn't have to compromise. We will still need something like Cassandra as we do want to collect more data points and we do need a place to effectively warehouse this stuff across regions.
Also plain Postgres is not a good choice for continuous aggregates as I mentioned before (again I'm going to ignore theory of the relational model... the relational model fits for us because we make it fit... not the other way around). It was one of the reasons why we investigated other technologies.
> The multi-master global replication is a key feature that will likely never be reached by postgres
I'm sensing some bias here... never... maybe never for postgres core but certainly someone could build an extension or add on.
I never thought to try mixing the two since pipelinedb has their own commercial product for clustering. You should write a blog post on that if you were able to cluster pipelinedb with citus.
I mix the two by using a message bus (Kafka + RabbitMQ)
I spent a lot of time with Cassandra. It is probably great tech I just don't have petabyte data yet. And I know all the stats I want a prior.
Thank you for recommending pipelinedb, I haven't come across them before.
I have kafka and postgres and in need of real-time analytics, this may be the solution I'm looking for.
I've used citus in the past, whilst it's excellent for data-warehousing. Unless you scale up the servers, then it's not suitable for real-time counts. I found that lacking for my use-case.
If you see this, can you reach out at my name @ gmail dot com, I'd like to chat about issues you have come across.
The solution has been to just buy more nodes (if you don't want long repairs, store less than 1TB of data per node) and faster disks. Read Repair maintenance is probably the only thing I hate about Cassandra - and seeing benchmarks that Scylla does these operations on the order of minutes rather than hours is attractive enough for most people (I don't think most deploys are even coming close to the benchmarked txn/s in real-world workloads, for both databases). Both compaction and repair tend to be CPU intensive (both work by essentially reading a ton of data), so I'd imagine the move to C++ and the core-per-thread design is more efficient.
In short, the operational efficiency is far more attractive even if you aren't pushing a trillion writes/sec.
I've been thinking about testing Scylla for a while, but unfortunately they don't support the features we support, and while our Cassandra deployment is a rather comparatively large cost, there are enough things on my plate right now where trading my current set of evils for other unknown ones isn't very attractive.
See this post by Discord App - https://blog.discordapp.com/how-discord-stores-billions-of-m... - where they are mentioning moving to Scylla from Cassandra for similar reasons. Performance is fine, but repair efficiency is more of the driving factor.
I'd also add that Cassandra advertises itself as a relatively high performance database for distributed workloads. If something like a faster Cassandra doesn't entice you, chances are you'd be better served by something like Postgres anyways.
We are currently working on first finishing materialized view support, which will hopefully be completed in the upcoming months. Secondary indices will be implemented after that and we're hoping to reuse MV infrastructure for that. So I personally expect both features to land into a release later this year.
Please note that @glommer is talking about the classic secondary index implementation in Cassandra, which is very simple but also broken. I don't know the details but we probably did have "half-ready" code for that. We decided against going forward with it because Cassandra had already moved to SASI (which is also much more complex). As I said, we're currently focusing on materialized views, and tacking secondary indices after that.
Btw, I highly recommend subscribing to our user mailing list or engaging on Github for questions and comments about features. You'll get better and up-to-date answers there.
Update: reading @glommer's reply carefully, he explicitly says that he's unsure if we'll move forward with that specific implementation: "_if_ we do implement it, it should land in our main version in a couple of weeks" (emphasis mine).
Secondary indexes will be implemented on top of Materialized Views. Patches for Materialized Views already exist, and are soon to appear in preview releases.
Any word on native JSON types ? Scylla will have another reason to drop Cassandra.
There's an open issue about it on Github:
https://github.com/scylladb/scylla/issues/2058
Please feel free to upvote and comment on the issue to voice your interest in the feature.
Cassandra is actually schemaless but since the shift to CQL from Thrift, it's unlikely that it'll go back to a schemaless model again.
In the meantime, the Keen.IO crew has a nice model for storing lots of arbitrary json if that's something thats needed. It takes some work but a very clever strategy and they've made it work well.
1. Let someone else solve the hard distributed-system problems.
2. Re-implement the local pieces for higher performance.
3. Profit!
Note that I'm not saying anything is wrong here. Reimplementations of existing ideas are a time honored tradition, and often lead to their own innovations. Linux was a reimplementation of UNIX, and seems to have been good for a lot of people. Most web servers and browsers are reimplementations of things that had existed previously. From compilers and databases to filesystems and hypervisors, a lot of software we all rely on today - especially in open source - is a reimplementation of something or other. I'm pointing out an opportunity, not a flaw.
What happens after 1M operations? The nodes catch on fire?