How Discord Stores Billions of Messages (2017)
blog.discord.com
blog.discord.com
All this happens on the aforementioned MongoDB cluster and just two server nodes. And the two server nodes are really only for redundancy, a single node easily fits the load.
What I want to say is:
-- processing a hundred million simple transactions per day is nothing difficult on modern hardware.
-- modern servers have stupendous potential to process transactions which is 99.99% wasted by "modern" application stacks,
-- if you are willing to spend a little bit of learning effort, it is easily possible to run millions of non trivial transactions per second on a single server,
-- most databases (even as bad as MongoDB is) have a potential to handle much more load than people think they can. You just need to kind of understand how it works and what its strengths are and play into rather than against them.
And if you think we are running Rust on bare metal and some super large servers -- you would be wrong. It is a normal Java reactive application running on OpenJDK on an 8 core server with couple hundred GB of memory. And the last time I needed to look at the profiler was about a year ago.
I have written experimental datastores that can hit in excess of 2 million writes per second on a samsung 980 pro. 1k object size, fully serialized throughput (~2 gigabytes/second, saturates the disk). I still struggle to find problem domains this kind of perf can't deal with.
If you just care about going fast, use 1 computer and batch everything before you try to put it to disk. Doesn't matter what fancy branding is on it. Just need to play by some basic rules.
Primary advantage with 1 computer is that you can much more easily enforce a total global ordering of events (serialization) without resorting to round trip or PTP error bound delays.
The other thing that is helping a lot here compared to Discord: Trading is very neatly organized in trading days and shuts down for hours between each trading day. So you don't have the issue that Discord had where some channels have low message volumes and others have high, leading to having scattered data all over the place. You can naturally partition data by day and you know at query time which data you want to have.
Decentralized Acyclic Graph based networks (e.g. Hashigraph, which are not technically blockchains) can reach effectively infinite TPS but suffer in time to finality.
Solana is a blockchain with zero downtown (and a Turing complete smartchain), mind you-- nota centralized exchange.
This is both wise and stupid at the same time.
It is wise if you mean "be ready for servers to crash at any time by thinking they are going to crash at the worst possible moment".
But it is stupid, because people think they need massive parallel deployments just because servers will be constantly crashing and it is just not true. The cost they pay is in having couple of times more nodes than they really need to have if they got their focus right (making the application efficient first, scalable later)
The reality is, servers do not crash. At least not the kind of hardware I am working on.
I have been responsible for keeping communication with a stock exchange for like 3 years in one of my past jobs and during that time we haven't lost a single packet.
And aside from some massive parallel loads which used tens of thousands of nodes and aside from one time my server room boiled over due to failed AC (and no environmental monitoring) I never had a server crash on me for the past 20 years.
So you can reasonably assume that your servers will be functioning properly (if you bought quality) and it kinda helps a lot at design stage.
In the big data world the "complexity" of the data doesn't really mean much. It's just bytes.
3x12TB
> In the big data world the "complexity" of the data doesn't really mean much.
Oh how wrong you are.
It is much easier to deal with data when the only thing you need to do is to just move it from A to B. Like "find who should see this message, make sure they see it".
It is much different when you have large, rich domain model that runs tens of thousands of business rules on incoming data and each entity can have very different processing depending on its state and the event that came.
I am writing whole applications just to data-mine our processing flow just to be able to understand a little bit of what is happening there.
At that traffic you can't even log anything for each of the transactions. You have to work indirectly through various metrics, etc.
Complexity of data and running business rules on it is not a data store problem though, that's a compute problem. It's highly parallelizable and compute is cheap.
For reference, my team runs transformations on about 1 PB of (uncompressed) data per day with 3 spark clusters, each with 50 nodes. We've got about 70ish PB of (compressed) data queryable. All our challenges come from storage, not compute.
In order to be able to run so much stuff on MongoDB, we almost never run single queries to the database. If I fetch or insert trade data, I probably run a query for 10 thousand trades at the same time.
So what happens is, as data comes from multiple directions it is being batched (for example 1-10 thousand at a time), split into groups that can be processed together in a roughly similar process, and then travels the pipeline as a single batch which is super important as it allows amortizing some of the costs.
Also the processing pipeline has many, many steps in it. A lot of them have buffers inbetween so that steps don't get starved for data.
All this causes latency. I try to keep it subsecond but it is a tradeoff between throughput and latency.
It could have been implemented better, but the implementation would be complex and inflexible. I think having clear, readable, flexible implementation is worth a little bit of tradeoff in latency.
As to storage being source of most woes, I fully agree. In our case it is trying to deal with bloat of data caused by business wanting to add this or that. All this data causes database caches to be less effective, requires more network throughput, more CPU for parsing/serializing, needs to be replicated, etc. So half the effort is constantly trying to figure out why they want to add this or that and is it really necessary or can be avoided somehow.
By batching 1-10 thousands records at a time, your use case is very different from discord, which needs to deliver individual messages as fast as possible.
This takes about 20 seconds. The process opens about 200 connections to the cluster and transfers data at about 2-4GB/s.
Did you consider an explicitly bitemporal database like crux[0]?
I got into programming through the Private Server (gaming) scene. You learn that the more you optimize and refactor your code to be more efficient, the more you can handle on less hardware, including embedded systems. So yeah, it's amazing how much is wasted. I'm kind of holding hope that things like Rust and Go focus on letting you get more out of less hardware.
Rust is slower (usually) because Rust does not revolve around making the most use of the hardware it has.
both are fine choices; nothing wrong with either direction.
Zig performs very well by default because it was designed to be efficient and fast from the start, without compromise. it has memory safety, too, but in a way that few seem to understand, myself included, so it's difficult for me to describe with my rudimentary understanding.
Computers are fast, basically. ACID transactions can be slow (if they write to "the" disk before returning success), but just processing data is alarmingly speedy.
If you break down things into small operations and you aggregate by day, you can always have big numbers. The monitoring system that I wrote for Google Fiber ran on one machine and processed 40 billion log lines per day, with only a few seconds of latency from upload start -> dashboard/alert status updated. (We even wrote to Spanner once-per-upload to store state between uploads, and this didn't even register as an increase in load to them. Multiple hundred thousand globally-consistent transactional writes per minute without breaking a sweat. Good database!)
apenwarr wrote a pretty detailed look into the system here: https://apenwarr.ca/log/20190216 And like him, I miss having it every day.
Not sure how many people would be interested. Reactor has quite steep learning curve but also very little literature on how to use for anything non-trivial.
The aim is not just enable good throughput, but also achieve this without compromising on clarity of implementation. Which is where I think reactive, and specifically ReactiveX/Reactor, shines.
But there is a lot of things that CPU can do even faster than that, because this limitation only relates to actual instruction execution (and even then there are instructions that can process multiple words at a time).
((channel_id, bucket), message_id)
The primary key consists of partition key + clustering columns, so this says that channel_id & bucket are the partition key, and message_id is the one and only clustering column (you can have more).They also cite the most common cassandra mistake, which is not understanding that your partition key has to limit partition size to less than 300MB, and no surprise: They had to craft the "bucket" column as a function of message date-time because that's usually the only way to prevent a partition from eventually growing too large. Anyhow, this is incredibly important if you don't want to suffer a catastrophic failure months/years after you thought everything was good to go.
They didn't mention this part: Oh, I have to include all partition key columns in every query's "where" clause, so... I have to run as many queries as are needed for the time period of data I want to see, and stitch the results together... ugh... Yeah it's a little messy.
Well, here it is. The partitioning in manual upto the SQL level.
Using Cassandra tends to mean pushing costs to your developers instead of spending more money on storage resources, and your devs will almost certainly spend a ton of time fixing downed nodes.
Apple supplied some of the biggest contributors to Cassandra who were optimizing things like how to read data in a partition without fully reading the partition into memory to avoid the terrible GC cost. They put in a ton of engineering effort that probably could have been better spent elsewhere if they’d used a different database.
Also, your partitions should never get that large. If you're designing your tables in such away that the partitions grow unbounded, there's an issue. There are lots of ways to ensure that the cardinality of partitions grows as the dataset grows. And you actually control this behavior by managing the partitioning. It's really easy to grok the distribution of data in on disk if you think about how it's keyed.
You've basically listed a bunch of examples of what happens when you don't use a wide columnar store correctly. If you're constantly fixing downed nodes, you're probably running the cluster on hardware from Goodwill.
This is a pretty good list of what not to do with Cassandra, or any similar database. https://blog.softwaremill.com/7-mistakes-when-using-apache-c...
In that case the overhead of processing and collecting all the parts of the data you need spread across different sstables and then do tombstones can lead to a lot of memory stress.
But let's not pretend that cassandra isn't almost always a bear. The other problem is that cassandra keeps things up (and never gets the credit for it) but that creates a host of edge cases and management headaches (which makes management hate it).
Most competitors abandon AP for CP (HBase and Cockroach and I think FoundationDB) in order to get joins and SQL, but the BFD on cassandra is the AP design.
Scylla did a C++ rewrite to address tail latency due to JVM GC, but after an explosive release cycle, they basically stalled at partial 2.2 compatiblity. Rocksandra isn't in mainline and doesn't appear to be worked on anymore.
I follow the Jepsen tests a lot: they don't seem to have found a magic solution.
I think Cassandra stopped short of some key OSS deliverables, and I think they could simplify the management as well, both with a UI for admin and with some re-jiggering of how some things work on nodes. The devs are simply swamped with stability and features right now.
And Datastax won't help that much, what admin UI cassandra had was abandoned, and I half think the reason they acquired TLP was that TLP was producing/sponsoring useful admin tooling.
I would love to try something new. What appeals to me about cassandra is the fundamentals of the design, and the fair amount of tranparency there is (although there is still some marketing bullcrap that surrounds it like "CQL is like SQL" and other big lies).
So many other NoSQL's are bolt-on capabilities for handling distribution that Jepsen exposes (MongoDB famously) and have sooo much bullcrap in their claims. All the NoSQLs are desperate for market share, so they all lie about CAP and the edge cases.
Purely distributed databases are VERY HARD and are open to exaggeration, handwaving, and false demonstrations by the salesmen, but those people won't be around when you need a database like this to shine: when the shit hits the fan, you lose an entire datacenter, or similar things.
Could you explain this more? Because Scylla has had pretty steady major release updates over the past few year. See the timeline of updates here:
https://www.scylladb.com/2021/08/24/apache-cassandra-4-0-vs-...
We have long since passed C* 3.11 compatibility. In fact, if anything, Scylla, while maintaining a tremendous amount of Cassandra compatibility, now offers better implementations of features present in Cassandra (Materialized Views, Secondary Indexes, Change Data Capture, Lightweight Transactions), plus unique features of its own — incremental compaction and workload prioritization.
But if there's something in particular you're thinking of, I'm open to hear more on how you see it.
One of the most impressive softwares that I've seen and use after years of using ventrilo/mumble/teamspeak.
It's not bad per se but there's plenty of crap in there.
The shortcuts situation is absolutely dreadful for one, I don't understand how gamers can cope with it:
* there are all of 5 actions you can bind to custom shortcuts
* discord defines dozens of built-in shortcuts you can not rebind or disable, if any of those conflicts with something you need you better hope the OS has a way of overriding it
Large chatrooms as well, the moderation tools seem rather limited, maybe it's better for administrators but as a user all you can do is block someone and you still have to see that they're posting comments. It' incredibly frustrating.
Then the linking and jumping to old message works half the time, maybe, search is absolute dogshit, and I've rarely seen a less reliable @-autocompletion, half the time I have to find old messages of the person I'm trying to ping before discord remembers they exist and lets me actually @ them.
And I don’t think support actually exists. You just post into the black hole that are tte support forum thing.
Find the UI very confusing. Perhaps I'm just old; but damn, my intuition in using discord's interface constantly lets me down.
Even years later it's still the only platform I know of that combines text chat rooms, voice chat rooms, and video streaming into one place, all accessible from your 'server' as they call it.
It also has clients for many platforms, including a web client, all of which look and function the same.
Any alternative out there does one of those things decently well, but either completely lacks or is utterly awful at the other things.
and unlike skype (and probably teams too) supports PUSH 2 TALK which gaming oriented voice chats had close to 2 decades ago and is even more useful now, during WFH.
The big issue they have is building up a large enough network effect. I really can't see the discord communities I am a part of moving over there any time soon. Also, they were recently acquired by roblox, and nobody knows for sure what the new ownership will end up doing to the platform.
All the features afterwards are mostly just them throwing stuff at the wall and seeing what sticks.
The ability to easily create servers, invite users to your server, and then make that server your homebase with its own channels and emojis, is pretty novel and perfectly fit into the gaming community which is basically a loosely connected graph of friend groups.
Mumble is also self hosted, and nobody wants to host anything anymore when there's a free alternative that's good enough and hosted by someone else.
Skype is just universally terrible.
EDIT: since I'm part of that tiny minority exception that proves the rule: s/nobody wants/the vast majority do not want/
* Where the server is hosted / quality of server
* Poor client UI
The client UI issue is how easy it is to work-around bad audio from other users. It's possible to do, the UI just completely sucks.User interface and end user fulfillment just aren't great generally for OSS. I think it would take a commons improvement project with either government grants (infrastructure) paying for results AND/OR a university spearheading the development project.
1) They send magic links. Pretty easy.
2) They make all known workspaces you've logged into before discoverable and allow for a one-click "add to desktop Slack" option, which makes dealing with the whole "different users" issue. And to the extent that I use different emails for different workspaces, Slack accommodates that and allows me to do so within the same desktop instance, so not really sure what the concern is there.
I despise discord quite a lot btw :)
I do have to open the official client whenever I do voice calls though, because there's currently an issue that can cause incoming audio to sound terrible. But for text chat, it's great.
Discord is literally the only x86 application that is still installed on my MacBook Pro M1.
It's not the simplest tool if all you want to do is PM a friend or two.
I really prefer IRC.
Yes, they do.
https://github.com/Bios-Marcel/cordless:
> Hey, so I know this is somewhat of a bummer, but I got banned because of ToS violation today. This seemed to be connected to creating a new PM channel via the /users/@me endpoint. As that's basically a confirmation for what we've believed would never be enforced, I decided to not work on the cordless project anymore. I'll be taking down cordless in package managers in hope that no new users will install it anymore without knowing the risks. I believe that if you manage to build it yourself, you've probably read the README and are aware of the risks. I'll keep the repository up, but might archive it at some point. And yes, you'll still be able to use existing binaries for as long as discord doesn't introduce any more breaking changes. However, be aware that the risk of getting a ban will only get higher with time!
https://github.com/atlx/discord-term:
> Disclaimer: So-called "self-bots" are against Discord's Terms of Service and therefore discouraged. I am not responsible for any loss or restriction whatsoever caused by using self-bots or this software. That being said, there's no one stopping you from risking using an account, so go head!
It's again TOS and people have copped bans for using alternate clients.
"All 3rd party apps or client modifiers are against our ToS, and the use of them can result in your account being disabled. I don't recommend using them."
I remember my friends and I kept bickering who would pay for this month's bill for the vent/mumble servers. That kept on for years until I had enough and hosted my own in a droplet in digital ocean. None of my friends knew how to do that since they're not very technical.
Discord you just had to click a couple buttons and its free.
Discord's a pretty good product, and they've got the engineers and money to get better, but the only reason they won is because of timing. Same for Slack; there were identical products to Slack that tried for decades to gain traction, but they weren't free, because that business model didn't exist at the time.
The ux of Slack is essentially screen+irc implemented in JS with emotes. It enabled technical and non-technical people to use the same tool. The key to success is not technical, it's that they tailored the product to a specific group that would then lock itself in.
I didn't understand Discord's success, but comments here point that gamers couldn't find free group-voice apps at a critical time. Here again, they tailored the product to a group that would then voluntarily lock itself in.
Later, they sell the companies with valuations based on the captured user bases.
Another big thing that the current crop of winners has going for it is that cloud hosting allows applications to launch literally for free and scale quite a bit without paying much of anything in infrastructure costs. That also wasn't an option 10-20 years ago.
How long can they keep paying for that bandwidth and message data storage while keeping the thing essentially free?
I wonder about this a lot. I wonder if they have some big 'whales' that help sustain their business OR they're just selling all of our data (is that enough to make money at discords scale??).
> was party to a class action suit with allegations including computer fraud, invasion of privacy, breach of contract, bad faith and seven other statutory violations. According to a news report "OpenFeint's business plan included accessing and disclosing personal information without authorization to mobile-device application developers, advertising networks and web-analytic vendors that market mobile applications" [1]
Of course that doesn't mean anything about the current model of Discord, but good to be aware of.
Discord declined to share how many Nitro subscribers it has, but the Wall Street Journal reported that Discord generated $130 million in revenue last year, up from $45 million in 2019. In the same time period, its monthly user base doubled.
0: https://qz.com/2034087/chat-app-discord-is-shedding-its-game....
It's a bit of an odd model for paying for businesses, but works well in the gaming world where multiple people can essentially help pay for a server (if you want the extra toys)
Discord makes you the product. It's gratis in exchange for letting them spy on you. If you don't know why that's bad...
That seems easy to you. That would be easy for me too and most likely 90% of the people on HackerNews.
But the average person doesn't have a "random Linux box" in their house. Most people don't even know what Linux is. Most people would be overwhelmed just looking for the terminal emulator on their computer, before they even typed a command into it.
Most people don't want to manage an always-on linux box for a voice server. Most people don't want to manage port-forwarding on their firewall/router. Most people don't have static IPs at their house and wouldn't know how to setup dynamic dns to solve the problem. Most people don't even know what DNS is.
MOST PEOPLE just want a program they can launch when they want to talk to their friends. That is why Discord has been successful.
I'm not saying that's good. I am just saying that its the way the world is.
It's no wonder so many projects and FOSS tools fail to gain large userbases when it seems that most developers seem to be living on another planet entirely.
its not a technical challenge, its a moral challenge. it means doing what is good for the users even if they don't really know it
if you have public IP or use stuff like hamachi (at least that's how we did it decade ago)
https://www.pcgamer.com/how-private-is-your-private-discord-...
But it's important to remember that Discord is not that. Discord holds all your data, in luxurious detail, with no option to delete. They go as far as ignoring GDPR when people ask for their messages to be deleted. "Deleting" your account will not even anonymize your ID, it unsets your avatar, renames you, kicks you from all guilds and disables logging in. That's it. And if they ban you there is no place to move on to.
These days? Well, most server providers have some sort of basic flood mitigations in place now, and even more advanced protection has become affordable.
Hmm
I didn't meant your server being DDoSd, but you being DDoS (but probably that's what you meant with Skype P2P example?)
Skype (at the time, no idea now) was a very shoddily written piece of software. It was trivial to query the IP of any online user, even if they were not on your contact list or appearing offline.
You had to use a VPN or carefully conceal your Skype ID, I did work with a somewhat popular live streamer back then (so a VPN wasn't feasible), and their ID was a very random string that was not to be shared under any circumstances.
It really has helped the social factor of moving nearly everyone in the office to remote working. Every department that has adopted the "virtual office" Discord setup loves it over Slack and basically never uses Slack anymore. It's way less awkward to call people, it's easier to not incidentally disturb them when they're busy, during breaks/lunch you can go to the "breakroom" and hang out and chat with everyone else. And it was all very easy to setup and with the Discord server template stuff we can even clone it for each department with very minimal work (renaming channels to that departments' people).
Slack does not support syntax highlighting of code blocks.
Discord uses proper markdown and supports syntax highlighting.
These are two things that make me think Discord is better specifically for engineers, aside from it just being generally way better.
It does, but only if you make your code block into its own post as a "text snippet." (I assume this is because Slack's internal markup doesn't allow regions to have parameterized metadata, but there is parameterized metadata at the chat-post-event level.)
You also get other benefits of doing this, e.g. being able to collapse the snippet, download it, etc. Code pasted into Slack should really always be pasted as a snippet. I just wish it auto-detected you were trying to do that and offered to make a snippet.
Not even remotely.
Discord supports:
* fenced code blocks (but not indentation)
* quoting, a single level (nested quotes don't work, properly replying to other comments is painful)
* inline decorations (italics, bold, underline, strikethrough, code)
* inline spoilers (an extension)
* disabling autolinking (an other extension)
It doesn't support: headings, paragraphs, lists (ordered or not), labelled links, tables, images, footnotes, images (you can only use the image upload feature which puts a single image below a comment).
It also has a limit to 2000 char (4000 with nitro), which can be rather low when posting code snippets.
No bulleted lists... that's disappointing.
Maybe worth upvoting:
https://support.discord.com/hc/en-us/community/posts/3600400...
Discord has great tools around moderation and membership tiers; it's designed for users you don't trust.
Slack is much more for a community where everyone knows each other (or at least trusts each other a bit, like you'd trust a coworker).
Also paid discord is 100x cheaper than paid slack, for non-corporate entities. You can get top tier discord for like $100/m while slack price goes up with each user. Not to mention that discord allows users to easily assist in upgrading your server while slack doesn’t have that functionality at all.
Discord School Hubs page: https://support.discord.com/hc/en-us/articles/4406046651927-...
https://www.reddit.com/r/discordapp/comments/p37s7s/so_disco...
Once you're in an enterprise space, Slack's features become actually useful.
One of the best hammers I own is a screwdriver.
To be fair, Mumble is FOSS, and Ventrilo and Teamspeak have literally not iterated since 2005. Discord is pretty mediocre software (remember when they accidentally allowed iframe XSS RCE attacks? A very amateurish mistake), but the incumbents were an absolute dumpster fire.
If you had a mic that had issues in any way (buzzing, volume, balance), "The Wizard" and "AGC" were supposed to fix it for you. Do not fret little one, for you do not need nor want to manually fiddle with settings, The Wizard will make everything right [1]!
The pivotal feature that was the reason so many people I know stopped using it is the ability to change the volume of an individual person [2]. It has been a requested feature since the beginning of time, yet it took until 2016 to implement in dev branch and didn't actually make it into a release version until 2020! Too little, too late.
[1] https://web.archive.org/web/20200223143654/https://wiki.mumb...
True, but to be fair: the next iteration of Teamspeak will be based on the Matrix protocol, which is quite a big iteration. See https://news.ycombinator.com/item?id=25743874 .
https://www.scylladb.com/press-release/discord-chooses-scyll...
https://www.scylladb.com/2019/03/20/discord-on-the-joy-of-op...
The product we built using Cassandra was widely known as our buggiest and least maintainable, and it died a merciful death after several years of being inflicted on customers.
We didn't have a good handle on the exact perf implications of different values of read/write replication. Writing product code to handle a range of eventual consistency scenarios is challenging. The memory consumption and duration of compactions and column/node repair jobs is hard to model and accommodate. It's hard to tell what the cluster is doing at any given moment. Our experience with support plans from Datastax was also pretty dismal.
Maybe the situation has changed since 2016. In my experience with several employers since then, it seems like every enterprise architect fell in love with Cassandra around 2014-2015 and then had a long, painful, protracted breakup.
The good news: C4.0 is a far better performing database than C3.11. The new GCs definitely get rid of the long tail latency nightmares:
https://www.scylladb.com/2021/08/19/cassandra-4-0-vs-cassand...
However, we also compared it to Scylla's latest release, and though C4 is better*, you can still find other CQL-compatible databases that outperform it. Especially around compactions and topology changes:
https://www.scylladb.com/2021/08/24/apache-cassandra-4-0-vs-...
Just published these numbers today.
I was following Scylla since the beginning (only because I like their mascot), and it's actually sort of interesting to see what's going on with the company. I've spent the past few years designing things where transactional systems backed by Cassandra. This is the first time I've been able to use Scylla on someone else's dime, though. The unpleasantly big company I'm at right now is looking to replace a bunch of infrastructure with ScyllaDB (Couchbase, Cassandra, Elasticsearch, DynamoDB). It's catching on for sure, but it still doesn't return any results when I search Dice. It looks like Discord is hiring, though...
In this 2018 benchmark, we were able to calculate that a sustained, provisioned of only 160k write ops / 80k read ops for DynamoDB would cost >$500k per year:
https://www.scylladb.com/2018/12/13/scylla-vs-amazon-dynamod...
That was a few years ago. These days, according to our most current pricing you could do DynamoDB provisioned, 1 year reserved for $38,658/month, which is "only" $463,896 annually (pop up the "Details" button and choose "vs. DynamoDB"):
https://www.scylladb.com/pricing/?writes=160000&reads=80000&...
The same workload on Scylla Cloud would be only $7,442/month, or $89,304 annually.
If you wanted, say, 1m ops — 500k write / 500k read ops — on DynamoDB, that'll run you $131,078/month, or $1,572,936 per year.
https://www.scylladb.com/pricing/?writes=500000&reads=500000...
The same workload on Scylla Cloud would run $29,768 reserved/month, or $357,216 per annum — 77% cheaper.
Of course, all of this is just pure list price. Depending on volume you might be able to negotiate better pricing. However, you'd need a really steep discount for DynamoDB just to get back to Scylla Cloud's list price.
Let me know if you spot any math errors or omissions on my part.
I think 2012-2014 was peak marketing from DataStax. There would be some new major feature with every new blog post, and it would mostly never work as expected. Between 2017 and now, things have settled down.
If you can't magically put out production fires, on huge high-throughput systems, potentially in the dead of night, we are unlikely to pay you $300-400K.
From what I recall from using it a few years ago, it's pretty damn fast, very low latency. HBase had speedy p50s as well but tended to get quite slow at p99 due to GC.
> we knew we were not going to use MongoDB sharding because it is complicated to use and not known for stability
But then goes on to describe using Cassandra and overcoming sharding and stability issues. I.e., changing the key, changing TTL knobs, adding anti-entropy sweepers, and considering switching to a different cassandra impl entirely.
Are these issues significantly harder to solve in MongoDB than Cassandra?
Cassandra and Scylla also use hinted handoffs so if a node is unavailable temporarily (up to a few hours) you can store "hints" for it when it comes back online. Handy for short admin windows.
Increasing top-end write throughput or replication in Cassandra is just adding more nodes, where in Mongo its not just adding nodes, its adding replica sets (which consist of 3 or more nodes). So there's a few more layers of complexity to that story. You need more replica sets to increase write throughput and need more nodes in replica sets to increase replication.
Im hand waving some details here, but I've worked with both platforms can definitely understand the choice at least from a pure infra lens.
The article mentions hot partitions becomming a problem with max partition size, but they're also a problem with scalability. Say, if your writing a very high throughput of logs into the table (contrived example), then your bottlenecked by the rate at which you can write to one partition.
Adding the bucket id (say, the current day or hour), is a common solution, and solves the max partition size issue, but not the scalability issue of hot partitions.
Does what it says on the tin for the primary key.
That said, hotspots are 100% the reason why Cockroach encourages UUID primary keys. The disadvantage to UUID is you want sequential data, you then need a secondary index which you'll have to bucket anyway.
- No out of the box horizontal sharding, according to the post they had 4TB (compressed) data in the cluster in 2017. Looking at their growth I think it is safe to assume that today they would have >50TB which can't be done on a single node. You could use Citus but this is not exactly vanilla Postgres anymore. For such a simple data model wasting time implementing your own sharding solution and (more importantly) shard migration makes no sense.
- Discord is storing text data, in Postgres this will be stored in TOAST tables which has some drawbacks.
- Their workload is mostly inserts, almost no updates. Vacuum only operates on complete tables so you would wast I/O and CPU processing data which you don't even touch. You can partition tables but it's a manual process and you have to make compromises. In 2017, Postgres partitioning still had many performance drawbacks.
- No out of the box redundancy.
- Once your data doesn't fit in memory, Postgres performance becomes unpredictable.
Personally I would have chosen ElasticSearch for this project.
My understanding was PG only uses TOAST when the data is too large to fit in the row, and since PG compresses data before inserting wouldn't user messages be fine?
Testing with Postgres, a 2000 char random sequence doesn't result in TOASTing, but a 4000 random sequence does get TOASTed
And for kicks, 4000 chars that aren't random compress well enough that they don't end up in TOAST.
Given that they said their requirements were "linear scalability, automatic failover, low maintenance, predictable performance", I don't think I'd go that route.
I'm not saying they shouldn't do that though - especially given regulations like GDPR. Designing systems for deletion is important! But it's also really hard, especially if you didn't design for it from the start.
There's also no way the tiny fraction of users who want to delete their data would make up a significant enough proportion of the messages that it would impact their scaling strategy.
---
It was a joy to use, we created channels left right and center and knew everyone needed would be in them thanks to the centralized "role-based" permission system. (We would create project-specific channels and an accompanying role, or client-specific roles for the few high-throughput clients that had lots of small projects)
At the time it did not have threading, which was one of the biggest pain points on the text-chat front.
---
The voice chat is very good, and having dedicated voice channels means you can emulate meeting rooms or desks and have people join as desired/needed. You could be working and idle at "kroltan's desk" voice channel, but even if you weren't, joining one is trivial (a single click, can be done independently by many people) compared to Slack (find the call button somewhere different each time because they redesign the UI every week, then wait for your peer to join the call).
Screen sharing is 720p on the free plan, so for meetings, it was hard to read documents, requiring zooming and whatnot. At the time there was also no setting to optimize for framerate or definition, so even 720p felt closer to 480p. Nowadays you can lower the framerate and also select the desired optimization, so you can ask Discord to optimize the stream for quality which is much better for documents, even in 720p.
---
The client is also much more responsive than other Electron-based chat programs, especially with big workspaces with close to a hundred channels (yes, for a 10-person team, we sure type a lot), search is basically instant and has very useful filters, mentioning roles is great and the notification settings are fine-grained enough to please everyone.
Edit: It seems they have moved to Scylla
Cockroach is very, very good for a distributed SQL database. But it's still performance-limited in its very nature.
More here on the difference between NoSQL/NewSQL performance, using Scylla (a CQL-workalike) as a point of comparison:
https://www.scylladb.com/2021/01/21/cockroachdb-vs-scylla-be...
v1.0 was released on May 10, 2017 [1], so I doubt it was even on their mind when they started working on the project.
I'd imagine a hash like SHA256 would be tricky because if that image was compressed an additional time at all throughout it's internet journey, then we'd get a different resulting hash, but maybe there is an effective way to fingerprint images. I have a utility on my machine (czkawka maybe?) that does really good image de-duplication with what seemed like a common algorithm (based on a quick look at the source).
No idea though, just spit balling.
A little history lesson is in order:
https://www.scylladb.com/2019/02/01/meshify-and-scylla-an-in...
ETA: Going back to the original thread, the whole question of encryption seems to be dodged and that usually means the answer isn't the one people are looking for: https://news.ycombinator.com/item?id=13440921
More on the latter here:
https://docs.scylladb.com/operating-scylla/security/encrypti...
/disclaimer/ I used to work at Scylla.
NoSQL scalable stores like Cassandra basically only work well if you have a very strong model of the queries that you will need to make.
In this case, that's exactly what they had: they knew what their read/write patterns looked like and they knew that they would be growing at hundreds of millions of rows per month, so easy horizontal scalability was a hard requirement.
The biggest weakness of classical relational databases like PostgreSQL come when you have a super high volumes of inserts (as opposed to updates) which will continue to grow your database over time, and you need to keep all of that data accessible for real-time queries.
They might have been able to achieve something like this using a PostgreSQL extension such as Citus, but it really does look like what they are doing fits Cassandra's sweet spot.
[1] I've used RethinkDB, Postgres, MongoDB, MySQL, Cassandra, CockroachDB, TimescaleDB, SSDB, and others
https://datastax-oss.atlassian.net/browse/PYTHON-891
With all the issues I encountered using in prod, it gave the impression of an overly complicated key/value store.
that's a prime target for acquisition by big-tech / big-data companies
1. Hinted Handoffs - if a node has a transient failure, the other nodes store up messages, like your buddy might take notes in class if you had to go to the bathroom. They'd pass you those notes when you got back. "Here's what you missed." When the node comes back online it processes all new operations and works through its backlog of hinted handoffs to get caught up. Because of the backlog it creates, hinted handoffs are only stacked up for a few hours. If the node never comes back up, or comes back after that window...
2. Repairs - in an eventually-consistent database you might miss an update or two over time. Or maybe you're a replacement node that has to fill in for a failed node. The replacement will get streamed data from the other replicas to get it started, or you might restore sstables from a backup, but then you should run a repair job to make sure all your replicas are properly in sync.
(That's my understanding. Let me know if that sounds correct from the hands-on experts.)
`timestamp = snowflake_id >> 22`
Thanks :)
* id is composed of: * time - 41 bits (millisecond precision w/ a custom epoch gives us 69 years) * configured machine id - 10 bits - gives us up to 1024 machines * sequence number - 12 bits - rolls over every 4096 per machine (with protection to avoid rollover in the same ms)
thanks :)
How Discord Stores Billions of Messages Using Cassandra - https://news.ycombinator.com/item?id=13439725 - Jan 2017 (155 comments)
I do use discord for a few groups, too bad they will not allow 3rd party clients because a discord terminal app similar to irssi would be awesome
In any of those domains, if you are trying to solve your problem with MongoDB you are in for a world of hurt.
That's generally when people start looking at other options. Whether an in-memory system for pure speed, or a horizontally scalable system for raw size or throughput.
Can't wait for OpenFeint 2.0 to have its scandal lawsuit too.
If y'all are, I'd like to get in contact! My email is: anthony75025[at]gmail[dot]com, and my resume can be found at https://www.anthonyjiang.com/pages/resume.html
Do you use Cassandra for all your access patterns or do you use something else (elastic search or something) also?
I'm just curious as in my professional career I recently switched to platform engineering from full stack mobile/web software engineer.
Thanks!!!
[0] Senior Site Reliability Engineer: https://discord.com/jobs/4004051002
https://db-engines.com/en/ranking
MongoDB is ranked #5 on the list at present; Cassandra comes in at #11. (And Scylla, which they moved to most of their workload from Cassandra, is currently #88.)
DB-engines also have specific rankings for what are known as 'NoSQL wide column stores' — which is what Cassandra and Scylla are classed as:
https://db-engines.com/en/ranking/wide+column+store
Note that MongoDB is a different class of NoSQL entirely. It is a "document store" — MongoDB is the most popular document store.
https://db-engines.com/en/ranking/document+store
But what this means is that even though both MongoDB, Cassandra and Scylla are all "NoSQL" making this move for Discord required significant data modeling and migration.
(Note that the difference between Cassandra and Scylla is far narrower. Both use the same data model and Cassandra Query Language (CQL).
Hope that helps give you some orientation in the NoSQL database field.
https://www.scylladb.com/press-release/discord-chooses-scyll...
- a company amasses a large trove of sensitive information
- it is exposed to adversaries or political enemies
- the information is used against the people
should we shame ycombinator for storing the messages, accounts and comments on hacker news then?
I am still unable to delete my account here even though the CCPA and the GDPR exists. But here we are.
Nobody is making the argument you should be forced to delete your messages.
In any normal world, messages that are not used would be deleted as a matter of privacy. They're kept, because they can be kept, and they can be monetized. That monetization has zero benefit to the user, it's just an artifact of our odd way of doing business where we continue to externalize a lot of things. I think over the next 10 years we might see a regulator shift , which also means costs more directly exposed, meaning Discord may cost $1 month, i.e. the externalization 'costed in' like carbon tax on fuels.