The visibility you will get after the capex when there’s a truly disastrous outage will be interesting.
The visibility you will get after the capex when there’s a truly disastrous outage will be interesting.
The biggest hardware price point is that you need insane amounts of RAM so that you can mmap the bloom hash for the mapping from word_id to document_ids.
And at the qps they've described, it's not a throughput issue either. So I'm pretty confident in saying that this is a case of premature optimization.
And at some point the increase in parallelization of scans dominates mmap speed, unless you're redundantly sharding your mmaped hash table across multiple machines. And there are cases where network bandwidth is the bottleneck before disk bandwidth, though probably not this case. But yeah basically, the answer is something like "if this is the optimal choice, it probably didnt matter that much".
So if that alternative database takes on average 0.1ms per index read, then it's starting out roughly 65000x slower.
"than a DB lookup (in the right style, with a reverse-index)"
Unless, of course, you're managing petabytes of data ;)
"at the qps they've described, it's not a throughput issue either"
It's mostly a cost thing. If a single request takes 2x the time, that's also a 2x on the hosting bill.
"parallelization of scans dominates mmap speed"
Yes, eventually that might happen. Roughly when you have 100000 servers. But before that your 10gbit/s node-to-node link will saturate. Oops.
What kind of servers are you running? What's your max QPS?
The fact is with your mmap impl. you probably use ram + virtual memory, and have more ram than needed to compensate for the fact that you don't keep the most used keys in memory, which a DB will do for you.
Point is if you have petabytes of data and access patterns only mean you access a subset of it, even Mongo might be cheaper to run.
So we are comparing here "just mmap" with "mmap + all that connection handling, query parsing, JSON formatting, buffering, indexing, whatever stuff that MongoDb does".
And no, MongoDB is effectively never a cheap solution. They are used because they are super convenient to work with, with all things being JSON documents. But all that conversion to and from JSON comes at a price. It'll eat up 1000s of CPU cycles just to read a single document. With raw mmap, you could read 1000s of documents instead.
In MongoDB conversion to and from raw JSON into BSON (Binary JSON) is done on the client (aka driver) so the server cycles are not consumed.
Mongo doesn't convert to and from JSON. The driver uses a binary protocol.
Are...are you saying that you've purchased petabyte(s) of RAM, and that that multi-million dollar investment is somehow cheaper than...well really anything else?
> But before that your 10gbit/s node-to-node link will saturate. Oops.
Only if you're returning dense results, which it sounds like you aren't (and there are ways to address this anyhow), which is why I said the issue of saturating network before disk probably wasn't an issue for you ;)
BTW, this is precisely how "real databases" also handle their storage IO internally. So all of the performance cost I have to pay here, they have to pay, too.
But the key difference is that with a regular database and indices, the database needs to be able to handle read and write loads, which leads to all sorts of undesirable trade-offs for their indices. I can use a mathematically perfect index if I split dataset generation off of dataset hosting.
It's really quite difficult to explain, so I'll just redirect you to the algorithms. A regular database will typically use a B-tree index, which is O(log(N)). I'm using a direct hash bucket look-up, which is O(1).
For a mental model, you can think of "mmap" as "all the results are already in RAM, you just need to read the correct variable". There is no network connection, no SQL parsing, no query planning, no index scan, no data retrieval. All those steps would just consume unnecessary RAM bandwidth and CPU usage. So where a proper DB needs 1000+ CPU cycles, I might get away with just 1.
Not using ES here is actually nuts.
A custom cache manager will always perform better than mmap provided by the kernel.
The problem is you haven't explained how the overhead of a DB is too much. Sure, it sounds like a lot of work for your servers and the DB compared to reading from a hashmap.
Where I work right now we fire around 1.5B queries a day... to Mongo.
There are tradeoffs — cloud removes much of the physical security risks and gives you tools to help automated incident detection. Things like serverless functions let you build out security scaffolding pretty easily.
But in exchange you do have to give some trust. And I totally understand resistance there.
Doesn't cloud increase the physical security risks, rather than decrease/remove?
We don't use AWS, because our use cases don't require that level of reliability and we simply cannot afford it, but if I needed a company to depend on IT that generates enough revenue... I probably wouldn't argue about the AWS bill. So long, prepaid at hetzner + in-house works good enough, but I know what I cannot offer with the click of a button to my user!
I run two critical apps, one on-prem and one cloud. There is no difference in people cost, and the cloud service costs about 20% more on the infrastructure side. We went cloud because customer uptake was unknown and making capital investments didn’t make sense.
I’ve had a few scenarios where we’ve moved workloads from cloud to on-prem and reverse. These things are tools and it doesn’t pay to be dogmatic.
I wish I would hear this line more often.
So many things today are (pseudo-) religious now. The right frsmework/language, cloud or on prem, x vs not x.
Especially bad imho when somebody tries to tell you how you could do better with 'not x' instead of x you are currently using without even trying to understand the context this decision resides in.
[Edit] typo
Might have always been that way? We just have so many more tools to argue over now.
But my point wasn't about how precisely the hardware is managed. My point was that with a large cloud, a mid-sized company has effectively NO SUPPORT. So anything that gives you more control is an improvement.
umm, what happens when one fails?
With large cloud my startup had excellent support. We negotiated a contract. That's how it works.
- A large ZFS pool of SSDs is much faster than any cloud storage.
- Cloud storage failed much more often than the SSDs in our pool.
- "Noisy neighbor" is an issue on the cloud
That redundancy, and the performance that scales due to it, place cloud services in an entirely different class from on prem servers.