Four billion messages an hour: benchmarking Deepstream throughput
deepstream.io
deepstream.io
We're handling 4m to 6m 1.5k log lines per second using Apache Kafka on a cluster of around 100 nodes.
2x Intel Broadwell E5-2630 v4 (10 core, 2.40 GHz)
128GB - 4X DDR4 32G ECC Registered 2133MHz DIMM
6 x 2.5" 240GB Intel SSD 540
x1 S2S 1G LAN/B W/OCP BRKT ASSY(1IN1)
x1 Solarflare 10G SFP+ SFN7002 PCIe card
TPM Module
Dual Power Supply 1100W each (x2 / chassis)
We have this stuff built by Quanta for us.Each log line (in JSON) is roughly 1.5kB.
Or, in other terms 700mbit/s per host with your kafka setup versus ~30mbit/s per host in the benchmark. Allthough your machines seem to be quite a bit beefier (I wonder if all that RAM is actually used?).
More details: https://blog.cloudflare.com/scaling-out-postgresql-for-cloud...
Sadly it does not give a figure/order of magnitude of the amount of data that's stored in citus after aggregation, but I guess it's just not public information. [I'm working on a system that is somewhat similar to CitusDB (eventql.io) and am always really interested in these numbers]
EDIT: I can't reply to your other comment for some reason but many thanks for digging that up, it's very interesting info to me!
These high level messaging frameworks attempt to provide some degree of reliability and easy clustering out of the box while keeping things simple. For specific domains where performance is paramount, you can always find light-weight solutions that outperform them. For instance, with market data it's almost always better to drop a quote if it's out-of-sequence or got dropped somewhere than to add latency waiting for a retransmission. This simplifies things since you no longer need durability.
That's ~25MB of data per second for an in-memory workload over 6 machines. I think they missing a zero on the size of the messages?
Deepstream relies on garbage collection to free up dereferenced
memory. If a machine's CPU is overutilized above 100% for a
consecutive time, garbage collection will be delayed and memory
can add up. If this continuous for a prolonged period, the
server will run out of memory and eventually crash - so be
generous enough when it comes to resource allocation to make
sure that your processors get some breathing space every once
in a while.
"Be generous when it comes to resource allocation".--
> The costs of running a six-instance cluster for an hour on AWS are 36 cents (6 x t2.medium @ 0.052$/h + 1 x cache.t2.medium @ 0.068$/h)
AFAIK, pricing in AWS world depends upon the region. Bandwidth, hard disks and so on also contribute to the price.
I'm not sure what to make of such conclusions.
It similarly provides data-sync, pub-sub and request response with no opinion about your frontend framework or technology stack and has an open ecosystem of connectors that make it work with all sorts of databases, caches and message buses.
It's also significantly faster than meteor, making it possible to also use it for multiplayer gaming, realtime trading etc...
I see a lot of advantages on using a more standard protocol such as MQTT over deepstream
It's closer to a self hosted version of Firebase or Parse than MQTT
It's not about 260$ it's about tens of thousands of dollars saved by companies i work with after migrating from AWS.
AWS is overpriced service for people that do not have time or resources to do things on their own and pay massive prices for that.
Isn't that the point of AWS?
The time / resources cost to build (monitor and maintain..) your own infrastructure isn't zero.
Did these companies have spare teams lying around the place, such that no new hires were needed to make the transition? If not, then it's not so much of a saving at the current cost of dev / ops personnel, just now on a different balance sheet.
Then you come back when you have due diligence and compliance and availability requirements. And you see how much that is really costing you with the arrangements you have.
Finally if you're super successful, you could self host or otherwise. But "rack costs" (naively) trade off against people and auditors.
Summary: AWS is expensive, especially for bandwidth. Practically if your business requirements are not bandwidth intensive, it ends up being paradoxically cheaper than the alternatives if your business involves legal agreements and things like availability commitments to customers.