Kafka in a Nutshell
sookocheff.com
sookocheff.com
The reason I don't just use kafka is because the "quorum" style scaling is overkill for our needs since the app will be for internal use only and will likely never exceed more than 1000 simultaneous users. Also, I've heard that zookeeper (required to use kafka) has its own technical overhead that I'd like to avoid dealing with if possible.
https://engineering.linkedin.com/kafka/running-kafka-scale
When combined, the Kafka ecosystem at LinkedIn is sent over 800 billion messages per day which amounts to over 175 terabytes of data. Over 650 terabytes of messages are then consumed daily, which is why the ability of Kafka to handle multiple producers and multiple consumers for each topic is important. At the busiest times of day, we are receiving over 13 million messages per second, or 2.75 gigabytes of data per second. To handle all these messages, LinkedIn runs over 1100 Kafka brokers organized into more than 60 clusters.
I'd also be careful when saying "these types" of pub/sub messaging systems as Kafka, while not unique, is pretty lonely in the space it targets. Most pub/sub systems don't make nearly the data garauntees that Kafka does (and therefore Kafka isn't appropriate for some pub/sub patterns).
In practice a broker uses very little CPU even with compression and compaction happening, and is bound by the disk or network speed--our brokers on AWS run about 80MByte/sec with room on top for bursts to ~110. Interestingly and in addition, blocks of messages can pass through Kafka without being decompressed because of this design.
Everything happens from one port on the brokers, who are backed with Zookeeper for cluster management; consumers may also connect to Zookeeper or similar to track the latest offset that they have processed.
implying overworking
what do you replace it with?
what problem was 'use the same number of partitions as consumers' trying to fix?
Producer/Consumer semantics are pretty similar. Partitions in Kafka are Shards in Kinesis terminology.
One big difference is retention period in Kinesis has a hard limit of 24 hours (no way to request increase on this limit).
Kinesis IMO is easier to use being a managed service. I have performed a Kafka to Kinesis migration & have found Kinesis easier to use. Plus, AWS Lambda makes consuming Kinesis a breeze (if your usecase suits it).