Apache Kafka – Publish-subscribe messaging rethought as a distributed commit log
kafka.apache.org
kafka.apache.org
TL;DR: "Kafka’s replication claimed to be CA, but in the presence of a partition, threw away an arbitrarily large volume of committed writes. It claimed tolerance to F-1 failures, but a single node could cause catastrophe."
In my experience network partitions are not the most common kind of partition although they are well represented, especially on some network topologies.
IMO you should never assume that any two things won't be partitioned as part of your expectation of correctness or availability.
Broadly, I think sacrificing P is pretty deeply misguided. If you run a distributed system at any scale for any period of time, you're going to experience network partitions, even within a single DC. (I speak from experience: I work on a distributed system which runs a whole bunch of machines in a whole bunch of datacentres worldwide. We see a reasonable number of non-trivial within-DC network partitions every year.)
A much more detailed argument of the above: http://codahale.com/you-cant-sacrifice-partition-tolerance/
You really do need to think about what happens to your system when it gets partitioned. This paper has much of interest to say on the topic: http://cs-www.cs.yale.edu/homes/dna/papers/abadi-pacelc.pdf
"we can tolerate N-1 Kafka node failures, but only N/2-1 Zookeeper failures. This actually is sensible, though, as Kafka node count scales with data size but Zookeeper node count doesn’t. So we would commonly have five Zookeeper replicas but only replication factor 3 within Kafka (even if we have many, many Kafka servers—data in Kafka is partitioned so not all nodes are identical)."
(Zookeeper is used to track node membership)
http://engineering.linkedin.com/distributed-systems/log-what...
http://engineering.linkedin.com/distributed-systems/log-what...
Previous HN discussion: https://news.ycombinator.com/item?id=6916557
If you process that much data, Kafka is one of the last things which you'll need to scale out.
I believe Kafka can still lose some data if all the active machines fail. It's a deliberate design decision (it's the right thing to do if you want to remain available and can tolerate some data data loss). I believe the Kafka team are working on it, but it's non-trivial to fix.
http://research.microsoft.com/en-us/um/people/srikanth/netdb...
By "hard real-time" I was referring to latency, rather than throughput. Achieving very low latencies is difficult in Java because you don't know exactly when the GC will kick in, but nevertheless it's possible to get very high throughput.
With Java, you don't have the unmanaged or struct support, so doesn't that really add up? If you go "native", isn't there significant overhead since you can't have pointers in Java (right - the bytecode doesn't support it?)?
People pull it off, but it seems that GC overhead would be a killer.
There's still GC from objects allocated by Kafka in the JVM, but the actual message data doesn't even go through the JVM.
RAII in C++ is a good pattern, but it does not avoid performance penalty of cascading release of resources, impossibility of using exceptions in destructors or thread contention with shared resources.
The only thing to do, is get on your coconut radio and wait for the planes to bring you some cargo. While you are waiting, Kafka is a useful MQ for event processing and is being used to solve real world problems today.
Not sure on what basis you claim the choice of language to be "questionable", but keep in mind that Scala's type-safety and many other features are much more difficult to achieve in C/C++. Cleaner code is sometimes more important than some tiny gain in performance.
Also in terms of scaling, Kafka cleverly takes advantage of many aspects in their design to ensure low-latency high-throughput.
* Little random I/O
* Relying heavily on the OS pagecache for data storage
Performance-wise, Kafka can outshine some of the in-memory message storing message queues.
Source: http://kafka.apache.org/documentation.html#design
http://research.microsoft.com/en-us/um/people/srikanth/netdb...
> Java garbage collection becomes increasingly fiddly and slow as the in-heap data increases.
So they had to work around that. In my view it's not a tiny issue. I'd say, instead of working around such inherent limitations, it's better not to have them to begin with when making high performance systems. That was my main point above. Time spent dancing around such problems defeats the purpose of supposed easiness of development.
The next line reads: "As a result of these factors using the filesystem and relying on pagecache is superior to maintaining an in-memory cache or other structure—we at least double the available cache..."
And if you don't know about pagecache, it's an in-memory cache managed by the OS and has nothing to do with JVM's memory at all.
And you forgot that C++ isn't the easiest language when it comes to designing a distributed system. Scala, as I mentioned, offers many other features that suit the needs of the team. Of course if you're a good engineer you'll know that there are trade-offs such as compiling time, but that's the same for every engineering decision.
When you are designing a high performance system you have to consider everything. Different platforms have different tradeoffs, but the tradeoffs on the JVM have been well proven over time.
BTW, I've talked to finance people who have JVM applications that haven't had a major GC in months. Really, it's not a big deal.
Is this difficult in Java? Yes. It's also difficult in C++. Just because it is difficult doesn't mean it is impossible.
So if I am resorting to managing my own memory anyway, why would I use the JVM? Because typically the code that is on the critical path is a small percentage of the entire code base, and the other advantages of the JVM (tooling, language features, libraries etc.) out way the downsides.
That's not always the case, and I don't have any specific knowledge of Kafka, but just because something needs to be low/consistent latency doesn't mean it can't (or shouldn't) be written on the JVM.
This is a persistent message queue for log messages. Messages are coming in sequentially and subscribers read them sequentially. It makes zero sense to keep tons of messages in memory inside complex data structures as they are not indexed or searched or analyzed.
So in this particular case it's not actually a workaround. It's just sensible design and I wouldn't do it any differently in C++ either.
There is a reason many of these kind of systems are written in JVM based languages. Examples: Hadoop and all of its siblings, Cassandra, Storm, Kafka. So either all of those people in successful projects make "questionable decisions" or your knowledge of Java/JVM performance has been outdated by the new developments of the past few years...
Java compilers are lacking a bit behind, but are quite comparable for distributed computing, up to the point I seldom use C++ on the job nowadays, other than when replacing existing systems with JVM/.NET ones.
For Java 9, there are already some features being discussed that will help Java compilers to improve code generation, like value types, better FFI and making unsafe an official package for those cases when there is no way around it.
Kafka's use of the JVM does not really impact throughput. Very little portion of the messages' lifecycle is spent in the JVM heap, Kafka makes aggressive us of the OS page-cache and avoids copies (e.g., using sendfile() whenever possible).
I agree that, e.g., if Kafka had to make heavy use of in-process memory (this is appropriate for databases) as opposed to OS managed buffers, then a language without a garbage collector (or perhaps a garbage collected language that gave you an option not to generate garbage in the first place...) would help.
That said, there's other advantages to using C++, but they won't bring significant performance improvements -- although I suppose one could experiment with, e.g., using AIO but that might end up causing more harm than good in case of Kafka.
As the "slow compilation times" feature, scalac and sbt do a much better job of it than cmake and g++/clang :-)
Some of the highest performance large scale systems in the world are written in Java. Notably, in the messaging market Oracle's MQSeries message bus and StreamBase (now TIBCO StreamBase) run on the JVM. Both are used in some of the largest messaging implementations in the world.
The high-perf messaging bus in the TIBCO lineup is FTL[1], not StreamBase which is a CEP engine.
[1] http://www.tibco.com/products/automation/enterprise-messagin...