Infrastructure for Data Streams
vilkeliskis.com
vilkeliskis.com
Given this, why do people deploy it to AWS? It seems like an invitation to disaster.
http://docs.aws.amazon.com/kinesis/latest/dev/service-sizes-...
As far as the API itself though, as is usually the case you might need to get a bit creative to design around AWS imposed constraints due to their design decisions. It might end up not feasible, or not worth it, for your particular use case. But I believe it's at least worth investigating/discussing.
It's a pity, because it would be great to have an open source component for the Kafka architecture, but where data-loss is unacceptable. Hopefully this will be fixed in Kafka itself; I've also been working on something that is based around Raft.
Most of the downsides of Cap'n Proto also don't apply here. Compressing with Snappy will elide all the zero-valued padding bytes. The format of an HTTP message is relatively stable, so you don't get a lot of churn in the message layout. HTTP doesn't have a lot of optional fields, so that's another potential source of Cap'n Proto bloat that doesn't apply to your use case.
It obviously happens some times [1] [2], but it should be more common...
[1] http://alvinhenrick.com/2014/08/18/apache-storm-and-kafka-cl...
But in it's current state, that patch is a starting point that is really intended more for Kafka developers than for Kafka users. I really like what the Mesosphere folks have done -- great variety of OSes and cloud platforms, plus they do all the heavy lifting of bringing the cluster up for you.
A complex kafka setup is pretty involved (relatively speaking) and becomes domain specific pretty quickly, at which point, it probably becomes less usable / understandable to someone who is trying to learn / get interested.
On the whole, I agree with you. Some of the open source software that exists now is truly amazing and I think lots of people are defaulting to less-the-optimal solutions because they just don't have exposure to the latest and greatest.
Kafka supports replication and fault-tolerance, runs on cheap, commodity hardware, and is glad to store many TBs of data per machine. So, retaining large amounts of data is a perfectly natural and economical thing to do and won’t hurt performance. LinkedIn keeps more than a petabyte of Kafka storage online, and a number of applications make good use of this long retention pattern for exactly this purpose.
From http://radar.oreilly.com/2014/07/questioning-the-lambda-arch...
> By having a notion of parallelism—the partition—within the topics, Kafka is able to provide both ordering guarantees and load balancing over a pool of consumer processes. This is achieved by assigning the partitions in the topic to the consumers in the consumer group so that each partition is consumed by exactly one consumer in the group. By doing this we ensure that the consumer is the only reader of that partition and consumes the data in order. Since there are many partitions this still balances the load over many consumer instances. Note however that there cannot be more consumer instances than partitions.
Kinesis provides the same ordering guarantees. They use different terminology (Kafka topics == Kinesis streams; Kafka partitions == Kinesis shards) but have the same system interface. The details of the APIs used for consumption differ, but they provide the same basic functionality of Kafka's "consumer groups".