How is it compared to kafka?
I am confused by this. The format of Kafka's log files is designed to allow reading and sending to clients directly using sendfile, in sequential reads of batches of messages. http://kafka.apache.org/documentation/#maximizingefficiency
When consumers fall behind, they start to request data that might not be in the page cache, causing things to slow down.
Pulsar separates storage into a different layer (powered by Apache Bookkeeper) which allows consumers to read directly from multiple nodes. There's much more IO throughput available to handle consumers picking up anywhere in the stream.
I just want to clarify this - you're limited to N concurrent consumers for N partitions per consumer group.