Amazon Kinesis Firehose – Simple and Scalable Data Ingestion
aws.amazon.com
aws.amazon.com
I want to be able both reading at different offsets in the stream AND backup it to S3 or ingest into Redshift.
With current offering I need to duplicate data into two different services with different APIs.
If you're already doing all the Kinesis shard management/KCL willy nilly why not just dump it to S3 yourself? Firehose seems to be targeting users who don't want to deal with sharding.
There are competing products, like Google Cloud Pub/Sub where there is no need to manage shards manually or run your own workers, like KCL.
$ sudo pip install awscli --upgrade
...
$ aws firehose help
FIREHOSE() FIREHOSE()
NAME
firehose -
DESCRIPTION
Amazon Kinesis Firehose is a fully-managed service that delivers
real-time streaming data to destinations such as Amazon S3 and Amazon
Redshift.
AVAILABLE COMMANDS
o create-delivery-stream
o delete-delivery-stream
o describe-delivery-stream
o help
o list-delivery-streams
o put-record
o put-record-batch
o update-destination
FIREHOSE()Is dodging those HTTP POST fees the value-add over simply using the S3 HTTP API yourself ?
> Storage
> You will be billed separately for charges associated with Amazon S3 and Amazon Redshift usage including storage and read/write requests. However, you will not be billed for data transfer charges for the data that Amazon Kinesis Firehose loads into Amazon S3 and Amazon Redshift. For further details, see Amazon S3 pricing and Amazon Redshift pricing.
But to me it wouldn't make much sense to use Kinesis Firehose unless the fees were cheaper than what it would cost to utilize AWS Lamdba for the same work. I mean it can't be all that many lines of nodejs code to pop events off a stream, flush them into a tempfile in batches, HTTP POST those batches to S3 with appropriate error handling/retry logic.
Admittedly I haven't crunched the math but just by eyeballing their pricing I suspect it may be more expensive than using your own lamdba. I wonder if Kinesis Firehose is implemented internally as an AWS Lambda, it wouldn't surprise me.
I suppose that eventually the pricing on this service will drop once someone open sources such a lamdba, especially since installing a lambda is so remarkably simple.
In other words, you can make small writes to Kinesis and then read out in larger amounts and write larger files to S3. This is a huge optimization for any job that runs across the data in S3. Many small files can really undermine performance in something like Hadoop MapReduce because of the additional request overhead.
Or, if you're using the new Firehose product, you could also use Lambda and attach it to the S3 event source using the bucket that the Firehose dumps into to perform batch processing on these records.