I haven't had a chance to actually read your post yet, but I'm really looking forward to it.
I haven't had a chance to actually read your post yet, but I'm really looking forward to it.
We used Terraform instead of CloudFormation, although there are a few places it didn't/doesn't cover.
We're also doing protobuf -> parquet instead of JSON; Scala instead of Python, and one of our major feature/issues is that the incoming data is out-of-order; Firehose outputs to partitions based on when the event arrived, and we're repartitioning based on when the event occurred (according to a timestamp in the message)
I see you ran into the capitalization issue too :) I ran across docs somewhere that said something along the lines of Glue downcasing column names, which definitely fits observed behavior.
I'm hoping to turn what we've done into a similar post, although I think it'll be a month or two before I can get to that.
Edit: Oh, I also got a lot of mileage out of the Zeppelin Notebooks. Way better than a raw dev endpoint, but watch out on the cost for both :)
The notebooks can be provisioned with almost a single click from the Dev endpoint console, but they changed the recipe halfway through my work on it, and now you have to also SSH into the box and run a script to setup some of the security. :/ Still totally worth it, tho.
Thumbs up on the Glue Dev endpoint. It's been killer. I had trouble setting up a Notebook (I wanted to get fancy with Docker) and I usually use the Python repl link that's provided.
I'm working on a follow up post that removes the Data Pipeline -> Lambda and uses the new Glue DynamoDB integration.
Looking forward to your blog post!
What was the trouble you had with a notebook? I can probably post up some of our Terraform code (which includes notes on the parts Terraform doesn't cover).
Oh, yeah - there was also the S3... VPC? endpoint. That needed to exist.
There were a lot of wires, and Amazon documentation is decent as a reference but rubbish as a tutorial :/