Amazon Redshift Now Available to All Customers
aws.amazon.com
aws.amazon.com
We're going to work on SnowPlow-Redshift integration next week, using the COPY command + SnowPlow S3 event files. It's great timing as we've been hitting the limits of what we can do in Infobright (which inherits MySQL's limit of 65532 bytes per row - an unfortunate restriction for a columnar database).
The Postgres JDBC driver when you try and do batch inserts runs each statement individually and you end up inserting 10s of rows a second.
I wish they had gone with something like Vertica.
Column stores are destined to have slower inserts, due to how the data is stored on disk. But if they are actually committing each statement individually, that is a problem.
Why do I have to write code to perform an extra step and pay the extra cost and latency of pushing data through S3 just to get it into Redshift?
Not supporting trickle loading is a leaky abstraction IMO. It's not a ton of code to log statements until you have enough to justify an import and you shouldn't push that complexity on every database user.
Postgres supports copying from a binary stream, why not support that?
Our goal is minimal architecture complexity, and to upload log files or other data to a file system before loading it into a data warehouse just doesn't make sense.
We're currently looking into Hadoop/HDFS/Impala due to cost constraints (Vertica would have been our primary choice). If anyone has any other suggestions it would be great to hear them.
I've looked at Snowplow (https://github.com/snowplow/snowplow) -- is that what most people are using, are you rolling your own, etc?
Seems like an "odd" product that kind of compete's with Amazon's existing offerings in many ways...
http://www.zdnet.com/amazon-redshift-paraccel-in-costly-appl...
How does speed/performance compare to something like that shown here: