Amazon Announces new Data Warehousing Product
aws.amazon.com
aws.amazon.com
Switching to Amazon would involve rewriting your ETL process, and retooling your reporting software, and converting all your existing, currently-used reports.
A huge expense in data warehousing projects isn't the hardware - it's the consultants, the time, and the people to support the thing. I'm sure this is a great solution for companies looking to start a data warehouse, or maybe companies looking to revamp their reporting environment completely... but other than that it'd be a hard sell...
I'm a big fan of AWS. But, like any other tool, it's not meant for every job.
From courts to records managers/custodians everyone is still trying to understand those questions. In my experience, when in doubt, big business decides the safest legal answer is "probably not".
Handing our customer lists, source code, finance and sales data to Amazon in plaintext form seems naive to me. There's lots of people at Amazon, and it only takes one ambitious middle manager who wants to get noticed by cleverly anticipating the competition. Most likely there's no audit trail, and no chance of getting caught.
What do you think, am I unreasonably cynical?
They will not have any skin in the "middle managers" personal game and so his only other resort is straightforward hacking which he could do in your data center anyway.
Nah. The cloud is as safe as your data center - with the exception of bad apples at Amazon (same diff at your data center). Its servers, in data centers, virtualised.
At this level, I suspect you will not even multi tenant with others above a certain price point.
Amazon might have good internal security procedures - but this stuff can't be audited effectively, we can only take amazon's word for it. Taking their word for it, with the security of all your customers' data, is a big ask.
Edit: " this stuff can't be audited effectively, we can only take amazon's word for it". Or a trusted third party. Go ask your aws sales rep about pci, fisma, etc.
Things you can do:
- Use another provider
- Disk encryption / DB encryption
- Audit audit audit
Hospitals & doctors outsource, and as long as the provider is HIPPA compliant (which AWS is[1]), your data is probably out there already.
[1] http://awsmedia.s3.amazonaws.com/AWS_HIPAA_Whitepaper_Final....
Careless or corrupt health staff releasing my information without my permission? Well, it doesn't really matter where the data is stored.
There were always only two defensible privacy fronts: keeping the data off electronic records or filling the records with shit.
This mindset screams "I AM IN THE VALLEY AND EVERYONE WHO ISNT UNDER 30 AND USES APPLE PRODUCTS DOESNT GET WEB2.0" (also caps lock is cruise control for cool).
Apologies for the negativity, I think I get it, I want my data to be in the cloud, and easily accessible and all that jazz, but I want to keep it secrete and safe and most importantly I want to be mine.
Says the guy who just signed up for the iCloud today... ;-)
A lot of other commenters immediately jumped to the medical records argument, but all I was saying is that for a LOT of companies that make the "we have to have everything on site" argument...it's just not true.
But I agree with you, the medical records argument is kind of boring. But, not everything needs to be outsourced; There is value in keeping things on site, if not for anything besides job creation!
My pet-peeve in this is that it has been now for a while (and is trending upwards, fast) that we don't see any problems at all, long or short term with simply "shipping it to the cloud", where it is everything from medical records, to phone contact lists to personal communications with our other significant other.
We as a community are quickly eroding any expectation of privacy and security all in the name of being agile. I guess it just rubs me the wrong way.
I should tweet about it, on my iphone and then copy it to a file for prosperity and upload it to my google drive...
The problem falls into two categories, on one hand you have non technical end users, and it takes a non trivial amount of time to train them to roll their own crypto if you will, and it's also hard to convince them it's worth it (This is a fair point, as security is a cost/benefit between ease of use and not getting caught with your ass in the wind).
On the other hand, you have companies using outsourced services, and with SaaS/PaaS/aa becoming all the rage, it's very important in my opinion that those service providers shoulder some of the responsibility to not let their users, serve their users etc in a manner that's not conducive to security/privacy etc...
Punting this problem up the stack, with it most often ending on the end users desks, is IMNHO a bad idea, since then, as it is now all those good things crypto promises are the exception, rather then the norm.
This is obviously much much much more complicated in practice, but I at least see this problem reflected in the "to the cloud!" mentality.
EDIT: Complete rewrite.
But can't we say that we have both a moral and an ethical obligation to protect our nontechnical users or our fellow developers from mistakes, lack of training or in the worst case maleficence ?
The business decision of using an externally hosted backend services, what ever they may be must take into account what data goes into it, out of it and how it's computed on by both you and the provider w.r.t. who the real end user is and how the data is going to live on.
And here I think is the crux of the problem, those questions and their solutions are generally very hard when put into practice (I dont have a silver bullet, or even a something vaguely resembling a mold for it) so it's not very conducive to being a "Fast" company.
For example, being European, it scares me a great deal that companies, schools and the public sector are increasingly punting the business decision of "how to handle email" to "let's use gmail".
That in no way takes into account my concerns (and often I do not have a choice in the matter of using these services) since my mail, and by extension a large part of my life is being handed to a for profit US corporation who "does no evil".
I use gmail privately though, since I did this particular cost/benefit and decided that i dont really care if google reads my mailing list traffic...
I find your observation that this problem is reflected especially in service oriented architectures questionable. By centralizing all resources (including documentation: http://aws.amazon.com/security/) It makes it easier to enforce best practices and standard interfaces. But just because they can doesn't mean it's always a good idea to do that.
So instead of building 1 cloud storage service, you need to effectively build a cloud storage factory, so you can deploy N cloud storage services on demand.
At that point you also potentially deliver a product (a rack of hardware) not just a pure service.
A Dell MD1220 with 24 Crucial M4 512GB SSDs will run you $12,600. That's 12TB. Multiply by 4, enable compression, etc etc.
You could buy two of those setups, pay for power, cabinet space, and bandwidth, have a ridiculous amount of IO available, with single-digit millisecond latency, pushing 6Gbps and still have money to burn. And it'll take you (much) less time to unbox and setup than it will to push that much data up to AWS.
Granted, this new service is probably only 50% more expensive than hosted your own, and if you have zero IT staff, it might make sense in some scenarios, but it's definitely not a no-brainer.
It doesn't need to compete with Terradata. It needs to compete with Dell, and in that field, it's still the more expensive option by a significant price margin, as well as being (odds are good) at least a couple orders of magnitude slower.
For most applications does it really matter whether the data sits in your data center or Amazons? Nope... cause the organisation your company contracted to manage already has full access to all your secrets.
So really Amazon is just another IT outsourcer except you don't need a long drawn out sales process.
Many startups like Qubole have been already working on providing such cloud based solutions for data analysis.
The bigger question for me is why Amazon has been able to figure out the technical details necessary to run this kind of service for this price. It's just ridiculous. Talk about taking the oxygen out of the market...
http://www.slate.com/blogs/moneybox/2012/10/26/amazon_profit...
I guess they grew the infrastructure for themselves, optimizing it bit by bit over the years. And then noticed that it could be sold too.
I'm building a tool that allows business people and non-technical analysts to query their data warehouses using natural language. (Currently, you must ask a technical person to write ad-hoc queries for you, or build you a dashboard. This bogs down your data people.)
Does anyone have insight into the demand for such a product?
[edit: I'd love to chat with anyone with insight into this topic. Reach me at Joseph at metaoptimize dot com]
They call me up, ask me to do a "quick report across the inventory db with the project cost data." I send it off to them. If they like it, we push a report (maybe with a couple of parameters) into production.
My gut is that we aren't lacking for good technical options in analytics and data warehousing. To be honest, the lion's share of my work in data warehousing is helping the users know what questions to ask.
But there is lots of room and probably several excellent lifestyle to 8-digit businesses for good BI.
Depends. Back when I did DW stuff my general workflow was to speak with the analysts about what they were trying to accomplish. From there I would create the cubes and additional metrics. I would also set up all the processing schedules at this time. The analysts would then use an Excel plugin that provided a pivot table interface to any cubes for which they had access. It worked pretty well.
For straight data access I would teach the them basic sql and/or build sql templates for them that they could extend.
My goal was always teach a man to fish and get out of the way.
At ExxonMobil, a place I worked, you're going to have VP's asking eachother and IT is going to hedge with, yea if we do this then project X will be late (it's going to be late anyway but they've kept quite about it and no one knows).
My personal solution when I needed a query was to bring a six pack of beer down to IT friday afternoon, mostly because I wouldn't be given access to write queries because we had BI software.
[1] http://en.wikipedia.org/wiki/Dimensional_modeling
[2] http://www.amazon.com/Data-Warehouse-Toolkit-Complete-Dimensional/dp/0471200247To speak to OPs point about difficulty in querying data warehouses, most business intelligence tools that I'm aware of provide semantic layer[1]-type capabilities, whereby the user interface of the tool is presented in the language of the business domain. Nevertheless, I still agree that this is still difficult work, unfortunately. That it is getting more complicated in some respects, such as through unstructured data, doesn't help either.
Enormous, and there are dozens of such tools available.
Most of them work best if you build an actual data warehouse -- dimensionally structured, not normalised. This is because they can easily build query forms using the DW dimensions in a language that makes sense to end users.
* Yahoo's Everest
* Greenplum
* Aster Data
All mentioned here: http://www.cubrid.org/blog/dev-platform/database-technology-...
The Register wrote that Amazon's solution is a column-oriented database possibly based on Postgres, like Yahoo's:
http://www.theregister.co.uk/2012/11/28/amazon_aws_redshift_...
In the last one, RDS and ELB had issues (and were flagged as such on their status board) due to needing EBS, but I don't believe DynamoDB was.
The discussion about Amazon Redshift begins at 52:50 http://www.youtube.com/watch?feature=player_detailpage&v...
This quote from the product page seems to indicate that EBS is not used for primary data storage: "it runs on hardware that is optimized for data warehousing, with local attached storage and 10GigE network connections between nodes."
Google has systems like this for analyzing its request logs. Think of how many HTTP requests hit Google's front-end servers per second or hour or day. Each one has a few dozen pieces of data associated with it -- URL, client IP, headers, etc. Suppose I want to make a bar chart of how many requests came from France containing a certain header, each day for the last year. The system can do this query quickly if the requests are already bucketed by time interval, organized by column, compressed, and stored so that exactly the information needed can be brought into RAM quickly.
It is a little funny, when you step back, that "storing," "archiving," and "warehousing" are different things and Amazon has services for each. Try explaining the difference between S3, RDS, EBS, Glacier, and Redshift to a layperson.
wow.. I just finished reading the sci-fi book a few weeks ago - "Redshift Rendezvous" by John E Stith. I wonder if this is where the name comes from? In the book Redshift is the name of the space ship that runs cargo mission through folded space, the obvious problem that since you are traveling within just a few m/s of the speed of light just walking on the ship while underway causes color shift - thus redshift.
I read that Stith has a physic degree and worked as an Engineer for NORAD Cheyenne mountain. That made me really interested in what novel he would come up with. http://www.neverend.com/short-bio-john-e-stith
I took a few screenshots from the keynote and included one showing the mention of Postgresql and ODBC/JDBC support. Included here if you want to see for yourself: http://wp.me/p2sRpx-1e
$1474 per TB per year for storage alone ($0.12 * 1024 * 12)
plus
$35.84 per TB queried
Amazon is definitely cheaper.
Can we run more complex in-database processes implemented as stored procedures on this platform or is it going to be limited to pure SQL querying/analytics?
And does anyone have an idea how to upload 1 TB of data to this service using Internet connection from your in-house company server? ;)
http://aws.amazon.com/importexport/
I am assuming that you have 1TB to start, not generating 1TB per day which obviously changes the equation.
Redshift is a different usage model. You upload your data once, then ask questions of it - but you don't update it. Google does have something similar to Redshift: BigQuery (https://cloud.google.com/products/big-query).