The big data disaster
thejonathanmacdonald.blogspot.co.uk
thejonathanmacdonald.blogspot.co.uk
We've always had data that defied structure, but now (with NoSQL, map-reduce, etc) we have a better idea of how to deal with it. We've still got data processing tasks that can easily be parallelized and others that are inherently linear.
Amazingly, when you actually have "Big Data", you need to analyze how best you could process it and architect your systems and software based on the facets listed above. In a corporate BI environment, you should know what questions you're trying to answer before you start. Unfortunately, in a research environment you don't always know what you're looking for.
We're trying to bring research scientists and computer scientists to optimize how data is analyzed at the Penn State University ... check out @HackingScience on Twitter.
I like to think that big data is more about aggregating the information available in multiple datasets, and then figuring out how you can read those entrails, as it were.
Honestly, big data today is mostly about a potential replacement for data warehouses. The real revolution around big data is that we now have a place to put it for analysis. That was a big hole that Hadoop filled. Now, the next 10 years will be spent actually figuring out how to analyze these big pools of data effectively. Frankly, that's the harder task, I think.
We now have the ability to ask questions of ALL the data. Too bad we have no clue what the questions should be.
Meanwhile, the asian grocery store down the road has old credit card terminals and no phone tracking or other gimmicky crap. They sell cheap vegetables and are PACKED and are huge and they rent out their space to other vendors, Walmart style. No place to move in peak shopping time. This is because of their greater selection and vastly lower prices.
There are other factors, but I don't think keeping track of what everyone buys will give them some insight to raising prices on a few items.
The IT edge was realized by companies like Walmart and Costco in their supply side and inventory, I don't think big data will give them a similar edge which is what they are aiming for. Like Target gave targeted advertisement to pregnant women, will that edge really vastly offset their number crunching IT costs?
If they can increase their revenue by 10%, and it cost them 5% of their revenue to do, then it's worth it.
You're missing what they'd do with that data.
Keeping track of what you buy means they can email/text/whatever to say 'hey, this %reallyinterstingthing% is on offer'. And really interesting thing really is interesting to you.
It's like an article I read a while back about how supermarkets can predict if you're pregnant by the things you're buying before you've even told anyone.
That's valuable to them, although extremely creepy. They can start bombarding you with vitamin supplement adverts or books or whatever and you're more likely to buy as you actually want those things.
If you've taken an introductory microeconomics course, you'll remember that almost all the models of a "rational" firm involve maximizing profit by picking a price point where people still buy yet will pay a maximum price per unit. How firms determine that demand curve is left as an exercise - it's assumed that firms who get it right will survive while firms that don't will go out of business, leaving only firms who get it right. This is cold comfort if you're a business owner and want your firm to get it right.
So what do the big retailers do? They A/B test. They divide customers into two groups (say, those who have a Safeway card and those who don't) and offer a different price to each. Then they measure how much of each good people will buy at each price, and presumably do some statistical corrections for demographics, the "sale" effect (where people buy whatever's on sale regardless of price), folks who won't use coupons no matter how cheap, etc. They end the sale, and then pick a different price next month. Over time, they build up extensive data about just how high they can raise their prices before people stop buying entirely.
So when you ignore the discount, that's Mission Fucking Accomplished for the store. It's telling them that the discount doesn't matter, and you'll buy regardless, and so they should try raising the price instead. They don't care about the sale, they care about the data, so that they can make more money from all their other customers.
(a) So that you do, in fact, buy them—turning that "hmmm maybe I should get this thing" into "yes, I will get this thing, especially since I can save money on it!"
(b) So that you buy them from me instead of from someone else who didn't give you a coupon. (And while you're here, you're probably going to decide to pick up some other stuff on your grocery list...)
If you raise the price too much I'm going to find myself a different store.
Did many dive head first into the social and mobile bandwagon? Sure. But to say there is no bandwagon or profit to be made in these niches is not realistic.
Everyone knows there is a hype, but that hype is based on reality. Big consumer data already has a value per GB.
Some say Big Data is just Business Intelligence over a large amount of data. I don't disagree and I see a solid future for BI.
Managers and consultants should prepare for the big data storm that is to come. Learn about the possibilities and impossibilities of cloud computing and Hadoop. I believe in a few years even tech-savy consumers (who now own an iPad) are able to use Big Data on their own data. Anyone can already rent a few Amazon resources to compute, and work with technology invented at Google a few years ago (MapReduce, BigTable). Or if you don't want to reinvent the wheel, go for Big Data as a Service, like Cloudera or Splunk.
In exchange for discounts, consumers will hand over their data to companies for free. Coordinates or purchasing behaviour: everything gets stored. Companies are busy exporting old data from tapes, so they can run completer aggregate queries.
When more big data analysts emerge, some companies will be ready: They'll have datawarehouses filled with data. If you do not invest in big data, you'll soon lose your competitive edge in regards to BI and consumer data.
The people that are extolling the virtues of big data are usually external consultants. They were the same that told companies to focus on mobile apps, and before that, to create a social presence. You can rail against that, but that is merely an epiphenomenon of every hype: Not what is really about.
All big tech companies are preparing for the future of big data (Amazon, Akamai, IBM, Microsoft). Facebook has the largest known Hadoop cluster in the world (over 100PB). Google might process around 25PB of data each day...
For any five predictions that are even remotely possible, any business newspaper from any day in history will contain at least one example of a company that meets one of those predictions.
for me "big data" is the emergence of scalable tools for handling data sets that are much bigger than you can handle with one computer. (ex. Hadoop)
there's definitely a feeding frenzy here because the spend on hardware could get high (say 200 powerful servers), therefore the spend on associated software and services could be high too.
a few issues stalk big data
for any IT project there's the issue of how much value you can create per bit. Imagine Facebook has a $10 ARPU, for instance. They can't afford to use more than $10 a year worth of hardware and electricity to offer you the service.
Thus, if your "big data" project is going to cost $500,000 you ought to have a plan to create more than that in value.
often the best way to deal with "big" data is to make it small. I had my eyes on a 50 TB data set a few months ago but didn't have the budget to work with the whole thing.
I realized I could take a statistical sample of just 50 GB which I could download to my house and process on my cluster here. I can't get an accurate answer to ~every~ possible question with my 50GB sample, but I got answers to the most pressing questions and discovered that a 1 TB sample would be good enough for the toughest questions I wanted to ask.
This article is garbage since it totally discounts all the other uses of big data. I, for one, work in the risk analytics sector where 'big data' is mined to try and identify which assets have the biggest impact in a risk profile.
It is baseless because it is passive and without direction.
I'm not bashing it, as a matter of fact it excites me as a programmer.
Edit to add: On re-reading the article I see that I misunderstood a general statement for a rhetorical denouement. Still, the example I gave is a strong one, although not a prediction since it already happened (indeed, it may have prompted this very blog post.)
What I mean by that is you can optimize one or several values to death, but if the core product is broken, you'll still lose all of your customers. So while you'll get a 1% lift over control for making the button blue over green, what you failed to notice is that the product sucks.
Big data (née "Business Intelligence"?) has its place, but beware of it becoming your entire product strategy.
Not very clear on what he's talking about.
This post should be called "The No Data Blog Disaster" for assuming we are mind readers.
To put it into perspective, people were working with 1M+ record datasets in the 1990s. I tried loading a 300k record file in Excel 2010 the other day and it crashed.