How to Build a Data Startup
forbes.com
forbes.com
exactly.
- totally making that up, but it is something you could feasibly derive from the data. Mostly it was along the lines of "teens get acne, american teens would rather use a pill than a cream so develop a pill for them, cream for the rest of the world."
1) Take other people's data and learn how to represent it in a way that will let them understand it better. See FlowingData, Stamen, and NY Times graphics http://www.nytimes.com/interactive/2009/03/10/us/20090310-im...)
2) Take other people's data, glean information from it, and offer a new service based off that information. See FlightCaster and TweetFeel
3) Develop tools that will allow people to play with their own data and do their own analysis in-house. See Datameer, Google Prediction API, and Palantir.
Whatever the route, you'll probably need someone who is comfortable dealing with databases, a graphic artist to make the information pretty, and someone with the algorithmic knowledge to capture new insights from massive amounts of data.
You're much better off reading the originals, which I think have already been posted on HN anyway:
http://datasyndrome.com/post/1375987697/analytic-product-tea...
and
http://petewarden.typepad.com/searchbrowser/2010/10/how-to-t...
I think a data startup needs two founders- a hacker to collect and analyze the data and a business guy to provide actionable recommendations and sell it.
But then again, I'm currently a single founder working on a data startup and wearing all of the hats, so what do I know.
Shameless Plug: Anyone want to analyze huge datasets and create recommendations with me? Email me!
http://datasyndrome.com/post/1375987697/analytic-product-tea...
Case in point: I'm currently hacking together an inefficient, unoptimized prototype analyzing pretty large datasets on probably the worst architecture for this kind of thing known to man, and the whole thing still runs pretty well on a single $50 VPS.
The startup I founded had analytics code in a ton of iPhone applications and was handling the load just fine right up until the day it suddenly wasn't. By that point we had customers who relied on us, and we had to deal with it very quickly. Not fun. And there's certainly more to scaling than just cheap architecture. We thought EC2 would handle the overflow until we unexpectedly became completely I/O bound. Firing up a few more instances can't fix that.
If you're just running some scraper and can control what you're taking in, that's a completely different story.
Some data startups I've seen as well as my own project take in existing data sets and simply generate reports from it for customers. Makes it a lot easier to scale.
While the former should not be allowed to impede one's progress toward a MVP, real customer feedback, and the potential need to adapt or pivot, neither should one ignore early optimization decisions where they are inexpensive and may only minimally impede (if at all) that progress.
Being able to recognize the difference is a talent that comes with experience.
- Datasets needn't be huge to be high value. In these instances, scale is not and should not be your primary worry (esp day one). - Figure out if people want your data before worrying about scaling it.
"Data startups need three bodies (hustler, designer, prodineer). Talk to customers early. Here are the levels of knowledge: 1) data, 2) charts, 3) reports, 4) actionable analytics; higher numbered levels are more valuable."