An Ode to Little Data
katsenblog.com
katsenblog.com
http://bayesian.org/sites/default/files/fm/bulletins/1106.pd...
1) Extrapolation of data where samples are lacking.
2) Taking big data and making it little data so that you can actually comprehend it.
My goal in practically every algorithm/tool I write is to take a several billion records and condense it into something that can be put in a spreadsheet, put on a chart, or rendered on a map. Analysts roll their eyes every time some big data guru releases another map with 7 trillion points on it. "Oh so you took 3 weeks rendering a map that looks like yet another population map, when you could have rendered this instantaneously with a choropleth?"
These types of statements always prompt me to ask "what is the specific problem you are trying to solve?". Replacing Excel? What is the unique advantage beyond a different UI? I think the real product challenge here is identifying why using a different UI or tool would have ROI for the average business user the author is describing. If it doesn't, why would they switch? What is your startup's tool going to do to increase my profits and justify the investment and switching cost?
I've often found it is the analyst and approach, not the tools that make a huge difference on these problems, and for those who care deeply about tools there are many open source options (R, SciPy, etc). Creating more generic "small data" tools runs the risk of solving a problem customers don't have.
I don't think it's a hard leap to consider that if the big-box statistical package companies realized how much of their money came from industry, they'd do what they could to make their software seem like an alluring proposition. Statistical software costs an order of magnitude more than Excel, so they'd need pretty good arguments on how to sell upper management that the business team actually needed an 8000 dollar piece of software.
I'm not sold that there's nothing in between Excel and R. From my experience, they require a slight learning curve (nowhere near the learning curve of going from Excel to R), but not an insurmountable one. What these solutions lack is the name recognition that Excel has, or decent integration into a MS Office stack (exceptions, of course--I remember seeing a statistics toolbox for Excel once), or they cost too much.
I think part of the real problem is that for a lot of companies, Excel is "good enough". There's plenty of stuff it can't do well. It chokes on larger data sets, has limited statistical functionality, poor scripting capabilities, and shaky random number generators. But it's good enough for people who don't want to do much with their data.
If they wanted to do harder-core analysis, they'd outsource it to an analytics team. This perspective comes from a viewpoint inside BigCo. The prohibitive cost of some of these solutions might be a harder pill for a small firm to swallow.
That said, I certainly think Paul's approach makes a lot of sense, and I hope someone builds that. Closest thing I can think of is http://anapsis.com/
</pitch>
We haven't figured about how to extend that to visualisations, which I think is what the OP is really looking for as well.