The Most Innovative Companies In Big Data
fastcompany.com
fastcompany.com
I've been to Big Data conferences where the main application was mean and standard deviation on a huge datasets, I'm just sick of companies inflating the Big Data bubble (and this comes from a "data scientist")
"Data science" is a bizarre job description to me. Some companies' data science teams are doing work that could be done in Excel, and others are doing sophisticated machine learning research. In many companies, though, it's a watered-down version of the (extinct, sadly) R&D job description. I've heard people opine that 95% of "data scientists" have never implemented anything more sophisticated than an SGD regression, and it wouldn't surprise me.
Granted, you can get some neat insights and visualizations with off-the-shelf tools, but I still consider it important to know how the algorithms actually work and what the basic assumptions are. For example, you can usually use linear regression for 2-class classification problems (as opposed to the mathematically more correct logistic model, since probabilities are in [0, 1] and linear models diverge at the boundary) get a reasonable predictive model, but it's worth knowing when (and why) that short-cut breaks down.
In software, you hear "data scientist syndrome" used to describe people who have a lot of short job tenures because they (a) leave if there isn't interesting work for them, and (b) tend to be the first laid off, not because they're bad but because R&D is first to bleed when things go bad. The software industry forces you to choose between short-term job security (people doing interesting work are most exposed to organizational changes) and long-term career health (people not doing interesting work turn into dinosaurs). It stands to reason that the selection process for organizational credibility (favoring tenure) would be something other than deep knowledge of data science (which requires a stream of interesting work, and the behavioral correlates such as job volatility).
It's sad that the typical corporate culture of reliable mediocrity, at the expense of excellence, has also learned to use the vocabulary of "Big Data". Reality: truly Big Data (> 10 TB) is a huge pain in the ass. It's what you have to deal with when there is too much noise, or the model's essential complexity is too high, to get a good model out of "small data".
Logistic regression can be implemented as linear regression predicting the input to the logistic function.
If most of your data are in the p ∈ (0.2, 0.8) range, then the logistic curve is approximately linear anyway. Where linear regression fails you on a logistic problem is if you have a lot of data where the linear model will give p's outside of (0, 1).
While the linear model's mathematical foundations are "incorrect"-- more precisely, it's not the maximum-likelihood model and may have a conditional likelihood of zero, that is, being impossible as "the correct" model on the data-- the truth is that, at 100k+ features, getting "the right" model is effectively impossible with most real-world data sets and you have to use simplifying assumptions and techniques (e.g. regularization, early stopping) anyway. Linear regression doesn't usually do as well as logistic regression, but sometimes it can, and most people are going to try both.
Yes, solving for your coefficents based on inverting the matrix can be O(n^3) (although wikipedia tells me there's ways of doing it that are O(n^2.373), but at the end of the day, it's slow for large n) but gradient descent doesn't have this problem, as gradient descent is O(n) with regards to your training set.
And if you're implementing gradient descent to build a classifier, you can equivalently do logistic regression by just changing your cost function (that is, you'll be solving a linear regression to give you f(x) such that S(f(x)) is the probability you care about, with S(x) being the sigmoid function).
So linear regression can perhaps be easier than logistic regression, but only when n is so small that you're solving your linear regression by inverting the matrix. Once you've switched to gradient descent (or anything that can solve for solutions to arbitrary cost functions) there is no difference.
I'd have to work this out on paper. I think that's essentially right, but most of your CPU cost is going to be in evaluating that function (or, more accurately, its derivative) and not all functions are, in terms of computational cost, equal. I think the logistic model might take longer to train, but I'd have to work it out to be sure.
You're right that the code isn't much different.
So linear regression can perhaps be easier than logistic regression, but only when n is so small that you're solving your linear regression by inverting the matrix.
The matrix grows with the number of features, so if you have 1 billion observations but only 100 features (dimensions) you can do that.
I suggest watching: https://class.coursera.org/ml-003/lecture/36 and you can jump to 7 minutes if you don't want to watch the whole thing.
The only difference is instead of merely doing the matrix multiplication of your observation times your coefficients (the scalar product of θ and x) you have to take the output and use it as the input to the logistic function. The matrix multiplication will grow in complexity with the number of non-zero features (either non-zero coefficients in your model or non-zero features in your observation, presumably the latter) and the evaluation of the sigmoid function will be constant time, and on many architectures will be a very small number of instructions.
If you're operating in very small numbers of dimensions what you're doing may still be interesting but I'm not sure it qualifies as "Big Data," regardless of the number of available observations. After all, if you're producing something with actual utility, you'll probably converge sooner rather than later...
It is starting to smell like a marketing term, where "Big Data" is used to get executives with purchasing authority to be interested in no-longer-new tech by putting a shiny wrapper on it, "Data Science" is doing something similar with job titles - perhaps in an effort to get those who know better to buy into the hype.
Still, the article doesn't have to be objective and you might be right but a little bit negative.
It's the Knewtons of the world that have to spend millions on PR. Actually, Knewton is solidly OK at hiring good people but they usually leave within a year because of the management.
Their management is known throughout the industry to be unethical. Some recruiters refuse to work with them, and you'll hear some amazing stories if you hang around New York. They've existed for almost 6 years, pivoted constantly and recklessly, and delivered next to nothing. They're great at using the founder's family connections to raise money and get partnerships, but they treat engineers like commodities and have extreme architectural instability. There's also one high-profile case of the execs spending months trying to ruin the reputation of someone after he left.
Ultimately, their true mission is fire teachers. That's a really terrible mission, if you ask me. Obviously, the engineers and data scientists (Knewton's management is scum, but they have good people at lower levels) are being told a different story about "democratizing education", but their real mission is to make teachers obsolete while making a few ed-pub incumbents (like Pearson) very wealthy.
They're really damaging the reputation of the ed-tech space. It's a shame, because there are some really good companies trying to advance the field, and they're struggling to raise money due to the behavior (and declining reputation) of a large no-longer-startup that has nothing to do with them.
0: http://berglondon.com/wp-content/uploads/2010/07/trough.png