176 karma · joined April 18, 2012
The tricky thing is that these insights wind up being kind of obvious from the first analysis. You will find things like "users who use the software more are more likely convert." Other times these types of analysis will confirm what you already know. The tricky thing is making sure you have the right tagging/events and place to make sure you're getting at the right level of detail to get something worthwhile. It's very a much a garbage in garbage out type of thing.
As far as the machine learning market goes, 90% of the projects require software engineering skills, the last 10% requires being able to go underneath the covers of linear algebra libraries, etc.
I just think the whole physics>cs degree for machine learning argument is not totally persuasive given my experience.
Sunshine
Blade Runner, gattaca, the matrix, alien etc.etc. basically just google "best sci fi movies of all time"
Motivation is ephemeral. Hard work, habit, consistently applying yourself regardless of your motivation, is key.
EDIT: Note that it's still EXCEPTIONALLY difficult to provide true, lasting value to a given organization with this stuff and takes years of experience (note I didn't mention deep knowledge and experience with the latest techniques aboard the hype train).
See http://www.theverge.com/2016/2/23/11097642/google-shuts-down....
The part I am having trouble with, with ~5 years of real world data science experience in industry, is the implicit assumption that we all need or want to become good at deep learning.
In my experience, most businesses, at best, are still struggling with overfitting logistic regression in Excel let alone implementing/integrating it with a production code base. And we all know that toy ML models that sit on laptops create ZERO value beyond fodder for Board presentations or moving the CMO's agenda forward.
The fact of the matter is that the vast majority of businesses, with respect to statistics/ML, aren’t doing super duper basic shit (like a random forest microservice that scores some sort of transaction) that might increase some metric 10%. This is due to lack of sophisticated analytics infrastructure/bureaucracy/ lack of talent/ being too scared of statistics. Ultimately when you’re rolling out a machine learning product internally (I’m not talking about Azure/ other aws-model-training-as-a-service type things), the hardest part isn’t: “We need to increase our accuracy by 2% by using Restricted Convolutionallly Recurrent Bayesian Machines!” The hardest part is convincing people you need to integrate a new process into a “production” workflow, and then maintaining that process.
For example, if your business is using any sort of machine learning model to govern a process, and you have some sort of idea of how you want this process to run throughout the day, you can build a PID system to change the model thresholds automatically to govern that plan.
An example might be any sort of retention/marketing strategy where you want to reach out to customers based on some factors X, and have a certain quota/capacity to do so during a day, and the expected number of customers with the best factors X can change throughout the day/week or is just erratic.
What was nice about working in scientific research back in undergrad in that whenever you got tired of coding the demon box, you could do some good physical labor in the lab.
Been there, done that.
This is the mantra of the BI system marketing machine that I've never actually seen work in practice.
But I'm glad you've seen BI enterprise systems succeed where the end users are happy and feel their toolset is flexible enough for the new daily challenges they encounter (sincerely).
I love Tableau as an exploration tool and think the UX and visualization capabilities are awesome. Just not the solution end users wanted. I've also rolled out a massive cloud based enterprise BI system (data warehousing, ETL, etc.) at a separate company. That wasn't the solution, either. Plotly was. But, just my two cents.