Practical advice for analysis of large, complex data sets (2016)
unofficialgoogledatascience.com
unofficialgoogledatascience.com
Some people are coming from the Kaggle / Text book/ Statistics perspective - the data set is ready for analysis so I think we should use analytical tool x or y to give us some information about the conclusions that we are likely to be trying to demonstrate from it - or the models that we will build.
Other people are looking from a Data Science pov. The data is made of bits and is likely to have pathological features due to it's history and the biases of the different collection and assembly methods. The task is to find these and mitigate them - at that point we have a decent data set and can move to analysis and modelling.
For any non-trivial data set, the efficiency and internal consistency of $tooling is critical. Throwing stupendous hardware resources at a problem because the $tooling is unstable rubbish is the wrong approach.
For any non-trivial data set, the data must have a cohesive reason to exist. In other words... Big Data Garbage In, Big Data Garbage Out. Telemetry is often BDGI.
If I had one piece of advice to give on the subject it would be PCA the crap out of everything and understand what the top components are doing. 9 times out of ten they will turn out to be very significant confounders and in the cases where they are not they will be confirming the structure of your data set in very useful and significant ways.
The coefficients will most likely be numerous and noisy, making sense of them will be impossible. You need to have an hypothesis first of what the principal components might be, and only then compare your idea with the PCA coefficients to see if that works.
Until then I don't think that PCA really tells you what you need to know - which is "is this data set what I think it is " and "is the information that I need to extract actually in here"?
There are at least four issues with this advice (wrt doing data analysis):
1. How do you link your PCA components to the original data? Let's say you are tasked to find the main drivers of sales on a given city. You run PCA on the data and find two main components on the dataset. What do you do next? How do you make this information actionable?
2. How do you treat categorical variables? There are PCA methods for dealing with categorical variables but by the time you apply these methods plus the issues in 1) your data has lost all actionable meaning.
3. PCA is _very_ difficult to explain to business stakeholders. The more difficulty business stakeholders have to understand the analysis, the less they will use it.
4. Data-driven business stakeholders will favour clarity and simplicity over sophistication (somewhat linked to 3)
> you are tasked to find the main drivers of sales on a given city. You run PCA on the data and find two main components on the dataset
Obviously it depends what comes out. But in all likelihood you will see some significant clusterings / divisions in PC1 and PC2, so you will try to interpret what properties of the points are driving those. You can do it in a data driven way (what are the significant coefficients in principal component vectors) or you can often do it in an exploratory way ... are they related to geography, are they related to age demographics, are they seasonal ... you color the data points by different possible explanatory variables to see what group things together. And you will very likely see things jump out (eg: you could find that the main reason a particular month was down in sales was due to a technical problem with the web site and you'll want to put that aside, because it doesn't have any predictive value).
https://www.efavdb.com/unsupervised-feature-selection-in-pyt...
There will be many problems where you can't easily get a "PCA that looks sane", but you can definitely get a UMAP/Ivis that looks sane. I guess I see where you're coming from in regards to like massive datasets and making sure you're not doing silly things (e.g. misreading an ordinal as continuous) - but outside of this I think PCA is antiquated.
The antiquation of PCA yet it's still ubiquitous usage is most likely one of the many factors holding back fields like bioinformatics dramatically. Please switch over to modern techniques. Your data will thank you!
Thanks for the suggestion!
(1) Slice your data, and check metrics in each slice
(2) Check for consistency over time, which is actually a special case of (1)
This is a type of as hoc bootstrapping. If your estimator is stable over various subgroups then it implies the estimator variance is low and you can be more confident in an observed effect.
But man, if there was one section that makes this a crucial read, it’s this:
> Be both skeptic and champion
Nothing makes me tune out someone’s presentation/proposal faster than hearing only upside. Everything has trade-offs and no analysis has perfect information. If you can’t acknowledge this, I can’t trust you.
Well, mostly about time series in general, not anomaly detection specifically. For more general advice about best practices for this type of ML, you might like “rules of ML”: https://developers.google.com/machine-learning/guides/rules-...