Algebraic Topology for Data Scientists
arxiv.org
arxiv.org
The idea is that you take your data points and expand them in space by putting circles around them. If there are features that are persistent over large variations of the sizes of the circles (or balls in higher-dimensional space), then those persistent features are said to be important reflections of the structure of your data.
The persistent features are characterized by homology, which is a tool that measures the basic topological shape of your data. Homology is actually extremely easy to calculate and doesn't require much advanced math. In fact, for data analysis purposes, you can focus purely on combinatorial aspects, in which case no advanced math is required.
In my opinion, it's a pretty idea and seems logical, but in practice it really seems to be limited in a few ways in terms of practical applications. For one, it often is weaker than more direct methods such as good experimental design and other cluster data methods that can be more directly used to tease out causal relationships.
For another, it seems most useful in the domain of comparing data changing over time or some other variable where the fundamental topological structure change in high-dimensional space in tidy ways. This isn't the case with a lot of real-world data.
In practial terms, you can think of this in the following way... there is only so much you can get out of data. And often, what you do need is not coarse topological features, which rarely have intrinsic meaning.
So my overall opinion of it is that while it might have some limited applications in highly specific fields, it will never become a general method that could be used in 99.9% of data science applications. (To be FAIR though, I also believe that a LARGE proportion of data science is also kind of useless too...data science has its fair share of snake oil.)
What can also be helpful is looking at the t-SNE creator's webpage[1], where you'll see examples given. Look at MNIST and pay close attention to the clusters and items in them. Are clusters that you'd expect to be near one another actually? Do numbers that are similar smoothly transition into one another? The top 4/9/7 structure is good but clearly 7 should transition into 1's (if we manually look at data you'll pick up on this in no time) We can say the same about some other numbers and structures. Of course, we've reduced a 784 dimensional object into a 2D representation, so we are losing a lot, but the most important question is what we are losing and if it matters to us. This question is surprisingly frequently absent despite being necessary.
One of the best ways to understand limitations is, unfortunately, to look at works that build upon the previous work. There are two main reasons for this: 1) we learn more with time, and 2) the competitive nature of publishing actively incentivizes authors to not be explicit about the limitations of their work as doing so often significantly jeopardizes the likelihood of the work being published as reviewers (have historically) weaponized these sections against the authors[3] (this is exceptionally problematic in ML (where I work, hence the frustration) and is unfortunately growing more problematic, and rapidly (hence the evangelization)). UMAP does an okay job though and I want to quote from the actual paper:
> In particular the dimensions of the UMAP embedding space have no specific meaning, unlike PCA where the dimensions are the directions of greatest variance in the source data.
We can also look at DensMAP[4] where we see they specifically target increasing density preservation. A critical aspect for data and local structure!
We can of course attempt to dive deep and understand all the math, but this is cumbersome and an unrealistic expectation as we all have many things. But the best thing I can say is to always be aware of the assumptions of the model[5]. If there is one thing you _should not_ be lazy about, it is understanding the assumptions. Remember: ALL MODELS ARE WRONG. But wrong doesn't mean useless! Just remember that there are nuances to these things and that unfortunately they are often critical, but our damned minds encourage us to be lazy. But you can trick yourself into realizing including the nuance is "lazier," if you account for future rewards/costs instead of just immediate (a bit meta ;)
I hope this wasn't too rambling... and did provide you with some answers you were looking for.
[0] In the words of Poincare: mathematics is not the study of data or objects, but rather the relationships between the data and objects. The distinction may seem like nothing, but it is worth mentioning.
[1] https://twitter.com/lpachter/status/1431325969411821572
[2] https://lvdmaaten.github.io/tsne/
[3] If you are a reviewer, stop this bullshit. It is anti-scientific. Your job is not to validate papers, you can't do that. You also can't determine novelty, this concept itself is meaningless and to provide meaning needs substantial nuance (99% of the time a lack of novelty claim contains more bullshit than this statistic). You can only invalidate a work or give it in-determinant status. Papers are the way scientists communicate with one another. The purpose of a reviewer is to check for serious errors, check for readability (do not reject for this if it can be resolved! SERIOUSLY WTF), and provide an initial round of questions from a third party point of view that the authors may have not considered. Nothing else. Your job is _NOT_ to reject a work, your job is to _IMPROVE_ a work and help your peers maximize their ability to communicate. We have a serious alignment problem and for the love of god just stop this shit. Karen, I know you're Reviewer #2. Get a real hobby and stop holding back science.
[4] https://www.biorxiv.org/content/10.1101/2020.05.12.077776
[5] Model is a much broader term than many people understand. Metrics, evaluation methods, datasets, and so on are also models. These are often forgotten about, and to serious detriment. All metrics are wrong, and you cannot just compare two things on a singular metric without additional context. Math is a language, and like all languages it must be interpreted. It is compressed information and ignoring that compression will burn you, others, and your community. Similarly, datasets are only proxies of real world data and are forced upon us due to the damned laws of physics that prevent us from collecting the necessary infinite number of samples as well as the full diversity of that true data (which is ever changing). As "just a data engineer" (no need for the just ;) it is quite important that you always keep this in the back of your mind. Especially when utilizing the works that my peers in AI/ML develop. There's a lot of snake oil going around and everyone has significant pressure to add it to their works.
If you're wondering how Algebraic Topology can be used to solve real world problems, you will still wonder 300 pages later.
I think one might be better of just reading about uMAP, Mapper and persistent homology directly . At least the last two are very simple and don’t require advanced maths.
https://people.maths.ox.ac.uk/nanda/cat/
Unlike what people commented elsewhere, TDA has growing application in computer graphics, 3D reconstructions and computer vision.
PS: I am told the new version of this course is narrated/illustrated by Robert Ghrist. He's famous for his foundational calculus course
Instructions to add:
- Click other formats
- Download source
- Add this to the includes
``` \usepackage[bookmarks,linktocpage=true]{hyperref}
\hypersetup{
colorlinks,
linktoc=all,
linkcolor={blue},
}\makeindex ```
- recompile
Presto, you got the same document except now we can click the links in the table of contents, the citations, and there's a bookmark section on the side that allows us to navigate the document instead of just scrolling.
I'm not sure what graduate math programs don't have Algebraic Topology. It's pretty fundamental.
The problem is this -- topological spaces on their own have a lot less structure than say a vector space or a metric space. Most of the real world data that ML applications deal with today have more structure and are thus handled using vector space methods or metric space methods or manifold methods. It is also the right thing to do. When you have structure you shoukd use it.
TDA would be useful for data sets where the members have no clear analogue of a distance, or the distances cannot be trusted, or where vector embeddings do not make sense.
constructive type theory and TDA are very different things even if they do both have `topo` in their names.
This is a bit unfair since Bayesian techniques are useful for when you want to reason about limited data.
But as soon as you have a fair sized dataset, frequentist techniques are typically computationally simpler and faster (no need for MCMC) and scale much better.
A good introduction to some additional problems with frequentist methods vs Bayesian and likelihoodist methods is this: https://gandenberger.org/2014/08/26/intro-to-statistical-met...
An interesting book on adapting frequentist methods to create confidence distributions that can better express uncertainty and can optionally incorporate prior information using likelihood functions is this: https://www.cambridge.org/core/books/confidence-likelihood-p...
For everything else there’s so much error. Being to quantify uncertainty is great — it’s a signal we need to collect more and better data. But so often we have to move ahead with uncertain data.
Interestingly, in business, taking action (even if wrong) produces outcomes that are much better signals to learn from than having statistically rigorous analyses, so many times there’s a bias for action rather than obsession over analysis.
But of course in some fields being wrong is costly (like clinical trials) so I can see UQ being more useful and prominent there.
Although there may be no data collected from the time something similar happened before in history, experts can reason through the situation to guesstimate the direction and magnitude of the effect in qualitative terms.
Bayesian formulations are very handy in such situations.
>I’ve found that in many applications, the difference between a frequentist analysis and a Bayesian one is unlikely to make a difference in the decision making (even with UQ).
In that case you may find the following interesting
https://en.wikipedia.org/wiki/Lindley%27s_paradox
"Lindley's paradox is a counterintuitive situation in statistics in which the Bayesian and frequentist approaches to a hypothesis testing problem give different results for certain choices of the prior distribution."
How likely is Lindley's Paradox likely to show up in practice ? well there is Bayes for that (tongue firmly in cheek).
I think it’s definitely possible that Bayesian and Frequentist approaches give different conclusions but in practice it doesn’t alter the final decision. Analyses guide decision making but in the end decisions are made on consensus, narrative and intuition. Statistics is only the handmaiden rather than the arbiter.
Indeed, but that does not make it right or rational. Bayesian method helps keep things rational. This is pertinent because human brains are terrible at conditional probabilities.
One can always argue that data analysis is usually just window dressing and decision making is mostly political and social. Empirically you would be mostly right if you take that position. One cannot argue against that factual observation.
The more interesting question is, if the decision makers aspire to be rational, which method should they use. I have used frequentist and Bayesian methods both. I made the choice on the basis of the question that needed answering.
For example, when we needed to monitor (and alert on) a time varying probability of error (under time varying sample sizes) -- Bayesian method was a more natural fit than say confidence intervals or hypothesis tests. Bayesian methods directly address the question "What is the probability that error probability is below the threshold now, considering domain expert's opinion about how often it goes below the threshold and how the data has looked in the recent past?"
I agree with you on Bayesian methods keeping things rational (consistent within a probabilistic framework).
I would say there are different kinds of rationality however: the 2 that I'm most interested in are epistemic rationality (not being wrong) and instrumental rationality (what works), and in the domain of business (but perhaps not other domains like science and math), we optimize for the latter. This is because not getting analyses wrong (epistemic) is actually less useful than getting workable results (instrumental) even if the analyses are wrong. In fact, some folks at lesswrong tried their hand at doing a startup, applying all the principles of epistemic rationality and avoiding bias, and it did not work out. Business is less about having the right mental model but doing what works. This article expands on this point [1]
The issue is in ill-defined (not just stochastic in a parametric uncertainty sense, but actually ill-defined) domains like business, the map (statistical models) is not the territory (real world) -- it's a very rough proxy for it. Even the expert opinions that Bayesian methods embed as priors -- many of those are subjective priors which are not 100% rational. Not to be cliched but to recycle an old John Tukey saying: "An approximate answer to the right question is worth a great deal more than a precise answer to the wrong question." Frequentist methods are often good enough for discovering the terrain approximately, and in business, there's more value in discovering terrain than in getting the analysis exactly right.
(that said, in these settings Bayesian methods are equally as good too, though their marginal value over frequentist is often not appreciable. One exception might be multilevel regression analysis where you're stacking models.)
[1] https://commoncog.com/putting-mental-models-to-practice-part...
Of course ! Like big bang it needs one initial allowable 'miracle' and does not let irrationality creep in through other back doors.
As I mentioned earlier, I choose the formulation that suits the question that needs answering.
Not sure about the 'what works' vs 'analytic correctness'. How would even one know that something works or have a hunch about what may succeed if they have no mental model to base it upon. Often that is implicit and not sharp enough to be quantitative. Bayesian formulation helps in making some of those implicit assumptions explicit.
Other than that I think we mostly agree. For example both the formulations have a notion of a completely defined sample space, the universe of all possible outcomes. That works in a game of gambling. In business often you do not know this set.
Anyhow, nice talking to you. I enjoyed the conversation.
This is true of everything, of course. Everyone knows that “practical applications” is just a flimsy excuse to get funding for your cool math project.
I’ve found that the main benefit of learning math is not the practical applications of the theorems and algorithms themselves, but just the practice you get from using this approach to thinking. This also frees you up to learn the math you find interesting and not what you think might be immediately useful in your day job.
As some other commenters mentioned, though, it was difficult to try to skim this book to find high level applications. For anybody else that's interested, I think there are some applications in chapters 5-7.
The results are spectacular, but the application set is indeed very limited.
You can extend some of the thinking in this paper to cover action spaces for deep RL models to really have some fun!
Could you expand on that?
Maybe I'm missing something but I browsed the PDF and am having trouble finding anything that would be deployed to a production system or produced as part of an analysis in a data science workflow.
I think there some very niche applications in manifold learning (e.g. UMAP) that is somewhat useful, but in my decade in this space, people have tried to find topology applications but ultimately we end up going back to the basic tools of statistics and machine learning.
Otter, N., Porter, M.A., Tillmann, U. et al. A roadmap for the computation of persistent homology. EPJ Data Sci. 6, 17 (2017). https://doi.org/10.1140/epjds/s13688-017-0109-5
There are classes of answers that I don’t see how you could find them without resorting to something like TDA. Stuff like characterizing the Betti numbers of neuronal circuits to profile signal routing redundancy. The economics of applying TDA at scale (at least at one that does not require eg nation-state level compute) don’t work well atm I think: small problems will quickly kill the beefiest cpu/ram combo you can find if you try to do basic persistent homology on them, so either you scale your compute like crazy for results that probably won’t give you jackpot margins wrt competitors, or just do whatever is cheap and quick and good enough.
A remark: I think tda is in an era not too unlike deep learning after the proliferation of backprop and little deep nets, but before backprop-on-gpus. People said all kinds of things about how deep learning was a gimmick, was too impractical to apply to real world problems, etc. I remain curious either way.
There are many applications of computational geometry today (which is not in data science but more engineering). If we can we can work with topological objects I can see we might be able to find areas where we can derive value.
NN is a little different because it has always had a very strong why (it’s a very flexible, highly parameterized nonlinear regression model — a fitting function for everything) but was hampered by the how for many years.
Topology’s whys are a bit less universal but I can see some very useful future applications.
--
0: https://en.wikipedia.org/wiki/Topological_data_analysis.
1: https://en.wikipedia.org/w/index.php?title=3LA&redirect=no