How UMAP Works
umap-learn.readthedocs.io
umap-learn.readthedocs.io
https://pair-code.github.io/understanding-umap/
UMAP is a really useful piece in the modern data science toolkit, and despite its power it's surprisingly simple and elegant. But as with all dimensionality reduction techniques, there's a lot of ways to misread the results. High dimensional data behaves very counterintuitively, and any reduction in dimensionality fundamentally distorts the original data in some way. Understanding the fundamental concepts behind UMAP and exploring how it works is the best way to develop an intuition of what the technique can and can't tell you about your data.
Also, is there some rule for choosing a) the number of dimensions the unreduced dataset should be mapped on to, b) the number of neighbors?
I assume the default parameters would work for most tasks.
It doesn’t really suit inference/prediction because you can’t really add new data without influencing the embedding values of the other data. It’s not like PCA where you can learn a projection once and then map new data points to the same embedding space.