What do numbers look like?
johnhw.github.io
johnhw.github.io
https://twitter.com/hippopedoid/status/1318917878364672001?l...
I say that the beauty (or value or worthiness) of the pictures of the Mandelbrot set comes from the transformations we apply to uninteresting complex numbers.
Similarly, the beauty of the pictures in the article may come from some hidden properties of the underlying prime numbers, or from the transformations themselves, and I don't think that either case would be better than the other.
I said this in reply to a comment that seemingly stated "what a pity, these images are 'merely' due to the transformations". I was objecting to the tone of disappointment that I read in that message.
So no, it's not like the Mandelbrot set. It's more like if you wrote a script to visualize the Mandelbrot set and created a bug that made part of the visual you created look like there was an interesting structure by accident, then shared your pictures with a wide audience going "look at this interesting structure I found in the mandelbrot set!" and then someone replies on Twitter with "you have bug in line 124 and when I correct it the structure disappears". Which is why it's a pity.
This is both an interesting/fun visualization exercise and a cautionary story. Apparently UMAP has a tendency to render blobs as rings or loops!
Example of what could cause a swirly chain in UMAP: if A related to B and B relates to C but A does not relate to C and so on. IMO that's a valid structure to visualize as a swirly chain. If re-run multiple times, of course you will get that chain in different locations and so on. But it is interesting that it is there.
But visualisations can always deceive you into seeing something that's not there, e.g. correlation vs causation.
However, all numbers from 1 to 1 000 000 forming a distinct clusters when being mapped to 2 dimensions with UMAP… I don't know. It might be nothing (like a representation of something trivial, like an observation, that multiples of 100003 are less common than multiples of 3 in the set of first 1 000 000 integers), and all these clusters may just disappear (converge) as we go closer to infinity. But there's definitely something a bit eerie about the possibility of it not being "nothing". Normally, you wouldn't expect any patterns to form like that.
Or, well, it may be more that a "nothing" but less than "interesting" for a mathematician — maybe there actually is some pattern that becomes more visible in this visualization, but it's already well-known among number theorists. I just have no idea, that's why I'm asking. It's just weird to see any clustering at all here.
(And, yeah, BTW, there's no such thing as "correlation vs causation" in number theory.)
Why should 2 and 4 both be [1 0 ...] instead of [2 0 ...], etc?
In many cases using all 2s or all -10s would produce the exact same result in theory, but with more work by the optimizing algorithm, possibly with adverse results as described above.
It's easy to reason about mathematically, too. If the input vector is all 1s and 0s, it's easy to read off the result of multiplying that vector with another vector. Norms of binary vectors and dot products between binary vectors are super-easy, and have a nice correspondence with counting the appearances of elements.
It also corresponds to an array of Boolean values which is conceptually appealing, because that's basically how it's constructed.
With a couple of specific exceptions, there's little reason not to use all 1s.
However this encoding technique is a bit more like choosing 1 and 0 to represent Boolean values in C: it's convenient, it's easy to reason about, it's mathematically simpler than any other option, and there's no compelling reason to choose anything else anyway.
For binary variables in particular, you sometimes see people using -1 and 1 instead of 0 and 1, to get symmetry around 0. I think this was mostly only used for encoding the labels/outputs of SVM models, where it's mathematically appealing as representing two sides of a hyperplane.
There are a handful of other schemes for encoding "categorical" or "nominal" data of this kind, used in certain statistical applications such as the design and analysis of experiments. These encoding schemes are called contrasts in the stats literature, because they emphasize the differences (the "constrasts") between categories.
https://gist.github.com/johnhw/dfc7b8b8519aac530ac97da226c17...
2 => [1 0 0 0...]
3 => [0 1 0 0...]
4 => [1 0 0 0...]
And map that high dimensional space back down to two dimensions (using some technique I haven't dug into yet). Colors are assigned by some scheme, later images help to illustrate how the particular clusterings happen like one where primes are rendered in white.- Richard Feynman