I'm not sure what momentum and dropout are, but I agree with Eleizer, without these things (which I didn't use) local minima are a problem.
[0] https://www.cs.cmu.edu/~venkatg/teaching/CStheory-infoage/ch...
however, being able to mix and match properties is really what most good plotting and visualization is all about. having done a bunch of ray tracing, my intuition around lighting and light is much better. I am not even good 'anechdata' so take that for what its worth, but I found visualization to be much more intuitive than reading e&m textbooks/lectures. I'm not sure if knowing a phenomena as bottom up or top down is really a guarantee (nor am I suggesting that symbolic reasoning is bottom up or down), but seeing something is just so efficient for some people. Like anything powerful, it just needs to be used judiciously and with asterisks.
For example, learning geography by flat map projections only. No matter what projection you use there is a trade off, and you end up instilling both the pro and the con of that trade off as intuition.
I found https://blogs.msdn.microsoft.com/ericlippert/2005/05/13/high... with a quick search.
the volume of an n-ball peaks at n=4 dimensions and quickly drops to zero around n=20. cf. https://news.ycombinator.com/item?id=3995930
The comment refers to lebesgue measure (I don't even what), but I'd intuitively and ignorantly assume we count all faces of all n-1 balls (recursively) whereas the Volumes overlap and so the total (in lebesgue ...space?) is less than the sum of it's parts (in euclidean space) - how far off am I? (will delete if too far)
A video from LEO or a rotating map projection provides very different intuition than a single static map. https://www.youtube.com/watch?v=EPyl1LgNtoQ
Very high zoom levels also work out nicely.
There are several key elements to effective visual thinking. The primary importance is to keep it grounded in proofs and theorems, so you know exactly what are your limitations. Often you can use a geometric argument on top of a few theorems and you get a very strong result intuitively, and then use this intuition with a tiny amount of algebra to prove it (which might take you forever to arrive from a purely algebraic perspective). Another key is that there are several ways of visualizing things. You can almost always transform a problem into an equivalent one that is easy to visualize (just need a little bit of care with the transformation, etc).
---
For example, you can show functions form a vector space, visualizing some interesting algebraic properties about them, even if it constitutes an infinite-dimensional space.
You can show several operators (such as d/dx) are linear, you can give it a norm, internal product, etc. This trick lets you use visual tools (and linear algebra tools) with arbitrary functions. You can visualize projection of a function into a subspace, or into some non-canonical basis -- yielding useful applications -- such as Fourier analysis.
Fourier analysis itself is a fertile ground for visual thinking. You'll be finding trivial arguments for seemingly difficult decisions such as "Does this linear system have a bounded output for any bounded input?". There isn't one right way of thinking about anything.
---
On the other hand, it can't be stressed enough the importance of keeping track of formal assumptions, axioms, definitions, theorems to construct valid, correct proofs. That way you minimize the risk of fooling yourself, and can safely use your intuition.
This 3B1B video exemplifies many of those elements:
well, the map is flat more or less at closer zoom levels, so the general problem seems to be purely about lossy compression.
Also anything related to topology, which is important when you are looking at decision boundaries, becomes counterintuitive in high dimensions, because so many things can be adjacent at the same time.
If true, you're very correct that lower-dimensional intuition does not transfer into higher-dimensional spaces: my intuition tells me that a Gaussian distribution drops off as you fall away from the mean, and it's quite easy for me to imagine that in 2 dimensions, 3 dimensions (e.g. by imagining a mound on a plane) and 4 dimensions (e.g. a cloud in 3-space with increased density around the mean).
Is my intuition wrong in any of those cases? If so, why? If not, how many dimensions do we need before it becomes wrong?
This is somewhat related to another 'curse of dimensionality' observation, which is that the volume of a hyperball / volume of hyperspace tends towards zero as dimensions grow -- there's just a lot more volume that's in some sense 'far' from the center.
Say you have an N-dimensional gaussian where each dimension has mean 0 and standard deviation 1. Define the center as the N-dimensional cube whose edges go from -3 to +3 in each dimension. A normally distributed value is within 3 standard deviations of the mean with probability 0.9973, so the probability that an N-dimensional point being in the center is 0.9973^N. With N=4 that's 0.989 which matches your intuition, but at N=1000 it's 0.067 and at N=10000 it's 1.81e-12.
It starts to fail really badly when dimension grows.
Two simple examples:
1) Consider 3 dimensional unit sphere centered at origin and unit cube centered at origin. Cube is clearly completely inside the sphere. Now generalize to n-dimensions. Hyperdimensional volume of hypercube with side length 1 moves almost completely outside the n-sphere with radius 1 when n-grows.
2) Alternatively almost all volume of n-sphere is close to the surface.
These are all very counterintuitive, yet simple to check toy examples. When you start to integrate over more complex multidimensional function, things get weird really fast.
How does this go against intuition?
Intuition from 1/2/3d tells me that the volume of an N-ball is O(r^N), and indeed it is the case in higher dimensions. Therefore it’s easy to see that the difference between the volume of an N-ball of radius r and an N-ball of radius (r + epsilon) will grow exponentially with N.
Density is different from mass. Namely, mass is the integral of density. So your intuition is roughly correct for density, but you need to make it accord with a good intuition for mass.
Since getting the mass requires an integral, getting the mass over N-dimensional distributions requires integrating an N-dimensional region, which means N integrations for N dimensions. Each integration is, intuitively, a kind of sum. Integrating out many dimensions happens recursively; looped or recursive addition is multiplication. So on some level, to take the probability mass of a region in N-dimensional space, you need to "multiply" a density.
Since the total probability mass is fixed (1.0), adding more dimensions means you need to "multiply" the density by a larger number to get the mass, which means you need to divide the mass by a larger number to get the density, which means that despite the density peaking at the mean, the available density at any given point gets smaller as the dimensionality rises.
I don't see how the numbers are supposed to work out for large mammals, with each female having under a dozen offspring. To have a decent chance of a back-mutation, the typical member of the species would need one twelfth of their genome to be deleterious mutations.
Meanwhile, people are thinking about using CRISPR to correct the human genome, creating unusually happy, healthy people. The underlying thought is that the correct genome is best. But why do we think that the correct genome works at all?
Most of the population is in a shell at a distance from the correct genome, the number at the center is actually quite low. Given the combinatorics, with two to the millions of possible genomes, but populations in the millions, the number at the center, or even close, is actually zero. Maybe the correct genome codes for a sickly, miserable individual?
My current guess is that the evolution of large mammals with few offspring is constrained by genetic load considerations. It is not sufficient, (or even necessary) for the correct genome to be any good. There needs to be a big blob of mediocrity in genome space. The species exists as a shell of individuals on the edge of the blob of mediocrity. The blob needs to be huge, so that individuals whose genome is one twelfth mutations are still in the blob. Then there can be an equilibrium between back-mutations, taking offspring towards the interior of the blob and other mutations, taking offspring out of the blob and out of the gene pool.
This potentially solves the Fermi paradox. Can creatures such as humans actually exist in this universe? It is not enough for natural selection to discover a good genome. Natural selection has to discover a huge blob of mediocrity. Such blobs might be vastly rarer than we realize.
This potentially shits on the CRISPR master race. There might be nothing special about the interior of the huge blob of mediocrity.
But what happens if you step back from black-and-white thinking and ask about mutations with ambiguous effects. Which is the mutation and which is the correct genome? It becomes unclear.
An alternative perspective asks: how well separated are the local minima in fitness space? Perhaps the typical separation is as large as the gaps between species. Then each species has only its own local minimum, which defines its correct genome. Or perhaps fitness space is littered with local minima, such that a single species has genetically healthy individuals in several different minima plus other individuals, perhaps not quite so healthy, nearby.
http://worrydream.com/refs/Hamming-TheArtOfDoingScienceAndEn...
(I assume Bret Victor has permission to host the PDF on his website, he is far from an anonymous pirate)