427 karma · joined January 19, 2013
"Put a box around any region of space (space-time, really, but I’m going to drop time throughout this essay for ease of visualization, as physicists often do). The holographic principle asserts that no matter what’s going on inside — from gas molecules pinging around to black holes colliding — you can decipher the entire contents of the box just by repeatedly measuring points on the surface."
Well, if we're bounding a region of space-time then, in a purely classical universe governed by deterministic ODEs or PDEs, the statement reduces to a triviality: having information about the boundary amounts to knowing all boundary conditions. Of course I understand that this isn't really the statement of the holographic principle, but the article's formulation is rather underwhelming.
To expand a bit: It's clear from the writing that the formatting is an intentional stylistic choice. My point wasn't about aesthetic preference. I meant precisely what I wrote. Language evolves and orthography evolves. We've been capitalizing less and less for centuries now. But (until very recently) it was a universal rule to capitalize the first letters of sentences. I believe this is partly because (again, until very recently) people read in large quantities and, to a fluent reader, sentence capitalization serves an important purpose: it helps the eye recognize where one sentence ends and the next begins. Periods alone can be easy to miss, or confuse for commas, when the eye is moving quickly.
Anyway, this is all a very minor point. You have a cool site and I enjoyed your article. I hope you keep writing.
Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.
So it's more useful to look at serious injury statistics as a proxy. For those, Waymo says that there is a significant reduction.
https://math.stackexchange.com/questions/2140493/counterintu...
If your model is different (y = Ax + b + e where the error e is not normal) then it could be that a different penalty function is more appropriate. In the real world, this is actually very often the case, because the error can be long-tailed. The power of 1 is sometimes used. Also common is the Huber loss function, which coincides with e^2 (residual squared) for small values of e but is linear for larger values. This has the effect of putting less weight on outliers: it is "robust".
In principle, if you knew the distribution of the noise/error, you could calculate the correct penalty function to give the maximum likelihood estimate. More on this (with explicit formulas) in Boyd and Vandenberghe's "Convex Optimization" (freely available on their website), pp. 352-353.
Edit: I remembered another reason. Least squares fits are also popular because they are what is required for ANOVA, a very old and still-popular methodology for breaking down variance into components (this is what people refer to when they say things like "75% of the variance is due to <predictor>"). ANOVA is fundamentally based on the pythagorean theorem, which lives in Euclidean geometry and requires squares. So as I understand it ANOVA demands that you do a least-squares fit, even if it's not really appropriate for the situation.
In any case, they are a bit more advanced, and out of scope for the undergraduate course I linked to.
This lecture by Dennis Freeman from MIT 6.003 "Signals and Systems" gives an intuitive explanation of the connections between the four popular Fourier transforms (the Fourier transform, the discrete Fourier transform, the Fourier series, and the discrete-time Fourier transform):
https://ocw.mit.edu/courses/6-003-signals-and-systems-fall-2...
How do you compute the fractional FT? My guess is by interpolating the DFT matrix (via matrix logarithm & exponential) -- is that right, or do you use some other method?
I also once made my own variant of this (just like gregfjohnson's idea): A "lucky minesweeper" where luck can be toggled on/off at any point during the game: https://github.com/yshklarov/minesweeper
"From the point of view of software engineering, the rapid spread of C represented a great leap backward. It revealed that the community at large had hardly grasped the true meaning of the term “high-level language” which became an ill-understood buzzword."
Source: Niklaus Wirth, A Brief History of Software Engineering, 2008 (https://people.inf.ethz.ch/wirth/Miscellaneous/IEEE-Annals.p...)
In my view, having a single lingua franca is nice. It better facilitates knowledge transfer. I wouldn't want to see a fracturing where each area of knowledge (or, say, every specialization/application programming) is best treated in a distinct linguistic community. That would be bad for everyone.
What, specifially, do you find awful here?
It's true that there is no intrinsic meaning to the scale, but you must specify at least a relative scale -- how you want to compare (or weigh) different units -- before you can meaningfully cluster the data. Clustering can only work on dimensionless data.
As for your last two points, I believe I agree! It seems that in the counterexample you give for consistency, some notion of scale-invariance is implicitly assumed -- perhaps this connection plays some role in the theorem's proof (which I haven't read).
This reminds me a bit of Arrow's impossibility theorem for voting, which similarly has questionable premises.
This is quite an uncharitable perspective!
Outside of a classroom setting, the way you learn from a textbook without external feedback is by engaging more actively with the material.
Treat each statement in the main text as an informal exercise. Each time you come across a proposition -- whether it's a formal theorem statement or a claim in the body of the exposition -- try proving or otherwise justifying it to yourself before reading on.
Take a look at Theorems 2.3.1 and 2.3.2 -- they are very similar. Once you have absorbed the proof of 2.3.1, you can attempt 2.3.2 on your own. If you can't finish the proof, you can read a couple of sentences from the included proof for "hints"... or, if you do finish a proof, you can compare it to the proof in the text.
If you read actively enough, you can learn the material quite well without doing any problems. Many people will claim that you need to do formal problems in order to learn math, but this is untrue. Many math textbooks at the higher level don't include formal exercises or problems at all, and people learn from them just fine.
Admittedly, reading mathematics is a skill in its own right, and you shouldn't expect it to come easily right away. Of course, the best thing is to have a one-on-one teacher, but few of us are so lucky.
That said, while I appreciate and admire free textbooks published online, I think the exposition would be much improved if the author had a better sense of who he was writing for (the most common writing advice...).
And I take issue with the view that the sample space is where the random phenomenon lives, as it were. In my experience, it's more common to use the random variable itself to model the (observable aspect of) the random phenomenon, and for the sample space to be either a hidden (i.e., more abstract) aspect of the phenomenon or else a purely abstract formalism introduced only for ease of mathematical computation.
It would be helpful also to see some more context, especially historical (who introduced the concept of a sample space, and for what purpose?).