HNHacker News
TopNewBestAskShowJobs

yshklarov

427 karma · joined January 19, 2013

submissionscomments
yshklarov··on Gravity seems holographic. What does that mean for reality?
You are correct if you mean a 2-sphere, but a spherical shell in space-time would be a 3-sphere, which gives quite a lot more information. To make things a bit simpler we could consider a 2-sphere × an interval -- the boundary of a 4-dimensional cylinder -- which is to say, your spherical shell at all points in time from t_0 to t_1, and also the entire interior of the spatial region at both the initial time t_0 and the final time t_1.
yshklarov··on Gravity seems holographic. What does that mean for reality?
Indeed. A "violation of logic and geometry"? Reading the author's own description, it's hard to see what the big deal is:

"Put a box around any region of space (space-time, really, but I’m going to drop time throughout this essay for ease of visualization, as physicists often do). The holographic principle asserts that no matter what’s going on inside — from gas molecules pinging around to black holes colliding — you can decipher the entire contents of the box just by repeatedly measuring points on the surface."

Well, if we're bounding a region of space-time then, in a purely classical universe governed by deterministic ODEs or PDEs, the statement reduces to a triviality: having information about the boundary amounts to knowing all boundary conditions. Of course I understand that this isn't really the statement of the holographic principle, but the article's formulation is rather underwhelming.

yshklarov··on Why I'm still bearish on LLMs after Navier-Stokes
Thanks!

To expand a bit: It's clear from the writing that the formatting is an intentional stylistic choice. My point wasn't about aesthetic preference. I meant precisely what I wrote. Language evolves and orthography evolves. We've been capitalizing less and less for centuries now. But (until very recently) it was a universal rule to capitalize the first letters of sentences. I believe this is partly because (again, until very recently) people read in large quantities and, to a fluent reader, sentence capitalization serves an important purpose: it helps the eye recognize where one sentence ends and the next begins. Periods alone can be easy to miss, or confuse for commas, when the eye is moving quickly.

Anyway, this is all a very minor point. You have a cool site and I enjoyed your article. I hope you keep writing.

yshklarov··on Why I'm still bearish on LLMs after Navier-Stokes
Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read.

Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.

yshklarov··on More questions about whether researchers can trust OpenAI with unpublished math
We love to do work that is useful and valuable to others, and we often form our identities around this. But identities are in large part socially constructed, so many of us need the recognition of others for our contribution. And it can be very painful when we perceive that the credit for our life's work got "stolen". Naturally, we fight against this. There's nothing shameful there. Sure, you can hold onto an ideal of egoless service. There's nothing wrong with that, either. But it's misanthropic to pass such harsh judgment on people for behaving in such a normal and natural manner.
yshklarov··on Growing proof that autonomous cars save lives
That observed fatality rate, 2, is too low to give a precise estimate of the mean. If we assume fatalities follow a Poisson process, then it's plausible (at 5% significance level) that the true expectation is only 0.36 deaths, not 2.

So it's more useful to look at serious injury statistics as a proxy. For those, Waymo says that there is a significant reduction.

yshklarov··on Ask HN: What are some good unintuitive statistics problems?
It sounds like you're looking for problems in probability theory (rather than statistics). I don't have anything specific for you but you might have better luck searching for problems, puzzles, and examples in probability. For instance:

https://math.stackexchange.com/questions/2140493/counterintu...

yshklarov··on Why does a least squares fit appear to have a bias when applied to simple data?
It has nothing to do with being easier to work with (at least, not in this day and age). The biggest reason is that minimizing sum of squares of residuals gives the maximum likelihood estimator if you assume that the error is iid normal.

If your model is different (y = Ax + b + e where the error e is not normal) then it could be that a different penalty function is more appropriate. In the real world, this is actually very often the case, because the error can be long-tailed. The power of 1 is sometimes used. Also common is the Huber loss function, which coincides with e^2 (residual squared) for small values of e but is linear for larger values. This has the effect of putting less weight on outliers: it is "robust".

In principle, if you knew the distribution of the noise/error, you could calculate the correct penalty function to give the maximum likelihood estimate. More on this (with explicit formulas) in Boyd and Vandenberghe's "Convex Optimization" (freely available on their website), pp. 352-353.

Edit: I remembered another reason. Least squares fits are also popular because they are what is required for ANOVA, a very old and still-popular methodology for breaking down variance into components (this is what people refer to when they say things like "75% of the variance is due to <predictor>"). ANOVA is fundamentally based on the pythagorean theorem, which lives in Euclidean geometry and requires squares. So as I understand it ANOVA demands that you do a least-squares fit, even if it's not really appropriate for the situation.

yshklarov··on Ask HN: What is better to use lead-free/leaded solder?
We have no evidence that the lead in solder makes its way into the body of the person doing the soldering (and we've been at this for quite some time!). The concerns about lead in solder are due to the environmental hazards of electronics waste, and the hazards associated with mining and smelting lead.
yshklarov··on What Is the Fourier Transform?
Really, do you think they've somehow fallen out of favor? If so, that's a surprise to me.

In any case, they are a bit more advanced, and out of scope for the undergraduate course I linked to.

yshklarov··on What Is the Fourier Transform?
As everyone in this thread is sharing links, I'm gonna pitch in, too.

This lecture by Dennis Freeman from MIT 6.003 "Signals and Systems" gives an intuitive explanation of the connections between the four popular Fourier transforms (the Fourier transform, the discrete Fourier transform, the Fourier series, and the discrete-time Fourier transform):

https://ocw.mit.edu/courses/6-003-signals-and-systems-fall-2...

yshklarov··on What Is the Fourier Transform?
I love the visualization! Thanks for sharing.

How do you compute the fractional FT? My guess is by interpolating the DFT matrix (via matrix logarithm & exponential) -- is that right, or do you use some other method?

yshklarov··on Minesweeper thermodynamics
That's pretty neat. I wonder how it works. It's not obvious to me at all how to build something like this, as the program doesn't know the sequence in which the player will reveal the tiles.

I also once made my own variant of this (just like gregfjohnson's idea): A "lucky minesweeper" where luck can be toggled on/off at any point during the game: https://github.com/yshklarov/minesweeper

yshklarov··on The Qweremin
For those who don't recognize the name: Linus Åkesson (lft) is the one who made "Nine", that C64 demo with the wizard and nine sprites that was popular a few months ago (https://news.ycombinator.com/item?id=42940553).
yshklarov··on Why Pascal is not my favorite programming language (1981) [pdf]
Apparently, this is a game that two can play. Niklaus Wirth, the creator of Pascal, had this to say in turn:

"From the point of view of software engineering, the rapid spread of C represented a great leap backward. It revealed that the community at large had hardly grasped the true meaning of the term “high-level language” which became an ill-understood buzzword."

Source: Niklaus Wirth, A Brief History of Software Engineering, 2008 (https://people.inf.ethz.ch/wirth/Miscellaneous/IEEE-Annals.p...)

yshklarov··on Most of the World Can't Code
This is by no means unique to programming. Many areas of knowledge are less accessible to those who don't speak English, and much more so to those who don't speak any of the dozen major languages. Because of this, many people will simply learn (enough) English. It's not such a big deal.

In my view, having a single lingua franca is nice. It better facilitates knowledge transfer. I wouldn't want to see a fracturing where each area of knowledge (or, say, every specialization/application programming) is best treated in a distinct linguistic community. That would be bad for everyone.

yshklarov··on Sky-scanning complete for Gaia
To nitpick with the grammar in the quote: It's capable of measuring to the accuracy of 120 μm at 1000 km. So it cannot accurately measure the diameter of a human hair (which ranges from around 20 to 200 μm) at that distance, but only to the accuracy of a human hair.
yshklarov··on The Missing Nvidia GPU Glossary
Not at all -- the usability and design are fantastic! (On desktop, at least.)

What, specifially, do you find awful here?

yshklarov··on The CAP theorem of Clustering: Why Every Algorithm Must Sacrifice Something
If different components of the dataset have different units, I would argue that it is a prerequisite of clustering to first specify the relative importance of each particular unit (thereby putting all units on the same scale). Otherwise, there's no way the clustering algorithm could possibly know what to in certain cases (such as the ::: example).

It's true that there is no intrinsic meaning to the scale, but you must specify at least a relative scale -- how you want to compare (or weigh) different units -- before you can meaningfully cluster the data. Clustering can only work on dimensionless data.

yshklarov··on The CAP theorem of Clustering: Why Every Algorithm Must Sacrifice Something
Actually, scale-invariance only refers to scaling all dimensions by the same scalar (this is more clearly specificed in the paper linked by the article, page 3). For arbitrary scaling on each coordinate, of course you're correct, it's impossible to have a clustering algorithm that is invariant for such transformations (e.g., the 6-point group ::: may look like either 2 or 3 clusters, depending on whether it's stretched horizontally or vertically).

As for your last two points, I believe I agree! It seems that in the counterexample you give for consistency, some notion of scale-invariance is implicitly assumed -- perhaps this connection plays some role in the theorem's proof (which I haven't read).

This reminds me a bit of Arrow's impossibility theorem for voting, which similarly has questionable premises.

yshklarov··on Discrete Mathematics – An Open Introduction, 4th edition
> This tells me a lot about how a teacher thinks of students.

This is quite an uncharitable perspective!

yshklarov··on Discrete Mathematics – An Open Introduction, 4th edition
I wouldn't count on it -- LLMs make lots of errors in reasoning, and errors in solutions are very frustrating to most math students.
yshklarov··on Discrete Mathematics – An Open Introduction, 4th edition
Not providing solutions is quite common in math textbooks, in part because professors (including the author!) want to be able to assign problems from the textbook to their class, and in part because making solutions is a lot of work!

Outside of a classroom setting, the way you learn from a textbook without external feedback is by engaging more actively with the material.

Treat each statement in the main text as an informal exercise. Each time you come across a proposition -- whether it's a formal theorem statement or a claim in the body of the exposition -- try proving or otherwise justifying it to yourself before reading on.

Take a look at Theorems 2.3.1 and 2.3.2 -- they are very similar. Once you have absorbed the proof of 2.3.1, you can attempt 2.3.2 on your own. If you can't finish the proof, you can read a couple of sentences from the included proof for "hints"... or, if you do finish a proof, you can compare it to the proof in the text.

If you read actively enough, you can learn the material quite well without doing any problems. Many people will claim that you need to do formal problems in order to learn math, but this is untrue. Many math textbooks at the higher level don't include formal exercises or problems at all, and people learn from them just fine.

Admittedly, reading mathematics is a skill in its own right, and you shouldn't expect it to come easily right away. Of course, the best thing is to have a one-on-one teacher, but few of us are so lucky.

yshklarov··on Money and Happiness: Extended Evidence Against Satiation
Check the "About" page. It's a self-published, non-peer-reviewed paper, written by a distinguished academic researcher.
yshklarov··on Ask HN: Coding Ability Evaluations
Could you elaborate?
yshklarov··on Do not confuse a random variable with its distribution
I agree that good definitions are important, but I don't know if this is a fair criticism (even if the sentence were wrong) -- the purpose of this page is to clarify a common point of confusion, rather than to lay out a carefully formalized framework. Besides, the practice of introducing very restrictive definitions in introductory material is universal and, I would argue, pedagogically sound.

That said, while I appreciate and admire free textbooks published online, I think the exposition would be much improved if the author had a better sense of who he was writing for (the most common writing advice...).

And I take issue with the view that the sample space is where the random phenomenon lives, as it were. In my experience, it's more common to use the random variable itself to model the (observable aspect of) the random phenomenon, and for the sample space to be either a hidden (i.e., more abstract) aspect of the phenomenon or else a purely abstract formalism introduced only for ease of mathematical computation.

It would be helpful also to see some more context, especially historical (who introduced the concept of a sample space, and for what purpose?).

yshklarov··on Do not confuse a random variable with its distribution
From what I saw as a recent grad student in probability, most texts do define a random variable to necessarily map into the reals, or the extended reals or perhaps a subset thereof, or occasionally the complex numbers, and the more general concept is a "random element" (when a more specific term is called for, there are "random vectors", "random graphs", "random processes", etc.). But this is certainly not universal even within probability. In any case, I don't believe it matters much -- it's hard to see how a mix-up here might cause any real confusion, though as always it is annoying that there isn't a common convention.
yshklarov··on Golden Rules of Interface Design (2013)
One problem with this: People often learn to touch multiple targets quickly in sequence, because one touch event predictably pops up the target for the next event.
yshklarov··on In every country people think others are less happy than they themselves say (2017)
This chart beggars belief. Could we have a link to the actual studies, and not an anonymous image, please??
yshklarov··on Noise is all around us
I'd love to see the teaching video!
Page 1 of 2Next →