HNHacker News
TopNewBestAskShowJobs

cscheid

1,837 karma · joined October 13, 2010

building http://quarto.org at Posit (fka RStudio).

https://cscheid.net https://github.com/cscheid https://bsky.net/profile/cscheid.net

submissionscomments
cscheid··on Fast
One thing that saved an estimated couple of million lives comes to mind.
cscheid··on I booted Linux 293k times in 21 hours
This is relatively subtle stuff, but here's an attempt at describing the general problem. I'm going to describe the deterministic case, but the probabilistic case is effectively the same.

Let's say you have a bug you suspect is from an interaction of any one pair of 10 features being "on" or "off", but you don't know which specific pair causes the problem. Encode each of the states you could set up your code by a 10-digit binary string: 0000000000, 0000000001, 0000000010, 0000000011, etc.

We could try the 45 possibilities in some order, and we would expect that on average it'd take us 22.5 tries to find the bug. But notice how your "target set" is smaller than the universe of strings: there's only 45 pairs of features, but 1024 strings.

What happens if we try a random string of ones and zeros? Now, instead of catching just one possible pair, we are covering many pairs. The only problem is that we now won't be able to know exactly which pair caused the problem when it does. But we can build a corpus of strings that don't trigger the error vs. strings that trigger the error, and a random sampling soon converges on the correct pair.

If you think about why this works, it's because any of these random strings has about a 1/4 chance to trigger the bug: wlog we can reorder the bits so that the buggy feature are the first two digits, and then we see that we have a 1/4 chance of hitting "11" on those two digits.

The problem is that as you increase the size of the subset that needs to be active, the probability that your random strings will actually catch the bug decreases exponentially. For any _fixed_ target size k (the number of features that need to be active), the overall complexity is still polynomial in n (the number of existing features). But if k is a constant fraction of n, then this technique takes exponential time in n.

cscheid··on I booted Linux 293k times in 21 hours
All-pairs is for _pairs of features_. For subsets you're in much deeper trouble because of the exponential dependence on N. For a fixed polynomial dependence, you can get clever and let tail bounds eventually work for you, but for exponentially growing hypothesis sets, that won't work.
cscheid··on Show HN: I created a game to memorize the fretboard
> After reading your comment I dug around to see how many Chapman Stick players there were (What's the TAM for anything targeted at Stick players..if you pardon my startup-speak). The stick subreddit seems exceptionally small: 596 members. I'm guessing there are non-reddit niche community forums with much bigger numbers to boast of.

So, each Stick has a unique serial number. I got mine in 2021 and it's 6596. So that gives you an idea. I totally understand if the market isn't there.

"There's dozens of us!"

cscheid··on Show HN: I created a game to memorize the fretboard
I added my email to your mailing list, but also please feel free to reach out directly to me over email (it's on my account info).

I put together an observablehq notebook with a Stick fretboard (https://observablehq.com/@cscheid/hanon-01-diagrams-for-chap...) for me to work on some fingering patterns, but something like fretboardfly would be so awesome (even if it were just one side at a time). A configuration file like "string tunings + fretboard markers" would totally do it.

cscheid··on Show HN: I created a game to memorize the fretboard
omg, I would pay $100 for something like this on the Chapman Stick fretboard. ("which Stick fretboard and which tuning" of course are the hard questions on said weird instrument... but 10-string Baritone Melody please? :) )
cscheid··on Optimization Without Derivatives: Prima Fortran Version and Inclusion in SciPy
> dummy argument aliasing

Thanks for the precise Fortran terminology; that's what I meant to say in my head but you're correct. From the linked website for everyone else:

"Fortran famously passes actual arguments by reference, and forbids callers from associating multiple arguments on a call to conflicting storage when doing so would cause the called subprogram to write to a bit of that storage by means of one dummy argument and read or write that same bit by means of another."

cscheid··on Optimization Without Derivatives: Prima Fortran Version and Inclusion in SciPy
The language semantics matter greatly.

GHC with `-fllvm` is not going make Haskell any easier to compile just because it's targeting LLVM. Fortran is (relatively!) easy to make fast because the language semantics allow it to. Lack of pointer aliasing is one of the canonical examples; C's pointer aliasing makes systems programming easier, but high-performance code harder.

cscheid··on Reduced cancer mortality with daily Vitamin D intake
> No subgroup is inappropriate unless you know all of the values of all of the parameters

I don't know how to put it less bluntly: you're incorrect.

> But sure, the intent they tried to demonstrate (take all sub-group analyses with a grain of salt)

That's not what they intended to demonstrate. They intended to demonstrate that you need a _reason_ to want to split, and that reason needs to be given _ahead_ of the analysis. If you see the results _and then choose_ a new data analysis (that includes new subgroup analyses), your procedure is no longer statistically sound.

This is a specifically bad statistical practice called HARKing https://en.wikipedia.org/wiki/HARKing

cscheid··on Reduced cancer mortality with daily Vitamin D intake
Yes, exactly. There's a real-life example of an author group who got sufficiently annoyed at a reviewer requesting an inappropriate subgroup analysis: https://www.thelancet.com/journals/lancet/article/PIIS0140-6... Reviewers asked for the subgroup analysis; authors said "no, this is statistically nonsense"; reviewers said "yes, but we'll reject the paper otherwise"; reviewers said "ok, but only if you let us also split on astrological sign".

Result: paper reports that aspirin has an effect, but only if you're not a Gemini or Libra. Too good.

cscheid··on Use Gröbner bases to solve polynomial equations
Sorry. You do mention a linear systems response, and that's what I meant.

In that setting, the eigenvectors work as a generalized forward and inverse fourier transform, and the eigenvalues form the transfer function you allude to in the bold sentence

"The attention mechanism’s role is the same as that of a transfer function in a linear time-invariant system, namely it calculates the frequency response of the transformer model,"

Specifically, it seems to me that this requires a _symmetric_ attention matrix. Which you get from the self-attention mechanisms (two of the three places where they're used in transformers), but not all of them, notably not the one that combines the output of the first two attention mechanisms (one input, and one output)

cscheid··on Use Gröbner bases to solve polynomial equations
Interesting paper.

I have a question about the claim in 6.2 that attention matrices are SPD, if you don't mind my asking.

It seems to me that accepting the empirical result that the eigenvalues are positive isn't enough to get a Fourier Transform interpretation. Specifically, I don't understand the assumption that all attention matrices are symmetric. (I'm sure you know that positive eigenvalues are not enough by themselves, but for other folks reading, [[1 1/2] [1/3 1]] is a simple concrete example.)

Consider Fig. 17 here: https://lilianweng.github.io/posts/2018-06-24-attention/ (this is Fig.1 in Attention is all you need). I understand that you get symmetric attention matrices for the self-attention matrix in the input stream, as well as the masked attention matrix in the output stream (the first block). But I don't understand how you claim symmetry for the final attention mechanism that combines input and output.

And if you don't get symmetry, you don't get the Fourier Transform interpretation and all the nice algebra that follows.

cscheid··on Weight loss relapse associated with exposure to perfluoroalkylate substances
I don't know why you're getting downvoted. I read the methods section and the stats model doesn't seem to adjust for "calories consumed after dieting stopped". Those calories, as you point out, can be correlated with PFAS because of their association with high-calorie foods. This seems to me like a potential confounding factor.
cscheid··on When Slide Rules Ruled (2006) [pdf]
> There used to be all sorts of cardboard "slide rules" that basically used slide rule-like principles to do calculations specific to some particular industry or function

Those are nomograms, and agreed - they're really freaking cool: https://en.wikipedia.org/wiki/Nomogram

cscheid··on Solaris movie review and film summary (1976)
> Does it know? Is it coincidence? Is it malicious?

In case you haven't read it yet, you might enjoy Peter Watts's Blindsight.

cscheid··on GPT4 is up to 6 times more expensive than GPT3.5
Ask it the other way around, "Why is the 8K context only twice as cheap as the 32K context?" and your original answer is clearer: because they think the demand curve supports it.

Price is not determined by cost, but by how much people are willing to pay.

cscheid··on The most boring number in the world is...
Not _any_ problem, but the _specific_ problem of determining how many arbitrarily placed points (cherries) can be split by a given shape of hypothesis classes/classifiers (knives) is literally the definition of VC-dimension, yes :)
cscheid··on The most boring number in the world is...
And if you think sufficiently hard about curved knifes and cherries on high-dimensional pies, you end up studying machine learning theory (https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_di...)
cscheid··on Academic urban legends (2014)
Read the whole paper! Seriously, do it. The author is being cheeky with showing how an academic urban legend gets made and it takes the whole paper to get the punchline.
cscheid··on GitHub Is Down
Looks like a larger one than usual: https://www.githubstatus.com/incidents/t2xwk9mz56f4
cscheid··on How inevitable is the concept of numbers? (2021)
You can't do graphs without numbers, so the history makes sense. Finite graphs are relations over finite sets. You can't do much with finite sets before inventing numbers.
cscheid··on Communicating with Interactive Articles
That would make it not really what distill was about. Concretely, the article my students wrote wouldn’t have been possible - it was literally proposing a new visualization method and showing a real implementation on the web.
cscheid··on Communicating with Interactive Articles
Distill was one of the best experiments in publishing of the last decade, no irony. Unfortunately, it’s worth reflecting on why they are on a hiatus that I fear will be (understandably) permanent.

When I was an academic, I had the privilege of participating in the process of producing one article for Distill, and the amount of work was equivalent to 3-5x the work of any one single publication in other venues. I’m not sure that’s avoidable to achieve the quality that Distill strives for, but it means the incentives are all pointed against it.

A direct consequence of working on an environment with bad incentives is that people there will burn out, which is part of what I think happened.

cscheid··on “A Handbook of Integer Sequences” Fifty Years Later
This makes me so happy to read. I had the privilege of working on the same lab as Neil (and Dave Applegate, another notable person in OEIS). No exaggeration at all to call them geniuses, you hang out with them for 5 minutes and know they're cut from different cloth. Nicest folk, too.
cscheid··on Ask HN: What sub $200 product improved your 2022
Just another upvote for these. They're great if you don't have room for a dumbbell rack, and (of course) they're lighter than the standard dumbbell set.
cscheid··on Against Method
Yeah, I mostly agree.

I would put it slightly differently, and I honestly don't know if I'm being kinder or not. I'd say that in this book, Feyerabend is being a troll. He's out to get a reaction out of you more than to argue in great faith. In my view it's actually to the detriment of his point.

I'm still happy I read it, but I think it's one of those finicky, "meso-scale" ideas that's useful, but doesn't apply at very small or very large scales. It's also interesting that it came out a good decade after moral particularism came out, and it feels to me that his principle is "simply" methodological particularism.

cscheid··on It’s not Tourette’s but a new type of mass sociogenic illness
> embarrassment Kessler syndrome

Thank you, this is a genuinely great turn of phrase.

cscheid··on What is an eigenvalue?
The values associated with each vertex on the _dominant eigenvector_ (the eigenvector associated with the dominant eigenvalue) are the long-term stable state probabilities. That's from a single eigenvector, not "the eigenvectors".
cscheid··on What is an eigenvalue?
That's not true. The page rank is read from the eigenvector, and is the value associated with the given vertex (ie web page). There are as many page rank values as there are web pages, but only one eigenvector from which to read: the dominant eigenvector of the transition matrix, which is the one with the largest eigenvalue. So, only a single eigenvalue for the entire pagerank computation.
cscheid··on Our brains 'time-stamp' sounds to process the words we hear
Antonio Damasio is a neuroscientist who's done work in the area, and has decent popsci books about this idea. I read the older ones, "Descartes' Error" and "The Feeling of What Happens", and they're fun, good reads. Apparently he's written more on the subject as well.
← PreviousPage 4 of 13Next →