1,837 karma · joined October 13, 2010
https://cscheid.net https://github.com/cscheid https://bsky.net/profile/cscheid.net
Let's say you have a bug you suspect is from an interaction of any one pair of 10 features being "on" or "off", but you don't know which specific pair causes the problem. Encode each of the states you could set up your code by a 10-digit binary string: 0000000000, 0000000001, 0000000010, 0000000011, etc.
We could try the 45 possibilities in some order, and we would expect that on average it'd take us 22.5 tries to find the bug. But notice how your "target set" is smaller than the universe of strings: there's only 45 pairs of features, but 1024 strings.
What happens if we try a random string of ones and zeros? Now, instead of catching just one possible pair, we are covering many pairs. The only problem is that we now won't be able to know exactly which pair caused the problem when it does. But we can build a corpus of strings that don't trigger the error vs. strings that trigger the error, and a random sampling soon converges on the correct pair.
If you think about why this works, it's because any of these random strings has about a 1/4 chance to trigger the bug: wlog we can reorder the bits so that the buggy feature are the first two digits, and then we see that we have a 1/4 chance of hitting "11" on those two digits.
The problem is that as you increase the size of the subset that needs to be active, the probability that your random strings will actually catch the bug decreases exponentially. For any _fixed_ target size k (the number of features that need to be active), the overall complexity is still polynomial in n (the number of existing features). But if k is a constant fraction of n, then this technique takes exponential time in n.
So, each Stick has a unique serial number. I got mine in 2021 and it's 6596. So that gives you an idea. I totally understand if the market isn't there.
"There's dozens of us!"
I put together an observablehq notebook with a Stick fretboard (https://observablehq.com/@cscheid/hanon-01-diagrams-for-chap...) for me to work on some fingering patterns, but something like fretboardfly would be so awesome (even if it were just one side at a time). A configuration file like "string tunings + fretboard markers" would totally do it.
Thanks for the precise Fortran terminology; that's what I meant to say in my head but you're correct. From the linked website for everyone else:
"Fortran famously passes actual arguments by reference, and forbids callers from associating multiple arguments on a call to conflicting storage when doing so would cause the called subprogram to write to a bit of that storage by means of one dummy argument and read or write that same bit by means of another."
GHC with `-fllvm` is not going make Haskell any easier to compile just because it's targeting LLVM. Fortran is (relatively!) easy to make fast because the language semantics allow it to. Lack of pointer aliasing is one of the canonical examples; C's pointer aliasing makes systems programming easier, but high-performance code harder.
I don't know how to put it less bluntly: you're incorrect.
> But sure, the intent they tried to demonstrate (take all sub-group analyses with a grain of salt)
That's not what they intended to demonstrate. They intended to demonstrate that you need a _reason_ to want to split, and that reason needs to be given _ahead_ of the analysis. If you see the results _and then choose_ a new data analysis (that includes new subgroup analyses), your procedure is no longer statistically sound.
This is a specifically bad statistical practice called HARKing https://en.wikipedia.org/wiki/HARKing
Result: paper reports that aspirin has an effect, but only if you're not a Gemini or Libra. Too good.
In that setting, the eigenvectors work as a generalized forward and inverse fourier transform, and the eigenvalues form the transfer function you allude to in the bold sentence
"The attention mechanism’s role is the same as that of a transfer function in a linear time-invariant system, namely it calculates the frequency response of the transformer model,"
Specifically, it seems to me that this requires a _symmetric_ attention matrix. Which you get from the self-attention mechanisms (two of the three places where they're used in transformers), but not all of them, notably not the one that combines the output of the first two attention mechanisms (one input, and one output)
I have a question about the claim in 6.2 that attention matrices are SPD, if you don't mind my asking.
It seems to me that accepting the empirical result that the eigenvalues are positive isn't enough to get a Fourier Transform interpretation. Specifically, I don't understand the assumption that all attention matrices are symmetric. (I'm sure you know that positive eigenvalues are not enough by themselves, but for other folks reading, [[1 1/2] [1/3 1]] is a simple concrete example.)
Consider Fig. 17 here: https://lilianweng.github.io/posts/2018-06-24-attention/ (this is Fig.1 in Attention is all you need). I understand that you get symmetric attention matrices for the self-attention matrix in the input stream, as well as the masked attention matrix in the output stream (the first block). But I don't understand how you claim symmetry for the final attention mechanism that combines input and output.
And if you don't get symmetry, you don't get the Fourier Transform interpretation and all the nice algebra that follows.
Those are nomograms, and agreed - they're really freaking cool: https://en.wikipedia.org/wiki/Nomogram
In case you haven't read it yet, you might enjoy Peter Watts's Blindsight.
Price is not determined by cost, but by how much people are willing to pay.
When I was an academic, I had the privilege of participating in the process of producing one article for Distill, and the amount of work was equivalent to 3-5x the work of any one single publication in other venues. I’m not sure that’s avoidable to achieve the quality that Distill strives for, but it means the incentives are all pointed against it.
A direct consequence of working on an environment with bad incentives is that people there will burn out, which is part of what I think happened.
I would put it slightly differently, and I honestly don't know if I'm being kinder or not. I'd say that in this book, Feyerabend is being a troll. He's out to get a reaction out of you more than to argue in great faith. In my view it's actually to the detriment of his point.
I'm still happy I read it, but I think it's one of those finicky, "meso-scale" ideas that's useful, but doesn't apply at very small or very large scales. It's also interesting that it came out a good decade after moral particularism came out, and it feels to me that his principle is "simply" methodological particularism.
Thank you, this is a genuinely great turn of phrase.