4,609 karma · joined July 6, 2015
Well they are doing something about it, just not the way the speakers had in mind.
It's that you can't even measure it, since the way it's defined as a subjective experience, no external measure could ever capture it. This is what gives rise to the p-zombie argument.
To get rid of that you have to accept "functional qualia" as basically equivalent to qualia, which solves the p-zombie issue and resolves half of the hard problem. From there, explaining consciousness is no "harder" than explaining other scale-depedent phenomenon in complex systems like LLMs: still hard, but at least tractable with scientific measurements and experiments.
Also missed is the pushback against AI art: the further devaluation of talent, and an associated loss of meaning many people have. I think this is probably still downstream of it threatening jobs though, since people would not react as violently if they could truly treat art as a hobby instead of as a profession.
Unfortunate, those types of refactorings are my favorite, since they're tightly scoped, easy to verify correctness, and it's like a little puzzle. Bonus points if you write your own collector or use some more obscure parts of the stdlib.
Ok wait I think I see what you mean. Although maybe it's not getting paris _into_ the value vector that's hard, but isolating the residual stream to _only_ that instead of things like other capitals.
So as a naive example maybe at the very first layer consuming your tokens: Q{France} would have high inner product with K{capital} and so our residual would now mostly contain V{capital}, which maybe contains embeddings of all the capitals of all countries. You need some way to filter out all the other stuff, but can't do that without a FFN + activation.
Just throwing in a relu by itself won't help since that would still work on all the elements uniformly, you need some way to put weight on "paris" while suppressing the others, i.e. mixing within the residual stream itself.
Although maybe if you really stretch it, somewhere in a deeper layer you could have 1-hot encoded values with a "gain" coefficient so that when you do the residual addition it's something like {<paris>, <tokyo>, <dc>} + 10000*{<1>, <0>, <0>} and then if you softmax that you get something with most of its mass on "Paris". But it seems like this would not be practical, or it's just shifting the issue to how that the right 1-hot vector is chosen
You actually do get some value, you can file two DTS tickets [1] a year which are (supposedly) looked at by a real apple engineer. Assuming they haven't outsourced it, that feels worth about $100 considering how badly documented their APIs are.
Now of course there's also jevon's "paradox" here, and the automation does allow us to support a larger population so in that sense not all the increased productivity is just "skimmed off the top" as profit. But on the flipside the crux of the other recent [1] HN post is that the wealth disparity is increasing. And if all the increased productivity directly translated to more "physical resources" in the world, that wouldn't be the case.
So something must be getting skimmed of the top, and intuitively you can feel the "rent seeking" layers in society have increased. Gains in efficiency are no longer resulting in surplus of physical products and decrease in prices.
I guess this is the part I find anachronistic. Why do we work in the source scene light for photography, but do the opposite for videos? It makes sense if you assume the viewing device is "dumb" (like a television or CRT, especially in the analog days) but by now I assume the workflows are all fully digital, and even the most basic output device can apply LUTs. When digital video container formats were introduced, why didn't they align with what ICC did? It would have saved a lot of headache for everyone, compared to limited NCLC tags and the mess around EOTFs.
I guess this is referring to https://www.youtube.com/watch?v=KXuZi9aeGTw ?
For some reason video workflows never adapted to the ICC system (probably because in CRT days you couldn't really adapt your decode gamma on the fly) which is basically where the whole debate in https://gitlab.freedesktop.org/wayland/wayland-protocols/-/m... comes from.
I'm not saying that the people using 2.2 EOTF are wrong, but all this just adds to the absurdity: in the modern day where LUTs are cheap and plentiful, instead of tagging content as an ambiguous sRGB it could simply be tagged as gamma 2.2 if it's actually intended to be decoded at that gamma.