>By April 2017, the updated specifications had diverged from the test such that the latest versions of Google Chrome, Safari and Mozilla Firefox no longer pass the test as written
4,606 karma · joined July 6, 2015
>By April 2017, the updated specifications had diverged from the test such that the latest versions of Google Chrome, Safari and Mozilla Firefox no longer pass the test as written
I guess that's probably why you still need some MLA layers in there.
https://x.com/waterloo_intern/article/2081762065392541951
and in particular this image which clarifies the key essential difference between liner attention and delta network by examining the case of two tokens with same key but different value
https://pbs.twimg.com/media/HOPCc7BaEAAQDtO.jpg?format=jpg&n...
I think for comparison it would also have been good to have how original quadratic attention handles it: since both keys are identical, the attention would be "evenly divided" between both values so the final output would be the average of both values, as opposed to the latest value
It's worse than that, the _same model_ from the same company can be quietly changed.
>A Pac-Man smart TV app from Bright Data offers users the choice between viewing ads in the game or agreeing to allow their TV to serve as a residential proxy node.
Why do you say this? I've admittedly never done a proper complex analysis course but I got the impression that that complex differentiability was a very strong condition that results in holomprhic functions behaving "nicely" in ways that real functions do not
Assuming you mean C^2 -> C^2, Do you have a link? If so it would be good to add to the wikipedia page. Also I'm not sure, but does the fact that there's a disproof for n=3 imply that it's false in all n>=3, or could there be higher dimensions where it still holds (I'd guess not since you could probably trivially "embed" this in higher dimensions in some way)
The Jacobian conjecture roughly asks what whether local invertibility gives you global invertibility when you restrict only to polynomials (which we might hope "behave nicely"). Apparently for polynomials over reals this was disproved a while back, but up until now the general case of polynomials over complex numbers was open.
(Just as a note for myself, I had to think of why the fact that such jacobians are constant is a byproduct, I guess it's because of lioville's theorem implying that any polynomial over C that never hits 0 must be a constant [because the reciprocal is bounded and thus must also be a constant])
Does anyone more familiar with this know why this _wasn't_ found earlier, when it seems like you could brute-force through some low-order polynomials?
Skilled craftsmen said the same thing, and yet everyone seems to prefer $1 Temu slop that breaks in a year.
This is probably the only reason why "naturopaths" and other "holistic" doctors are still in business, because they'll actually bother to consider a broader picture of a symptom as hinting at some structural imbalance rather than solely treating the symptom as a localized issue to resolve.
Peter Naur explained this decades ago
>Peter Naur argues that programming is fundamentally a human activity of building a mental "theory" - a deep conceptual insight into how a system's parts match the real-world problem it solves. He rejects the prevailing view that programming is merely the mechanical production of source code, specifications, and documentation. Instead, Naur posits that the true product of programming is the shared mental model held in the minds of the developers who built it.
As a human you're the one with the problem you want code to solve, so it's worth having an understanding the problem. Otherwise you risk an X/Y situation, where the LLM ends up solving a problem that may not actually satisfy what you need.
I think what LLMs allow you to do is better abstract away everything that's _not_ essential to the problem you care about, in the same way that libraries or any higher-level programming language does. Ultimately there's still a "core" of the problem that needs to be expressed formally though. When you see people "vibe coding" by prompting the LLM to add constraints at a time until they reach their desired end-goal, this is ultimately "programming" in a sloppy, non-formal way. Better to get the LLM to write everything else _surrounding_ the problem, so you can write and understand the core yourself.
A great variant of the gay jailbreak
A lot of pretraining is also choosing the right type of data, you don't want to just have it ingest garbage (although I read that some amount of garbage actually helps the model be more robust). Pretraining crystallizes a lot of the inductive biases that post-training builds on, so by crafting the right data mixture you can make it easier for it to start off with a good foundation. There is also a lot of focus on mid-training these days, which I understand is basically either the name for the synthetic data stage, or the SFT phase before all the RL
That's not to dismiss claims of there being an "inner world" or "conscious experience" (which isn't really a falsifiable claim, the whole p-zombie thing). But purely in terms of _why_ you'd expect J-space to contain those things, given that the j-space is a subspace of the residual stream with coordinates we can interpret, it seems like your priors should be that anything that could help accomplish its pretraining & post-training objectives would be captured in there.
And this also helps provide an explanation of some of their claims they observed. For instance, they way they present J-space ablation seems almost mystical, that ablating j-space suddenly turns a "ensouled" model into a robotic one. But j-space is really just a specific subspace within the residual stream, so ablating j-space is not much different than adding a steering vector. And presumably to ablate j-space they nulled out a lot of those dimensions, which would ikely involve nulling out some of of the concepts related to emotion. So their claim could be rephrased as "injecting a steering vector that removes emotional components, results in the model having a robotic voice".
Is it not possible that the interventions were the cause of the disease? There's a lot about the body we don't understand, if you're mainlining supplements daily and doing blood transfusions on the regular you're messing around with a delicate biochemical balance.
I like the clarification from https://x.com/XYHan_/status/2074478449020850623#m
>The “J-space” is not a separate, hidden space. It is an alternative coordinate system for intermediate layer activations. Using a Jacobian between the last layer right before unembedding and the intermediate layer, you can “move” rows of the unembedding matrix (corresponding to distinct tokens) into the space that the intermediate activations live in. So each unembedding row/vector has a corresponding vector in the intermediate activation space this way. They use those vectors to generate coordinates for the same intermediate activations (the “J-lens”). Since each coordinate in this alternative coordinate system is now matched with a token, they can now use it to interpret and manipulate the same intermediate activations
>LLMs think in a subconscious space using tokens narrative is completely misleading because >(1) It's a coordinate system. Not a new/separate space. >(2) Tokens only appear because they specifically built the coordinate system using the unembedding vectors of tokens
There is also a good companion piece by Neel Nanda [1] which answers "Why Jacobians rather than linear regression?" which was another question that came to mind
[1] https://www.lesswrong.com/posts/zFJ3ZdQwrTWE9jT5S/a-review-o...
The paper talked about this, the jacobian matrix corrects for the shift in basis from initial to final layer compared to logit lens which assumes that the residual remains in the same basis across layers. Maybe the latter is in fact true for some models/architectures so the J-lens doesn't do anything extra?