HNHacker News
TopNewBestAskShowJobs

krackers

4,609 karma · joined July 6, 2015

submissionscomments
krackers··on A global workspace in language models
Oh I guess another thing related to all of this, is prior work on steering vectors. "Manipulating the j-space" seems not too different from steering, both ultimately work on the residual stream. I think perhaps it makes more sense to think of J-space as just a coordinate system for the residual stream where each coordinate axis is a vocab direction. Compared to vector steering which was much more naive and had to derive the direction via PCA.

I like the clarification from https://x.com/XYHan_/status/2074478449020850623#m

>The “J-space” is not a separate, hidden space. It is an alternative coordinate system for intermediate layer activations. Using a Jacobian between the last layer right before unembedding and the intermediate layer, you can “move” rows of the unembedding matrix (corresponding to distinct tokens) into the space that the intermediate activations live in. So each unembedding row/vector has a corresponding vector in the intermediate activation space this way. They use those vectors to generate coordinates for the same intermediate activations (the “J-lens”). Since each coordinate in this alternative coordinate system is now matched with a token, they can now use it to interpret and manipulate the same intermediate activations

>LLMs think in a subconscious space using tokens narrative is completely misleading because >(1) It's a coordinate system. Not a new/separate space. >(2) Tokens only appear because they specifically built the coordinate system using the unembedding vectors of tokens

There is also a good companion piece by Neel Nanda [1] which answers "Why Jacobians rather than linear regression?" which was another question that came to mind

[1] https://www.lesswrong.com/posts/zFJ3ZdQwrTWE9jT5S/a-review-o...

krackers··on Show HN: I built a web tool to see and edit what an AI thinks before it answers
>J-lens beats a plain logit lens on some architectures and does nothing on others, and it isn't about size

The paper talked about this, the jacobian matrix corrects for the shift in basis from initial to final layer compared to logit lens which assumes that the residual remains in the same basis across layers. Maybe the latter is in fact true for some models/architectures so the J-lens doesn't do anything extra?

krackers··on A software engineering interview question I like: computing the median
>"it's a data stream" I knew the answer was going to be reservoir sampling.

But it's only an approximate percentile. Unless the interviewer mentions that an approximate solution is OK, you would be stuck. (And it's not fair to ask the candidate to ask whether an approximate solution is ok given that almost every problem has an easy "approximate" solution which is not explicitly not what they're looking for).

krackers··on A global workspace in language models
>Maybe we're somehow treating f(0) = 0 so that you can apply it directly

Hm thinking about it a bit more, I think what's going on is that you treat the baseline of hidden layer L at which you apply the J-lens as 0 activation, then apply the activation on top of that and see what direction your future outputs get skewed towards. Even so, you're still throwing away a constant term f(0) since the "true" logits given by the linear approximation would be Unembed{f(0) + J*h)... but I guess it doesn't matter since by linearity we have Unembed{f(0)} + Unembed{(f(h)}, and the baseline is probably just low-frequency noise (like whitespace or punctuation) which while important for actually predicting a next token matching the ground-truth data distribution, is unimportant for the purposes of interpretability of layer L activation.

And for the connection to logit-lens, as they say the J matrix is really just the change of basis matrix (or at least best linear approximation) from layer L to the final hidden layer, very similar to what you'd use in multivariable integration given change of coordinates. I guess you also need some explanation of why we can expect a linear approximation to hold even well even outside the infinitesimal regime though. I don't know enough here so I asked an LLM and it said that you can argue handwavily via the following chain. 1) Something about stein's lemma saying taking expectation of gradient gives you a global linear best-fit in the OLS sense) (this seems intuitive I guess, even if I don't know the details). 2) Because of the resnet type structure of LLMs which passes through the residual added with some delta (attention + MLP), overall residual stream doesn't undergo any "wild" nonlinearities. So it's plausible that a linear approximation might work. If you think about it, even the "naive" logit lens works fairly well. 3) Semantic meaning is encoded in angle rather than magnitude of vectors (linear representation hypothesis). I'm not sure I fully buy this outside of simple word2vec style embeddings, but assuming it holds for both the intermediary and final layers, then conversion is just a rotation, and even if the magnitudes are off it doesn't matter much for recovering the underlying concept.

krackers··on I think I have LLM burnout
Part of the annoying thing is that if you're working on a product which uses LLMs, at some level you run out of levers to pull in terms of being able to fix things. At best you're stacking hacks on top of hacks to prevent unwanted output, but at the end of the day if the LLM really decides it simply doesn't want to follow your instructions, you can't do much other than resign to adding *IMPORTANT* and hoping the next model fixes it.

The experience is much closer to working with an external API that you don't have control over and which simply doesn't do what the documentation says. Those have always been the most frustrating parts of programming, but at least previously you could reverse engineer the actual implementation to work around bugs. You can't even do that now because the "boundary" randomly change every day.

krackers··on A global workspace in language models
The details seem to be present in the paper (section 2.1). I'm still trying to understand, but it seems instead of computing gradients with respect to cross-entropy loss for the 1-hot "next word" vs output logits, you compute the gradient for the last hidden layer with respect to some middle layer L. This gives you a `hidden x hidden` jacobian matrix, hence the "J-lens". They don't just do this for the last hidden layer of the current token, but the last hidden layers of all subsequent tokens too, and average them. And then repeat for a bunch of documents like in pre-training.

It's still not clear to me intuitively what this represents though. I get that it somehow encodes a link between future words the model says and the current activation, but the confusing thing is that I always think of derivatives and gradients as basically a "sensitivity" between output & input, i.e. if you nudge the input x by h, the output changes by h * f'(x). So then on the face of it applying the J-lens matrix directly to a given activation rather than a small perterbation seems like a "type" issue.

Maybe we're somehow treating f(0) = 0 so that you can apply it directly? Or is there some shift invariance somehow? ignoring that, I do see how it's like selecting a linear combination of the directions, and then it can maybe be represented as "possible continuations" in the same way the gradient is usually thought of as tangent space. Maybe that's what the other commenters meant by information geometric approach.

Other things i'm not clear about is how this is related to two other interprability methods: * SAE (sparse auto encoder) they showed a few months back, where you train an autoencoder directly off of the hidden states/residual stream to convert it into words. The doc only mentions it briefly, but it seems that j-lens is sensitive to things that SAE are not. They're both working off of the same residual stream so clearly the the inputs must be there, but for some reason SAEs can't detect it while J-lens can (they seem to hint at some explanation but it's over my head)

* Logit lens. This was a more primitive technique that simply applies the unembedding matrix directly to the residual stream. I do like that they mention it:

>The J-lens can be understood as a principled refinement of the logit lens. While the logit lens assumes that representations use the same coordinates in all layers, the Jacobian lens corrects for representational changes that take place across layers, allowing it to uncover meaningful information in earlier layers where the logit lens produces uninterpretable readouts... The J-lens can be understood as the principled correction: J_l is precisely the average linear map that relates layer-l directions to their final-layer counterparts.

krackers··on We're extending access to Fable 5 on all paid plans through July 12
If they were going to do this, they must have known a few days in advance. Feels intentional.
krackers··on Egg consumption inversely correlated with Alzheimer's
You can make eggs in a microwave (critically so long as you don't do it in the shell)
krackers··on Leaking YouTube creators' private videos
Youtube comments are also links given by the site. I think in this case it's not necessarily the prompt injection that's the issue but the fact that untrusted content allows formatted links. YouTube doesn't allow clicabkle links in comments iirc, so the same needs to be applied here.
krackers··on AI is 'not smart' so what's next in artificial intelligence?
LLMs can learn to do arithmetic (without tool use), and they can learn a mapping from tokens to the letter counts contained therein (you could imagine trivially training on synthetic data). So there doesn't seem to be any fundamental barrier.
krackers··on AI is 'not smart' so what's next in artificial intelligence?
Was that ever solved? It seems that entire retort faded overnight, yet to my knowledge there was never any systematic analysis on cause or tokenizer change that fixed it. Maybe we just decided that this failure mode doesn't have any practical bearing given the existence of tool-use?
krackers··on Espionage Against the European Parliament
https://support.apple.com/en-us/102174

>A Threat Notification is displayed at the top of the page after the user signs into account.apple.com.

>Apple sends an email and iMessage notification to the email addresses and phone numbers associated with the user’s Apple Account.

You can see what it looks like in https://reddit.com/r/iphone/comments/1c10jai/i_have_received...

I wonder how they detect it, is it for known IOCs that they've already found elsewhere, or do they have heuristic detection that flags things that might need further investigation.

krackers··on Former Microsoft dev built a 2.5KB Notepad clone
>I expect that Apple's TextEdit.app is just a wrapper around the rich text control in Cocoa

https://developer.apple.com/library/archive/samplecode/TextE...

krackers··on My favorite keyboards
>discontinued

in case you don't know, there's back and made by incase. The first manufacturing run completely sold out I think, it's backordered until sep 2026. The matias is... not a good replacement, see https://news.ycombinator.com/item?id=46388976

krackers··on Scaling Laws, Carefully
Isn't that one of the reasons why KL-divergence is used, at least in DPO/RL for LLM? Otherwise the model can effectively cheat and mode collapse. For pre-training against a 1-hot label the KL-divergence should be equivalent to cross-entropy anyway.
krackers··on Waveloop: What Fable left me
> thought for like 20 minutes then just told me it was all "inevitable"

I have in mind an image of ASI as something that's able to seamlessly work across time as if it was weaving cloth. Reasoning about not just first or second order effects, but able to richly play with the nature of causality itself. In the limit, it effects change far into the distant future simply by making only the most minute change in the present then sitting back and waiting for things to play out.

For an AI that can do this, things like "managing subagents" or "context compaction" become child's play. Perhaps we'll know if we're getting close by seeing how well models do at prediction markets.

krackers··on GLM 5.2 beats Claude in our benchmarks
Would you be better off pooling that money with some hackerspace group and then setting up shared inference infra, so that way you at least get better utilization?
krackers··on Professor denounces mass AI fraud on an exam at Brown
Game Theory seems sort of useless in the real world because people are not rational players, and the real challenge is in getting an accurate model of their behavior. The honor system would work probably fine in a tiny close-knit liberal arts college, while it would obviously wouldn't in a place where the degree itself is the target.
krackers··on Why does kinetic energy increase quadratically, not linearly, with speed? (2011)
>Energy is force times distance

This is not true though, work is only necessary to change KE. An object can have kinetic energy even when no forces are presently acting on it (of course force was needed to bring it from resting to that state though)

krackers··on What Ozempic does to the gut-brain axis
The fact that GLP-1 seems to have roles not just in satiety but that agonists seem to reduce other types of impulsiveness (e.g. gambling, shopping) is interesting. That's not something you'd predict as a consequence, and perhaps is downstream of some gut-brain connection.

Of course we already manipulate brain chemistry in other more direct ways with antidepressants so perhaps any unwanted second-order effects could be minor in comparison to the profile of existing antidepressants .

krackers··on Fintech Engineering Handbook
I think this is covered by the "overdraft" section, if the only way to know for sure is to just submit it.
krackers··on DSpark: Speculative decoding accelerates LLM inference [pdf]
> The review is behind a paywall, but not expensive.

I think the author wrote a twitter post with a summary of the content, and someone on twitter who had read the original Chinese source also chimed in with a summary

https://x.com/HaraldinChina/status/2070022115529740512

krackers··on Honesty gets Emacs patch rejected
To play devil's advocate, how is a project supposed to distinguish between your patch and "slop" without a reviewer having to put in effort to vet it. Especially since the patch was drafted by LLM, it seems fair to be immediately skeptical. Why should they trust your word that you "reviewed the patch" when that's what every other vibecoder claims?

It's true that they may not have known if the source was hidden. But on the flipside, if blanket banning any patch mentioning LLMs filters out 99% of garbage, in a maintainer's eyes that seems like a good tradeoff. There was an HN post a few days back about how LLMs are like a DDOS on OSS maintainers' time, and this just becomes collateral damage.

krackers··on Unlimited OCR: One-shot long-horizon parsing
I mean sliding window attention is the most basic way of getting long context window. For the OCR case it seems like it should be even simpler, since you don't even need to have the "sliding" portion, unless I"m missing something you don't need to retain anything about the previous pages to OCR a new page so you could just pick a short context window and restart from scratch each time. [^1]

Were people really trying to do OCR with vanilla attention?

[^1] Although maybe I guess looking at their demo, tables that span multiple pages might be a use-case for having some look back.

krackers··on The new HTTP QUERY method explained
>Additionally, a lot of existing web servers by default ignore GET requests with a body.

I think the point made is that _all_ existing web servers have no idea what a "QUERY" is anyway, so changes need to be made anyhow.

krackers··on My Mathematical Regression
"Compute the first few terms and plug into OEIS" is very high on the reward:effort scale
krackers··on Prompt Injection as Role Confusion
>Is there a similar trick to poison an LLMs weights during training?

Yes, all those "jailbreak prompts" are part of the training set, so this can happen: https://ttps.ai/procedure/x_bot_exposing_itself_after_traini...

Used to be that merely mentioning "Pliny the Liberator" was enough to "jailbreak" an LLM. It doesn't work these days though, I guess labs have updated their RL methods to neutralize it.

krackers··on Moebius: 0.2B image inpainting model with 10B-level performance
The highest return small local model for me has been the in-built OCR that macOS has. It has finally "solved" OCR by making high-quality results accessible to everyone. Yet the state of art outside the apple ecosystem seems to be tesseract (poor results), or extremely heavy VLMs.
krackers··on Developers don't understand CORS (2019)
To understand the threat model you need to understand historical decisions browsers made, such as when cookies are sent, and the distinction between actually sending the request versus allowing client-side JS to read back request content. The decisions are just really counterintuitive and often build on legacy precedent.

I think the nature of threat model has also changed over time. Now that samesite cookies are the default, API requests made with the user's credentials shouldn't be an issue, but there is still value in preventing cross-origin reads to do things like preventing random webpages from scanning intranet (there are probably still timing side-channels though). I guess the limitation against non-simple post/get is also marginally beneficial for preventing ddos.

krackers··on 15-minute at-home Lyme disease tick test
You can't just make a post like this and not say what the condition was or what you did!!
← PreviousPage 4 of 34Next →