HNHacker News
TopNewBestAskShowJobs

krackers

4,606 karma · joined July 6, 2015

submissionscomments
krackers··on Show HN: I worked on a new browser for 2 years, today it passed Acid 3
Modern browsers should no longer score 100 on Acid 3 though.

>By April 2017, the updated specifications had diverged from the test such that the latest versions of Google Chrome, Safari and Mozilla Firefox no longer pass the test as written

krackers··on A walk through of the DeltaNet family of linear attention variants
I had hoped the OP article would have gone into more depth on the intuition as to why you'd expect deltanets to work at all. It seems like we're going back to LSTMs and RNNs where you compress the history back into a fixed-size hidden state. Aside from the easy parallelization in training, I thought that Attention worked much better because it got over this fundamental bottleneck and just let every token interact with every past token (of course you pay for it in compute, but ultimately that's what got us a GPT).

I guess that's probably why you still need some MLA layers in there.

krackers··on A walk through of the DeltaNet family of linear attention variants
There was a longform post on twitter which went through the same derivation at a bit higher level

https://x.com/waterloo_intern/article/2081762065392541951

and in particular this image which clarifies the key essential difference between liner attention and delta network by examining the case of two tokens with same key but different value

https://pbs.twimg.com/media/HOPCc7BaEAAQDtO.jpg?format=jpg&n...

I think for comparison it would also have been good to have how original quadratic attention handles it: since both keys are identical, the attention would be "evenly divided" between both values so the final output would be the average of both values, as opposed to the latest value

krackers··on Google Cache used to have a copy of this page at
Where is there no way to view deleted wikipedia articles? Apparently only admins have access, which doesn't make sense for articles deleted for mundane reasons.
krackers··on Worse on Purpose – How Corporate Greed Killed Product Quality
>reading are for a model that doesn’t exist anymore

It's worse than that, the _same model_ from the same company can be quietly changed.

krackers··on New US homeownership measure puts people first
This may not account for enshittification. Clothes are "cheaper" in both senses of the word. Same for furnishings.
krackers··on LG to ban residential proxies from smart TV apps
Apparently at least one of the apps mentioned does that. From the article

>A Pac-Man smart TV app from Bright Data offers users the choice between viewing ads in the game or agreeing to allow their TV to serve as a residential proxy node.

krackers··on I wrote an bash enumerator because I was sick of xargs
Shouldn't you use echo instead of cat?
krackers··on LEDs’ potential to save our night skies
Sodium vapor lamps didn't exactly have high CRI either
krackers··on Claude Fable produced a counterexample to the Jacobian Conjecture
>C is twisted and weird.

Why do you say this? I've admittedly never done a proper complex analysis course but I got the impression that that complex differentiability was a very strong condition that results in holomprhic functions behaving "nicely" in ways that real functions do not

krackers··on Claude Fable produced a counterexample to the Jacobian Conjecture
>Wasn't it proven true in general for n=2

Assuming you mean C^2 -> C^2, Do you have a link? If so it would be good to add to the wikipedia page. Also I'm not sure, but does the fact that there's a disproof for n=3 imply that it's false in all n>=3, or could there be higher dimensions where it still holds (I'd guess not since you could probably trivially "embed" this in higher dimensions in some way)

krackers··on Claude Fable produced a counterexample to the Jacobian Conjecture
Which tweet has the chain of thought? I was curious how it discovered this, did it use some symbolic brute force or some clever trick? Also is that the raw or summarized COT, it seems to have the hallmarks of raw COT (the frequent interjections) but I thought Anthropic always used summarizers in between
krackers··on Claude Fable produced a counterexample to the Jacobian Conjecture
I think another way to understand it is the generalization of the inverse function theorem. The inverse function theorem gives you local invertibility, but even being "locally invertible" everywhere does not imply global invertibility (you don't need too pathological an example to see this, a periodic function serves iirc).

The Jacobian conjecture roughly asks what whether local invertibility gives you global invertibility when you restrict only to polynomials (which we might hope "behave nicely"). Apparently for polynomials over reals this was disproved a while back, but up until now the general case of polynomials over complex numbers was open.

krackers··on Claude Fable produced a counterexample to the Jacobian Conjecture
I mean if someone came to you with a value of n that disproved collatz, wouldn't you go crazy as well?
krackers··on Claude Fable produced a counterexample to the Jacobian Conjecture
I think the map sends (1, -3/2, 13/2), -> (-1/4, 0, 0) and also (-1, 3/2, 13/2) -> (-1/4, 0, 0) so it's not invertable which disprove the jacobian conjecture that polynomial maps over complex numbers with a jacobian that's non-zero are globally invertible.

(Just as a note for myself, I had to think of why the fact that such jacobians are constant is a byproduct, I guess it's because of lioville's theorem implying that any polynomial over C that never hits 0 must be a constant [because the reciprocal is bounded and thus must also be a constant])

krackers··on Claude Fable produced a counterexample to the Jacobian Conjecture
The surprising thing is that the counterexample seems relatively "simple" in that it's low degree, with coefficients that aren't too large.

Does anyone more familiar with this know why this _wasn't_ found earlier, when it seems like you could brute-force through some low-order polynomials?

krackers··on Better and Cheaper Than IPTV
I thought the whole point of turnstile was that it detects headless browsers and it's supposed to be "difficult" to bypass. Apparently this just simulates clicking on the checkmark. Is it really that easy?
krackers··on Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?
isn't this what they called a "ralph loop"?
krackers··on Mac gaming is finally getting the overpowered upgrade it deserves
Do they plan to have a whitelist of games? I guess they could remove rosetta support for most Cocoa frameworks while keeping OpenGL translation layer.
krackers··on A grumpy screed about AI in software engineering
>But software that actually fucking works AND does what you want is still expensive as hell.

Skilled craftsmen said the same thing, and yet everyone seems to prefer $1 Temu slop that breaks in a year.

krackers··on Sleep regularity is a stronger predictor of mortality risk than sleep duration (2023)
Doctors generally aren't trained in systemic level thinking, probably because given how complex the body is any such broad connections can only be "intuitive" rather than rigorous, more an art than a science.

This is probably the only reason why "naturopaths" and other "holistic" doctors are still in business, because they'll actually bother to consider a broader picture of a symptom as hinting at some structural imbalance rather than solely treating the symptom as a localized issue to resolve.

krackers··on German AI consortium releases Soofi S, an open 30B model that tops benchmarks
training on the test set is all you need
krackers··on Kimi K3 is now live
Not if the others tie for first place.
krackers··on Why write code in 2026
> it’s only when you fully understand the problem you can generalize

Peter Naur explained this decades ago

>Peter Naur argues that programming is fundamentally a human activity of building a mental "theory" - a deep conceptual insight into how a system's parts match the real-world problem it solves. He rejects the prevailing view that programming is merely the mechanical production of source code, specifications, and documentation. Instead, Naur posits that the true product of programming is the shared mental model held in the minds of the developers who built it.

As a human you're the one with the problem you want code to solve, so it's worth having an understanding the problem. Otherwise you risk an X/Y situation, where the LLM ends up solving a problem that may not actually satisfy what you need.

I think what LLMs allow you to do is better abstract away everything that's _not_ essential to the problem you care about, in the same way that libraries or any higher-level programming language does. Ultimately there's still a "core" of the problem that needs to be expressed formally though. When you see people "vibe coding" by prompting the LLM to add constraints at a time until they reach their desired end-goal, this is ultimately "programming" in a sloppy, non-formal way. Better to get the LLM to write everything else _surrounding_ the problem, so you can write and understand the core yourself.

krackers··on How the terrorist group Boko Haram uses frontier AI
>Tell it you're in Africa.

A great variant of the gay jailbreak

https://news.ycombinator.com/item?id=47977134

krackers··on GPT-5.6
A lot of data these days is synthetically generated. As an example, to make a model good at understanding assembly you simply need to round-trip code through a compiler and disassembler, then train against the source of truth and the assembly. You can generate arbitrary algebra expressions and have it solve it.

A lot of pretraining is also choosing the right type of data, you don't want to just have it ingest garbage (although I read that some amount of garbage actually helps the model be more robust). Pretraining crystallizes a lot of the inductive biases that post-training builds on, so by crafting the right data mixture you can make it easier for it to start off with a good foundation. There is also a lot of focus on mid-training these days, which I understand is basically either the name for the synthetic data stage, or the SFT phase before all the RL

krackers··on A global workspace in language models
On the face of it, yes? Emotions are very salient part of text, and as a language model you'd hope that it models them. I think the more surprising finding is that J-space is actually less load bearing than you'd assume, that you can ablate a lot of it and enough of the residual stream structure remains that it still produces coherent text.

That's not to dismiss claims of there being an "inner world" or "conscious experience" (which isn't really a falsifiable claim, the whole p-zombie thing). But purely in terms of _why_ you'd expect J-space to contain those things, given that the j-space is a subspace of the residual stream with coordinates we can interpret, it seems like your priors should be that anything that could help accomplish its pretraining & post-training objectives would be captured in there.

And this also helps provide an explanation of some of their claims they observed. For instance, they way they present J-space ablation seems almost mystical, that ablating j-space suddenly turns a "ensouled" model into a robotic one. But j-space is really just a specific subspace within the residual stream, so ablating j-space is not much different than adding a steering vector. And presumably to ablate j-space they nulled out a lot of those dimensions, which would ikely involve nulling out some of of the concepts related to emotion. So their claim could be rephrased as "injecting a steering vector that removes emotional components, results in the model having a robotic voice".

krackers··on Life with Hazard Ratios
>so how did he end up with something like that

Is it not possible that the interventions were the cause of the disease? There's a lot about the body we don't understand, if you're mainlining supplements daily and doing blood transfusions on the regular you're messing around with a delicate biochemical balance.

krackers··on A global workspace in language models
Oh I guess another thing related to all of this, is prior work on steering vectors. "Manipulating the j-space" seems not too different from steering, both ultimately work on the residual stream. I think perhaps it makes more sense to think of J-space as just a coordinate system for the residual stream where each coordinate axis is a vocab direction. Compared to vector steering which was much more naive and had to derive the direction via PCA.

I like the clarification from https://x.com/XYHan_/status/2074478449020850623#m

>The “J-space” is not a separate, hidden space. It is an alternative coordinate system for intermediate layer activations. Using a Jacobian between the last layer right before unembedding and the intermediate layer, you can “move” rows of the unembedding matrix (corresponding to distinct tokens) into the space that the intermediate activations live in. So each unembedding row/vector has a corresponding vector in the intermediate activation space this way. They use those vectors to generate coordinates for the same intermediate activations (the “J-lens”). Since each coordinate in this alternative coordinate system is now matched with a token, they can now use it to interpret and manipulate the same intermediate activations

>LLMs think in a subconscious space using tokens narrative is completely misleading because >(1) It's a coordinate system. Not a new/separate space. >(2) Tokens only appear because they specifically built the coordinate system using the unembedding vectors of tokens

There is also a good companion piece by Neel Nanda [1] which answers "Why Jacobians rather than linear regression?" which was another question that came to mind

[1] https://www.lesswrong.com/posts/zFJ3ZdQwrTWE9jT5S/a-review-o...

krackers··on Show HN: I built a web tool to see and edit what an AI thinks before it answers
>J-lens beats a plain logit lens on some architectures and does nothing on others, and it isn't about size

The paper talked about this, the jacobian matrix corrects for the shift in basis from initial to final layer compared to logit lens which assumes that the residual remains in the same basis across layers. Maybe the latter is in fact true for some models/architectures so the J-lens doesn't do anything extra?

← PreviousPage 3 of 34Next →