HNHacker News
TopNewBestAskShowJobs

vatsachak

838 karma · joined November 5, 2025

submissionscomments
vatsachak··on Subquadratic 3SUM and Subcubic APSP
Nice. That's why I'm a former mathematician haha
vatsachak··on Subquadratic 3SUM and Subcubic APSP
Yeah but at least IMO TCS has little to do with real world optimization. Real world optimization uses the easiest possible algorithms with very simple ideas like min-cut flows.
vatsachak··on Subquadratic 3SUM and Subcubic APSP
I am not representative of all mathematicians and I definitely use LLMs to snag problems I couldn't in my previous life. But we now know that they are good at math.

I want lower energy bills, lower rent, better understanding of health etc. more than I want theorems.

vatsachak··on OpenTPU – An open-source AI accelerator, developed by AI
I feel like there is a lot to be gained from an experienced user pointing an LLM in a tasteful direction.
vatsachak··on Subquadratic 3SUM and Subcubic APSP
As a former mathematician, I'm kind of over them using the LLM for math. we know it works. I want them pointed at "data construction", like being libraries, theories and experiments. But I guess they are deduction machines and there is a lot of low hanging fruit with superhuman deduction in math.
vatsachak··on Dust: Pretraining Transformers Without Backpropagation
Random selection seems to favor post training

https://arxiv.org/pdf/2603.12228

vatsachak··on Dust: Pretraining Transformers Without Backpropagation
Models suffer from "catastrophic forgetting" if you train them on new data.

People are working on this field, recent results suggest that continual learning can be possible by converting the input data to "LLMese"

vatsachak··on Dust: Pretraining Transformers Without Backpropagation
I mean the reason why models are using less energy is because they are getting smarter per token and also engineering algorithms/chips that make inference cheaper.

If we could have success with spiking neural networks in silico they would take even less energy, because they don't require global co-ordination. Co-ordination is information and "information = energy by the second law of thermodynamics" is my crank proof

Also the brain has way more parameters than LLMs and also has different neurotransmitters, loops, branching etc so they probably have WAY more capacity than LLMs.

But coding output/W LLMs have us beat

vatsachak··on Dust: Pretraining Transformers Without Backpropagation
Yep. You need to transport all the weights at the boundary regardless of Backprop/NPC.

But the cool thing is that if your NN is split into mostly self contained chunks then you can go widthwise parallel.

An architecture like MOE exploits this fact so that the active weights during pre-training you're backproping only through active experts

vatsachak··on Dust: Pretraining Transformers Without Backpropagation
No real advantage over Neural Nets here; backprop matmuls can be calculated layer by layer so you can chunk backprop across different machines. The real advantage comes from energy savings, you require no global co-ordination
vatsachak··on Dust: Pretraining Transformers Without Backpropagation
I mean co-ordination requires energy though. The brain wattage looks at GPUs and says "skill issue". But you're right that we look at natural energy production techniques and say "skill issue"
vatsachak··on Dust: Pretraining Transformers Without Backpropagation
Not necessarily, backprop is highly parallelizable since it is just a bunch of matrix mults.

Something like Dust skips the backward pass on backprop. But other techniques like Neural Predictive Coding can be completely asynchronous, each "weight" can fire independent of those far away from it. Innocenti, et. al have shown that NPC gradients converge to backprop within a certain "regime".

The win with asynchronous techniques like NPC is that you do not need the extreme co-ordination that backprop requires and hence should be computationally much easier given the right device.

Although at this point the industry has so much money in the forward-backward pass system that I doubt a backprop successor would win unless someone makes NPC hardware feasible and can prove scaling up to billions of params

vatsachak··on Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates
I could have gotten this in one prompt lmao
vatsachak··on Plain text is still one of the best technologies we have
Probably because plain text can encode any form of distilled data. Technically our DNA can be plain text lol
vatsachak··on Differences Between `Foldl` and `Foldr`
Every time I spend a lot of thought on a problem I always go back to F-Algebras and F-CoAlgebras.

Most recursion is through fold and unfold

vatsachak··on LeCun has "zero concerns" about AI wiping out humanity, recent "rogue" incidents
By this definition everyone is at the right place at the right time.

ML research is weird because it's really about

- compute

- data

- architecture

You're at the right place at the right time for the first two and you're probably rediscovering a Schmidhuber for the third

vatsachak··on Sites in ChatGPT
Stuff like a wrapper over Plaid and some OCR is definitely dead for sure.

But currently they can't manage a vending machine as good as human, so I doubt they have our centuries long planning ability

vatsachak··on Sites in ChatGPT
I don't think that Claude or Codex could build and maintain a CMS as of yet.

Software is mostly in the last 5% of work

vatsachak··on Sites in ChatGPT
I think you're underestimating your own role in making sure they are not going off the rails haha
vatsachak··on Sites in ChatGPT
But then you quickly have tech debt build up because current LLMs cant plan over long scales
vatsachak··on Sites in ChatGPT
The idea is great but the websites shown as examples are clearly made by AI because they make no sense.

Current AI is NOT going to replace web devs or coders because it cannot tabula rasa make human sense. I suspect the reverse with a Jevons thingy happening

vatsachak··on Fixing GRPO's credit assignment problem without evaluating every step
How do frontier labs choose which ideas to include in a training run? It seems that there are too many to choose from
vatsachak··on Context Language Models
Eventually the CLM will be a separate model co-trained with the actual model right?

And there will be multiple contexts like hot vs cold pages in DBs.

Speaking of which I am predicting a "Context as a DB" paper within one year

vatsachak··on Context Language Models
Yeah, that would be the canonical solution. Modularity is better
vatsachak··on Adding Floating-Point Decimals for Fun and Profit
threshold = 1e-9 enters the chat
vatsachak··on Surprisingly complex waves reveal the brain's inner workings
I wonder what the "waves" in LLMs are; The global firing patterns that can help us interpret what the weights are doing en-masse
vatsachak··on Coding Is Not Solved
Coding is not solved but this article hasn't accounted for opus 5.5 yet.

Long term planning in LLMs has not been solved.

vatsachak··on Imp is a full port of DSPy to the BEAM
Yeah. Myself, a PhD in math, have less attention to detail than models.

Now their main weakness is creativity, planning and decision making

vatsachak··on The internet discovers TLA+. Now what?
If you don't even check what you're formalizing you must REALLY trust the agents.

I guess if they found real verifiable bugs then it's good

vatsachak··on We're gonna need a lot more mathematicians
The models get things wrong in the way that humans don't.

They will never make a logical error yet make terrible assumptions and poor long scale decisions.

Wake me up when an agent swarm can write gcc in a box sealed from the internet.

Page 1 of 16Next →