838 karma · joined November 5, 2025
I want lower energy bills, lower rent, better understanding of health etc. more than I want theorems.
People are working on this field, recent results suggest that continual learning can be possible by converting the input data to "LLMese"
If we could have success with spiking neural networks in silico they would take even less energy, because they don't require global co-ordination. Co-ordination is information and "information = energy by the second law of thermodynamics" is my crank proof
Also the brain has way more parameters than LLMs and also has different neurotransmitters, loops, branching etc so they probably have WAY more capacity than LLMs.
But coding output/W LLMs have us beat
But the cool thing is that if your NN is split into mostly self contained chunks then you can go widthwise parallel.
An architecture like MOE exploits this fact so that the active weights during pre-training you're backproping only through active experts
Something like Dust skips the backward pass on backprop. But other techniques like Neural Predictive Coding can be completely asynchronous, each "weight" can fire independent of those far away from it. Innocenti, et. al have shown that NPC gradients converge to backprop within a certain "regime".
The win with asynchronous techniques like NPC is that you do not need the extreme co-ordination that backprop requires and hence should be computationally much easier given the right device.
Although at this point the industry has so much money in the forward-backward pass system that I doubt a backprop successor would win unless someone makes NPC hardware feasible and can prove scaling up to billions of params
Most recursion is through fold and unfold
ML research is weird because it's really about
- compute
- data
- architecture
You're at the right place at the right time for the first two and you're probably rediscovering a Schmidhuber for the third
But currently they can't manage a vending machine as good as human, so I doubt they have our centuries long planning ability
Software is mostly in the last 5% of work
Current AI is NOT going to replace web devs or coders because it cannot tabula rasa make human sense. I suspect the reverse with a Jevons thingy happening
And there will be multiple contexts like hot vs cold pages in DBs.
Speaking of which I am predicting a "Context as a DB" paper within one year
Long term planning in LLMs has not been solved.
Now their main weakness is creativity, planning and decision making
I guess if they found real verifiable bugs then it's good
They will never make a logical error yet make terrible assumptions and poor long scale decisions.
Wake me up when an agent swarm can write gcc in a box sealed from the internet.