5,552 karma · joined September 11, 2019
So theoretically in another 20-30 years we can take a guess at who we’ll be fighting.
They’re basically just taking equations of the form f(x,y)=k where k=0, and asking “ah but what about when k!=0”?
Indeed, looking at the constant as a variable by looking at the equation f(x,y)=z IS quite useful. (I’m reminded of Feynman favorite trick for solving integrals). That generalization trick is actually useful in MANY contexts. But there’s nothing novel here.
Not just Silicon Valley but also OpenAI, specifically.
I.e. Instead of limiting autoregression to the token level, you introduce a persistent compressed global workspace latent memory vector that is fed back into the self-attention mechanism at every layer or every token step, allowing the network to attend to its own prior attentional states before computing the next token. Obviously that’s going to involve some compression steps.
Trouble is… I think the architecture there is much simpler a tweak than figuring out how to train it.
…that’s likely to just destabilize training for not much if any gain at first. You’re probably gonna have to resort to some really clever (and currently missing) tricks to figure out how to train the network to actually use that feature.
Also, for further insights on some of the unstated policy, I recommend reading the project 2025 document as well. [2]
Perhaps by “coherent policy” you meant “sound” or “effective”? But as it stands it certainly is coherent.
[1] https://www.whitehouse.gov/wp-content/uploads/2025/12/2025-N...
[2] https://www.documentcloud.org/documents/24088042-project-202...
Ordinarily in mathematics there’s a TON of papers to write just combining low level problems with different techniques. Better still, and often considered groundbreaking is borrowing techniques from other fields and adapting them or creating analogous methods to solve problems. A lot of landmark papers have been written this way. This is also what transformers are sort of good at within other contexts. They have super human breadth so I’m hopeful they’ll become real assets in math for a long time. Though the leaps necessary to adapt a technique in a nonobvious way might be too much for a while longer. We’ll see.
Truly novel techniques are quite rate indeed and I don’t know if LLMs can represent them faithfully in their embedding space or not. My inclination is that they probably can most of the time, but I don’t know. Mathematicians would describe such thins as “alien”.
I’m not sure how well the analogy holds up, or if there’s anything to be learned from it though.
Even a few light years away, it’s almost infeasible to build a radio transmitter that could be detected above the noise floor. The amount of energy you’d need would be absurdly wasteful.
I mean, new features I don’t use and useless UI changes have been added at a high rate over the same time period, but holy Christ, it’s pretty bad when some really basic functionality bugs are getting shipped by big names all the time. Like to name one: adobe acrobat is… an absolute hot pile of garbage. Basically any browser offers a MUCH better pdf reader. I don’t understand how the paid version of a product that has only the job of processing PDFs is so bad at it, and has actually regressed over time. And the PDF editing features are also terrible. I don’t know what that program even does well anymore.
It feels like the cognitive gaps on current LLMs are indeed structural, but also that if we solve that structural issue with a new or extended transformer type of architecture, we’ll be looking at a whole new ballgame.
I mean, basically we’re just looking at needing some type of new post training learning architecture. It’s very clear that extending context windows isn’t that. What’s needed is an honest to god, continuous learning and modification process.
I would have to look at the math more specifically, but I would think it would be impossible to both disguise the compromised key and take advantage of the key to interfere with the signal in the way its designed to prevent. And even if you did, I’m fairly sure the jamming would only affect users whose decryption runs through your branch, and in your geographic proximity so I don’t think it would have an advantage of countering the below the noise floor decryption designed to mitigate against jamming.
There were many words I didn’t know though.
When riding a motorcycle, you’ll encounter people that don’t see you almost every trip. The same is not true in a car.
Riding a bike is just a 100% engagement thing with higher risks and lower margins for error, for all kinds of reasons. And it’s not just traffic, minor pavement imperfections become relevant, the necessary skill floor is also higher. It just demands more attention, straight up.
In a car, you shouldn’t, and it’s not without risk, but you CAN occasionally get away with minor distractions: adjusting the radio, seat, etc. That just doesn’t work on a bike as well. I’m failing to properly articulate the why, but it really is fundamentally different in some ways. I’ve spent many years doing both, and the bike just demands more of your attention resources, independent of your vulnerability in the event of an accident.