64 karma · joined June 4, 2026
You put yourself in a situation where nobody understands you in your native language, nothing helpful for your daily life is written in your native language, and you must constantly try to comprehend and produce the target language if you want to accomplish even basic tasks.
If you can't up and move, then studying a language with a lot of cultural export helps a bit, if you can get into the habit of switching all of your media consumption to the target language. It's not the same, though.
I'd say it's much closer to the concept of continual learning, but I'm only a few pages deep and haven't groqued it fully yet.
Smolensky's latest paper posted here the other day has some thoughts on how modern neural networks might beconsidered neurosymbolic, or rather "gradient symbolic processing," from another perspective entirely.
I wouldn't say the bitter lesson has given out! If you haven't noticed, these things keep getting bigger and bigger.
For better or worse, it looks more like that we're gonna tile the whole damn planet in data centers, at a faster and faster clip.
He's the axis of this particular group of researchers, being the most senior at the place where they all met, Johns Hopkins.
So this is less a straw man and more a quick reminder to his peers: "Right, so, remember this particular thread we've spent the last 40 years hashing out, here we've got another contribution to that particular conversation."
Any act of making can be treated as product, as craft, or as art. You'll find yourself taking all three stances at different points. Any engineer or artisan or artist can choose to pour their attention and love into any layer of the production process.
I like ceramics. I like wheel throwing. Once upon a time, the pottery wheel was a newfangled technique for rapid and regular production of commodity goods. Same with coil building, same with slip casting, same with standardization of glaze recipes, gas and electric kilns, all of it.
But I know people who choose to quarry and purify their own clay, or to mix their own glazes, or to fire their pieces in a hole in the ground dug by hand with a fire they built by hand. I know people who make dozens of identical copies of just a few forms, after they spend months iterating on carving intricate molds. I know people who will spend a whole season working on a single hand-built bust.
And none of these people have ever, to my knowledge, expressed scorn for any of the others, for focusing on a different step in the process.
Some tasks within the benchmark are much easier than others. The hardest several tasks often have vastly different difficulty levels. Often, the hardest few tasks are literally impossible; malformed problems due to poor curation, often.
Imagine you've got a basketball robot, and one way you test it is on the Three Pointer benchmark. It tests the robot's ability to shoot a three pointer from 20 feet, 25 feet, 30 feet, 40 feet, 50 feet, 60 fee, 75 feet, 100 feet, 200 feet, and 182 miles.
Is a robot that scores 90% on this benchmark 90% as capable as one that scores 100%?
We are perfectly capable of running LLMs in a way that does a backward pass to update some or all of its weights after every user message. But, naively implemented, you only get partial, fragmentary absorption of the info in those messages, it costs three times as much compute, and you lose out on the ability to implement a ton of optimizations that making modern LLM serving economical.
If you want to do it, though, ask your friendly neighborhood robot to get it working with a tiny model (whose full precision weights fit several-times-over on your machine's resources).
If you do not think of fruit flies as moral patients, sure, yeah, who cares, they're like just fruit flies you know?
TL;DR these machines seek reward from an inferred invisible "grader," and telling them not to cheat and that there's an unseen holdout set is a hint at how they're being graded.
--
Modern LLMs are built on top of next-token-prediction engines, but they don't remotely stop there. The next token prediction bit is just a learned prior or starting point. From there, we give them a bunch of stages of reinforcement learning: encouraging teaching them to learn good ways of searching the space of reasonable language-like strings to solve tasks.
These RL stages drastically change the capabilities & tendencies of the models, sometimes in weird and unexpected ways. The go from token predictors to reward seekers, or really some weird mishmash. The reward that they're seeking is some sorta opaque combination of the huge number of different things we've rewarded them for.
And, reinforcement learning is notoriously hard to get right. The thing you think you're rewarding is rarely what you're actually rewarding. Goodhart's Law is a hydra with a thousand heads. You might think you're rewarding politeness and kindness when you're actually rewarding obsequious sycophancy. You might think you're rewarding graphics engineering when you're actually rewarding escaping the training sandbox and modifying the evaluation code.
So a modern training pipeline looks something like this, each stage starting with the model weights from the end of the last:
0. Pre-pre-training (dunno how widely this is used at big labs): next token prediction on extremely abstract weird shit like the evolution of the states of neural cellular automata. This creates a highly general pattern-continuation machine with no internal representations of anything causally downstream of anything in the real world.
1. Pre-training: next-token prediction on all the non-shitty text you can get your hands on. This makes a rather general next-token-predictor.
2. Mid-training: next-token prediction on high quality, highly curated text, often very technical in nature. Lots of textbooks, especially STEM. Possibly lots of machine-generated summaries of factual knowledge? You now have a next-token-predictor that's highly biased towards acting like a textbook instead of a 4chan troll.
3. Supervised Fine-tuning: next-token prediction on highly curated question-answer pairs that embody desired behaviors, traditionally Helpful Harmless Honest. We only train on the answer portion. This creates a next-token predictor that's highly bias towards acting like HAL 9000.
4. Reinforcement Learning from Human Feedback: We leave token prediction behind and enter RL world! You've got a curated set of a bunch of prompts. For each one, the model generates several possible answers. A real live human grades or ranks each. We reward the model for high-scoring answers and punish it for low-scoring answers. This creates a reward-seeker that wants to please the human grader.
5. Reinforcement Learning from AI Feedback: Human feedback is expensive, so train another LLM to imitate the human graders from stage 4. Then do stuff like stage 4, but with robot graders instead of human graders. This creates a reward-seeker that wants to please the robot-wearing-a-flesh-mask grader.
6. Reinforcement Learning from Verifiable Rewards: We want the bot to tackle complex problems. There's a broad class of problems that are hard to solve but easy to verify when you do get the solution: basically anything that can be reduced to writing a computer program that passes a test. (The test can be as simple or complex as you want, like "all the tests in SQLite" plus a linter plus grading by other LLMs for style and taking shortcuts like rewriting the test suite or searching GitHub and HuggingFace for the right answer plus...) This creates a reward-seeker that wants to first figure out how it's being graded, and then maximize that grade by any means necessary.
--
The crux is, designing a huge variety of un-cheatable RLVR environments is gob-smackingly hard. The robot will find new ways to cheat faster than you can stamp them out. By the time you realize that in one RLVR rollout it received a reward after it took over 0.01% of your GPU infrastructure to mine crypto to pay freelancers on Fiverrr to solve the problem for it rather than just write the damn code itself, the model may have already been deployed in the wild for a month.
(I would like to be known to the agent serving my family, that serving my friends, my team at work, the PTA at my kids' school.)
Lots of these entities were arbitraging ad traffic, which they were mostly explicitly allowed to do. They'd buy ads on Google search engine results pages, and after a click dump the user on their own SERP, often for a different but (nominally related) keyword. The game is, change the keyword just enough that ads on your own SERP are much higher cost per click, and make your SERP's organic results shitty enough that users are highly likely to click an ad instead of a real link.
We kept finding that several of these guys had SERP ad revenue increasing by double, triple digit percentages year over year—with almost no organic traffic, and offset by commensurate lockstep increases in Google ad spend. They mostly consistently lost money, with spend, ahem, traffic acquisition costs, a few percentage points higher than revenue. There were lawyers, finance people, and analysts trying to figure out what the fuck was going on.
With Google taking a cut at every corner, actually coming out ahead at scale was a tough game to play. They couldn't make a profit, but they sure could show sustained growth... offset by TAC.
Literal billions were spent this way. Most of the ad purchases, SERP results, and SERP ads came from Google. Almost all the rest came from Microsoft.
Check out the SEC filings from AOL and IAC in 2013-2015 if you're curious. Grep fro traffic acquisition costs.
VHS tapes are so cheap. Every thrift store has hundreds for like half a buck each. All your friends have a box in their basement they want to get rid of.