Anyone who tells you they know what the future looks like five years from now is lying.
Anyone who tells you they know what the future looks like five years from now is lying.
On a codebase of 10,000 lines any action will cost 100,000,000 AI units. One with 1,000,000 it will cost 1,000,000,000,000 AI units.
I work on these things for a living and no one else seems to ever think two steps ahead on what the mathematical limitations of the transformer architecture mean for transformer based applications.
Humans also keep struggling with context, so while large contexts may limit AI performance, they won't necessarily prevent them from being strongly superhuman.
OK, I will bite.
So "Sparsely-gated MoE" isn’t some new intelligence, it's a sharding trick. You trade parameter count for FLOPs/latency with a router. And MoE predates transformers anyway.
RLHF is packaging. Supervised finetune on instructions, learn a reward model, then nudge the policy. That’s a training objective swap plus preference data. It's useful, but not breakthrough.
CoT is a prompting hack to force the same model to externalize intermediate tokens. The capability was there, you’re just sampling a longer trajectory. It’s UX for sampling.
Scaling laws are an empirical fit telling you "buy more compute and data" That’s a budgeting guideline, not new math or architecture. https://www.reddit.com/r/ProgrammerHumor/comments/8c1i45/sta...
LoRA is linear algebra 101, low rank adapters to cut training cost and avoid touching the full weights. The base capability still comes from the giant pretrained transformer.
AlphaFold 2’s magic is mostly attention + A LOT of domain data/priors (MSAs, structures, evolutionary signal). Again attention core + data engineering.
"DeepSeek’s cost breakthrough" is systems engineering.
Agentic software dev/MCP is orchestration, that’s middleware and protocols, it helps use the model, it doesn’t make the model smarter.
Video generation? Diffusion with temporal conditioning and better consistency losses. It’s DALL-E style tech stretched across time with tons of data curation and filtering.
Most headline "wins" are compiler and kernel wins: FlashAttention, paged KV-cache, speculative decoding, distillation, quantization (8/4 bit), ZeRO/FSDP/TP/PP... These only move the cost curve, not the intelligence.
The biggest single driver the last few years has been the data so de dup, document quality scores, aggressive filtration, mixture balancing (web/code/math), synthetic bootstrapping, eval driven rewrites etc etc. You can swap half a dozen training "tricks" and get similar results if your data mix and scale are right.
For me a real post attention "breakthrough", would be something like: training that learns abstractions with sample efficiency far beyond scaling laws, reliable formal reasoning, causal/world-model learning that transfers out of distribution. None of the things you listed do that.
Almost everything since attention is optimization, ops, and data curation. I mean give me exact pretrain mix, filtering heuristics, and finetuning datasets for Claude/GPT-5 and without peeking at the secret sauce architecture I can get close just by matching tokens, quality filters and training schedule. The "breakthroughs" are mostly better ways to spend compute and clean data, not new ways to think.
Not necessarily a bad approach but feels like something is missing for it to be “intelligent”.
Should really be called “artificial knowledge” instead.
"They talk by flapping their meat at each other!"
It’s not that it knows grammar, it just was trained on a dataset that applied proper capitalization.
Humans learn from seeing patterns. I suspect AI only repeats them, more like a parrot.
However, the kind of first-order Markov-chain model you're describing does not form "somewhat reasonable" sentences, except very rarely. A second-order Markov model seems to do much better at first, but that's because it's copying big chunks out of its training set. When the training set gets big enough, it stops doing that.
Here's what the first-order model you described looks like trained on Moby Dick and the King James Bible:
> The Project Gutenberg-tm License terms of God, into a prince of the last days, and what things that he had an altar of the camp of the rest content, I laughed us into the sand which were upon the hope and made a snare shall continually flitting through their pains ye offer sacrifice. 50:6 And he saw that the forest of the threshingfloors. 23:2 Son of the fire, and came to anger of Israel. 48:20 And they cannot survive without money. 5:20 Despise not on fire. 1:5 And when he healed of legerdemain in the pitiless jaw; spilled it a chain work, and blood upon by letters. 10:10 And the chief of the Antothite. 12:4 And he hath given them he sprinkled her that the LORD, and hold of Bigthana and hill in the patterns of man, is your God. 3:4 And at the sword. 8:25 And the priest: and cast lots. 23:35 Hezrai the other rings of them competent to, these coffin-canoes were departed from the lower jaw; you to a vast bodies have I will not of which sat still in controversy and my people go. 16:7 And the pledge again with sea-water; which is to the glory, and make ye say, and carried to his disciples said unto the Philistines,
The only grammatical sentence in the bunch is "And they cannot survive without money."
Trained on my reproducible corpus of 49 megs of RFCs, I get:
> RFC 1213, defines a prototype algorithm, since the remote networks to have keys as during a standard is used by the 16-bit words, our first fragment. While designing X.500 and a CR NUL ::= "Community: " {<time> "-" (also known to that databases are not contain the causes the Internet community name and especially since in the data element of decrease in a few protocols are specified in their experimentation. 6. REFERENCES [1] is given in which point several kinds of the destination file is a request arrives, the abstract information present state of cfdpcln and error 38 0[0] Message Indication Not supported. If the textual reference shall be of 10.2.0.0 path down the interface's Area Networks Graphics - Size: The meanings for publication date information Agreements on the SNMPv2 entity and eases the header in datagram shortest-path trees are the Digest Authentication Value ----- ----------- ---------- Receive a list of this command to construct a sequence number of an error in the end-system to unblock the basic monitoring map well, and found and root dispersion <$Ephi tau>, where that would be sent by the array of
This contains zero grammatical sentences.
Here's the source code, so you can see if I fucked up the algorithm:
#!/usr/bin/python
import random, sys
model, last = {}, None
for line in sys.stdin:
for word in line.split():
model.setdefault(last, []).append(word)
last = word
words = list(model.keys())
while True:
last = random.choice(model.get(last) or words)
print(last, end=' ')
(11 lines of code is indeed "less than 100," but maybe you had in mind a more memory-efficient implementation.)I also suspect that current AI only repeats patterns without understanding them, as you evidently did in your comment about Markov-chain models, and as people evidently do almost all of the time.
But I know I don't know, as your earlier comments claim to.
Moreover, I think it's obviously foolish to extrapolate from my current experience with current LLMs to the ultimate limits of the Transformer architecture, much less whatever DeepSeek and Anthropic come up with three years from now. I haven't even implemented a Transformer! Have you?
I may seem a little harsh in this comment, but I think it's important that you stop claiming to know things you don't actually know.
It’s like asking a college student 4th grade math questions and then being impressed they knew the answer.
I’ve use copilot a lot. Faster then google, gives great results.
Today I asked it for the name of a French restaurant that closed in my area a few years ago. The first answer was a Chinese fusion place… all the others were off too.
Sure, keep questions confined to something it was heavily trained on, answers will be great.
But yeah, AI going to get rid of a lot of low skilled labor.
No, it's more like asking a 4th-grader college math questions, and then desperately looking for ways to not be impressed when they get it right.
Today I asked it for the name of a French restaurant that closed in my area a few years ago. The first answer was a Chinese fusion place… all the others were off too.
What would have been impressive is if the model had replied, "WTF, do I look like Google? Look it up there, dumbass."
What's the point of this anecdote? That it's not omniscient? Nobody is should be thinking that it is.
I can ask it how many coins I have in my pocket and I bet you it won't know that either.
> Five years from now AI might still break down at even a small bit of complexity, or it might be installing air conditioners, or it might be colonizing Mercury and putting humans in zoos.
do all these seem logically consistent possibilities to you?
> AI might still break down at even a small bit of complexity, or it might be installing air conditioners, or it might be colonizing Mercury and putting humans in zoos.
that each of these things, being logically consistent, have equal chances of being the case 5 years from now?
>There’s a significant difference between predicting what it will specifically look like, and predicting sets of possibilities it won’t look like
which I took to mean there are probability distributions around what things will happen, and it seemed to be your assertion that there wasn't, that a number of things only one of which seemed especially probable, were equally probable. I'm glad to learn you don't think this as it seems totally crazy, especially for someone praising LLMs which after all spend their time making millions of little choices based on probability.