332 karma · joined May 15, 2025
You're right though they largely are "smart" in the 80's computer sense. This is largely due to continual learning being unsolved.
BUT the more you look at them, research, and try experiments there's something there not in a 80s computer. If I had to guess maybe 1-5% of a humans ability but it's there. They are able to do novel things but ever step outside of their distribution takes exponential effort for every small addition. There is a true ability to adapt and learn new things on the fly, things never seen before. That is the the smart part. There something hidden in these things we don't understand that allows novel insights built from in context learning.
It's actually measurable in experimental settings but even there it's hard to tease out. I saw it mostly while doing CL training experiments. But I also see it while working with them for coding novel things.
But the more power we provide and farther down the road of this we go those 1-5% are things like solving unsolved math problems. No human solved these things. You say brute force, I say it needed massive effort to break out of it's distribution and get those small insights. It's very human like when taken at scale. The scary thing is that scale is getting smaller every day.
Today it cost massive effort but it's possible 10-20yrs from now an AI could solve a problem like this in under an hour with a single thread on a free subscription paid for by serving an ad.
These arguments are so weak because you'll then have to make the same one a few years from now when it does something else impossible. The argument only stands if we assume no progress will occur.
Sythetic data works in a bounded box we have mapped you can't just make up novel data and it works.
Model collapse is completely unsolved.
Astra goes farther than any model before but it cannot go forever.
"Given infinite thinking time a finite number of humans will solve all theorems"
I also love the angle that this was not intelligence just brute force. As if the mathematicians didn't reeaaally want to solve this they were just too lazy to give it a good try.
What does AI have to actually do before you realize these things are actually smart?
In a world of perfectly rational actors you could say "it's still better than the coal plant" but your point is true for the real world and makes much more sense.
Clearly it's the cost AI leaders will have to pay after spending years telling everyone they will die and lose their jobs. They wanted the publicity and with it comes what they sow.
Suprised and glad you responded.
If you solve it then yes humans are toast mental usefulness wise
Your suggestion on this is mitigation then?
Take an aluminum or concrete plant. I see it as a hard argument to give to these companies that they are somehow special and therefore should do more to their already low impact "factory" If we compared these to large scale steel mills for example. I agree with your point that mitigation can help public opinion but on the opposite side I don't see how you can make a reasonable argument to the builders that this is actually worth funding?
Their troubleshoot support tool must be the same team that makes it on local windows troubleshooter because literally every single time you use it, it spins for 10min then says it can't fix anything.
So there's truth in the statement afterall. Clearly a powerful PR campaign going on because there is soooo many dumb takes this kind of argument vanishes.
I do think that amoung all our water uses this is easily the nicest amoung them. I think datacenters are the new factories and as far as factories go they are the cleanest ones ever made. But I suppose if you have the impression AI is useless then it makes sense. But if this is america's future industry we should make room. These centers are even funding infrastructure upgrades. To me they are the closest thing we'll get to reindustrialization in our lives and I would still want a stronger push for even more aggressive buildout.
Only layer beyond this I want is the permission scoping, stronger sandboxing per session and centralized control. I have not seen a clean product around this though where I own the compute.
Couldn't you hook it up to a multiple choice exam?
It's still LLM like, how smart is it? I'm wary of something that the company states they don't want to benchmark it across public benchmarks.
Is this a pure TPU infra? Really high performance solid intelligence.
I don't judge you for not growing your own food when you hand me a burger.
You can make the argument that "high" math has no function value which is fair.
But I would not agree.
The simple fact that you're comparing art to mathematics is making my point for me.
The only thing I read from this is their ego being bruised by a machine.
If these people cared more about discovery and advancement of human knowledge the only thing they should be doing is celebrating. There's no proof of plagarism but that's an independent issue.
How are they not realizing that in the future children will be able to do impossibly hard math but they will be doing something we can't even think of as of now.
One world class mathematician in the future could be advancing mathematics the equivalent of one Riemann hypothesis A DAY.
How are they not celbrating this as the achievment of the centry? Who cares about plagarism at this scale. It has been solved and it wouldn't have been without AI.
I enjoy seeing the fastest human on earth run.
But I'm still getting in my car to go to the shop.
Nobody needs to throw a rock far, or run fast, nobody gets any value except entertainment.
I think your own analogy is actually much worse than my point.
I considered an automatic promotion path but decided against it I want to actually review the data myself first.
What's your model churn rate like? I was worried about customer experience by same day maybe same work getting a totally different model response (also caching is worse)
My basics are I have a test suite that: 1. Finds newest models of my versions 2. Inferences every single model and a few providers for each with a short problem 3. Analyze latency and if a model failed the stupid simple questions drop it and the provider 4. Run larger context haystack kinds of problems.
It's cheap and fast less than 5$ so I can do this daily, hourly, whatever depending on how much I care. If it's mission critical I would say you need a two day study running once an hour to know the STD of model variance.
Then lock a top 3 contenders via latency dropping routing.
Is this easy? No. Is it cheap? Also no. Is it better than just using a trusted labs api? Also not really.
But it does give you exponentially more flexibility. Being able to run 10 unique models at the flick of a switch on a problem for pareto front analysis is amazing. And giving a dropdown for customers for multiple model options is powerful.
If the answer is yes then atleast you're consistent if no then the question is why can't you scale this until breaks? Then never move beyond that limit?
My argument is there's a "break even" point when the power of the AI is larger than the problem you give it to the point it doesn' slop. You then build at that chunk rate and only try to increase it with next gen model. I usually keep a few "screw it" ideas in my back pocket when a new model arrives to see what happens.
"Go rewrite this entire pipeline in rust" "Go train me a custom x model for y"
Fable is the first model that did not just crash and burn on one of these tasks. Astra still can't do the rust migration (goodbye tokens). But I assume eventually it will. Then I'll have to make up a new ridiculous level.
The model training one was literally an identical pipeline I made before AI and it was like a 6mo process. Fable did it better than me in 1 week (with me helping of course). My theory though is that its datascience is massively higher skill than other systems.
You need to find the chunkrate for your problem and style that works.
The roman empire perfectly matched that and most powerful men seem to go that path.
It would also be kinda easy to argue many moderns countries are going down that path.
So I totally agree if AI also cannot do the second part better than a person.
Honestly though, I wouldn't want to take that bet. I never thought that the first thing AI would become super human AGI like is math.
You ask me 10years ago and I'd think the opposite. I think we all would have said we'd have super human HR employees before a super human mathematician.
But here we are.
The random engineer looking at a funny problem 10 years later now has the literal author of the math to talk to about it and implement it.
I have never even spoken to a world class mathematician and now I can have them design with me?
How is this not better in almost everyway?