So I refrer to LLMs as language extrusion confabulation machines. Language extrusion was a term I heard the linguist Emily Bender use. Confabulation because my observation is that talking to an LLM is very much similar to my experience of interacting with Korsakov syndrome patients some years ago.
I look forward to the hype settling down to see what we end up with.
What metric are you using to compare? By most counts, the energy budget of an instance of an LLM in a datacenter is lower than the energy a person uses. Of course, if you count energy per neuron connections vs weights then you'll likely get a quite different number. But then again LLMs do a lot of things with far fewer weights than the brain does neuron connections, even if you only count neurons in some parts of the brain. And of course you can point to capabilities that the brain has but LLMs lack, but on the whole it feels like it's pretty difficult to make a useful like-for-like comparison here.
Also your comment feels like the classic climate denial discourse - say something a bit complicated and a bit difficult to follow that looks at a very small out of context part of the story to cast doubt.
If you're looking at training costs, there is definitely one aspect in which LLMs are obviously significantly less efficient: the amount of information needed for the initial training. This does translate into some pretty high costs but it only needs to be paid once for the amount of work that any given LLM does. In terms of fine-tuning LLMs can get significantly more data-efficient than the initial model, which also means energy-efficient, and it's not obvious to me that it would be drastically worse than a human (though again, only thinking in terms of doing the energy input for the kind of work that an LLM is good at).
(The increased efficiency in comparison to humans is part of the reason why you see Jevon's paradox mentioned a lot whenever concerns about the resources used by LLMs are mentioned: more efficiency can easily mean more resource use in total)
Other than that I'd look at some of the more unique benchmarks for astra, like playing factorio or using blender. It's an entirely different beast.
By the time we can show you data that convinces you that it does work, the next generation would already be out & incrementally dismantling the old conjectures that were true in the previous generations.
You're fundamentally asking for a violation of how information passively disseminates amongst humans: To go any faster requires more effort on the receiver's part to move up on the adoption curve.
It doesn't matter what comes tomorrow, with the next generation, if the claims now can't be proven.
To preempt the response: The math proof, regardless of them using non-disclosed user data or not, they spent $30M do do something closer to a 1000 monkeys approach, rather than a singular inference being very intelligent.
I agree with the meat of your statement, but am very interested in the pre-emption, "they spent $30M do do something closer to a 1000 monkeys approach, rather than a singular inference being very intelligent". First, I think the $30M number is inflated -- that's what the general public would have paid, but presumably the internal cost is lower, perhaps it's more like $10M. But it is still expensive. Second, I'm curious if it's really the case that they did a 1000-monkeys approach? I haven't read much in-depth reporting about the proof, so it's totally possible I just don't know. What is it that they did which is more like 1000-monkeys? Also, I wonder if that distinction matters -- if 1000 monkeys can reliably make ground breaking proofs, and the approach generalizes to other tasks, I will happily become a circus owner. Maybe you're claiming that it won't yield other proofs? Or the proofs are too opaque to be useful to humans? Or it can handle proofs but not other tasks?
Yes, a proof is a proof regardless how you get there. We however don't hear about when they fail, and I doubt their 10000 agents (from their own statement) would necessarily reach another solution/proof (this by leaning towards using user data after finding out others were close). They could as well have attacked another Millenium problem, but they didn't. In whichever case, we will have to wait and see if they (either company) can reach novel solutions/proofs without significant amount of human provided data for the LLM to bridge the gaps.
Further, and this is more of a policy opinion/prediction: If the numerable obtainable (albeit very hard) problems are solved, assuming training data is needed, will it push out future researchers from entering the field due to lack of reachable goals, thus cutting off future training data? LLMs have been great at replacing gateway jobs. But those jobs are what leads to frontier training data (be it maths, physics, chemistry, economics, graphics, prose, etc).
It's so abysmally bad on Google search... and it's free. Isn't Google the great pioneer of the product is us?
That's definitely part of your problem.
In my recent experience, error rates for astra/fable are at or below human level. Just like when directing humans, it pays to ask probing questions ('Are you sure about X?', 'Did you check for Y?', 'Please run Z just to double check.') if you really care about the result being correct.
Weak models tasked with review can catch a decent amount of the mistakes that weak models make and help them be much better, especially if you have them verify against authoritative sources. Strong models make far fewer mistakes to begin with. And, for the mistakes they do make, a swarm of reviewers (same model or somewhat weaker, reviewed by the stronger model) can really help reduce the error rate further.
You’re using something that is very energy efficient; you cannot extrapolate that experience to conclude that SOTA models are not doing something much different.