We've only begun to build multimodal systems; GPT-4 already exhibits improved ability when trained with images vs without.
It isn't "crank science" at all. Despite their mockery and personal insults and dismissive handwaving, no AI capability-pusher has yet made a convincing argument why an entity with extremely strong cognitive powers would not be capable of transforming its surrounding environment on unprecedented scale. The closest one we've heard is "we just won't be able to make AI very smart." Which wishfully will be true, but we're about to spin up multiple Apollo-programs trying to prove that it isn't.
A recent interview with Paul Christiano is about the closest I've come to this. He does note some semi-accurate predictions at the linked timestamp, but the forecast for how things are likely to go is not exactly rosy, though he's quite a bit more optimistic than Eliezer.
https://youtu.be/GyFkWb903aU?t=1357
Also this whole interview was pretty interesting. Near the end he details how few people world-wide actually work on X-risk from AGI. He also outlines how the academic ML community in general just continually keeps getting predictions really wrong, and many aren't taking X-risk seriously.
Overall his is the most balanced take I've seen. A lot better than Eliezer.
> he details how few people world-wide actually work on X-risk from AGI ... and many aren't taking X-risk seriously
still sounds extremely dangerous.
Or you could take that as evidence (and there's a lot more like it) that AGI is a phenomenon so complex that not even the experts have a clue what's actually going to happen. And yet they are barrelling towards it. There's no reason to expect that anyone will be able to be in control of a situation that nobody on earth even understands.
In their defense, LLMs did come a bit out of the blue. In retrospect, Yudkowsky and his disciples were focusing too much on rationality as science/mathematics, bending Bayes to the point of breaking to try and gleam how perfect intelligence works, and how to somehow formalize the aggregate mess of our fuzzy morality.
They failed to predict that shoving gigabytes of random Internet text at a NN, and having it place parts of words as points in a hundred thousand dimension vector space, will suddenly reduce most of what we consider "thinking" into proximity search in said vector space. They were so focused on the theory, they failed to predict the brute-force, messy practice. But so did everyone else.
If anything, Yudkowsky & co. were the only people consistently taking the problem seriously, and got the outline right.