Turing Machines Are Recurrent Neural Networks (1996)
users.ics.aalto.fi
users.ics.aalto.fi
In fact, it's probably the case that you can build a NAND gate (and hence a TM) out of any non-linear transfer function. I'd be surprised if this is not a known result one way or the other.
I'm not sure why that's a problem because polynomial approximations are still useful.
https://en.wikipedia.org/wiki/Kolmogorov%E2%80%93Arnold_repr...
For one, only continuous functions can be represented.
Much more importantly, the theorem doesn't prove that it's possible to learn the necessary weights to approximate any function, just that such weights much exist.
With our current methods, only a subset of all possible NNs are actually trainable, so we can only automate the construcion approximations for certain continuous functions (generally those that are differentiable, but there may be exceptions, I'm not as sure).
Fit a pleasant day-trip velocity curve with a NN and the acceleration would kill you.
But all the other universality theorems refer back to it and don't have their own names, for example; Optimal approximation of continuous functions by very deep ReLU networks by Dmitry Yarotsky [1]. The reference for the original theorem would be On The Structure Of Continuous Functions Of Several Variable, David A. Sprecker [2],
[1] http://proceedings.mlr.press/v75/yarotsky18a
[2] https://www.ams.org/journals/tran/1965-115-00/S0002-9947-196...
This is the key part. "Turing completeness" is just a way of reasoning about infinities. (Anything that can recurse or loop infinitely is eventually Turing-complete.)
Turing Machines Are Recurrent Neural Networks (1996) - https://news.ycombinator.com/item?id=10930559 - Jan 2016 (12 comments)
People are very often not internally coherent over periods much shorter than 90 years.
I've yet to meet a human that doesn't eventually run out of steam. You just aren't well situated to notice decoherence on time scales longer than your own coherence.
With new implementations like xformers[1] and flash attention[2] it is unclear where the length limit is on modern transformer models.
Flash Attention can currently scale up to 64,000 tokens on an A100.
[1] https://github.com/facebookresearch/xformers/blob/main/HOWTO...
« Inflated » expectations doesn't mean NO expectations...
People are still throwing « AI » around as a buzzword like it's something distinct from a computer program, in fact the situation got worse because non-neural network programs are somehow dismissed as « not AI » now.
Autonomous cars are still nowhere to be seen, even more so for « General « AI » ».
The singularity isn't much on track either looking at Kurzweil's predictions : we should have had molecular manufacturers by now, and nanobots connecting our brains to the Internet, extending our lives, and brain scanning people to recreate their avatars when they are dead, don't seem like they are going to happen by the 2030s either. (2045 is still far enough away that I wouldn't completely bet against a singularity by then.)
(And Kurzweil doesn't get to blame it on misunderstanding the exponential function : how people have a too linear view of the future and tend to overestimate changes in the near future, but underestimate them in the long term !)
(Really what we should be thinking about is information complexity, not facile analogies.)