How to Pursue a Career in Brain-Based AI
numenta.com
numenta.com
I wonder how many members of their team have a background similar to what this article suggests (rather than PhDs and MScs in the field).
Edit: Didn't have to wonder too long, they're hiring interns and would like PhD students with a history of publishing at relevant conferences. So they themselves wouldn't hire an intern based on their own advice!
Is there any reason AI research has to run at fast speeds? Like, any modern learning model that you're researching with could certainly run on a million dollar gpu farm or whatever.... but it could also run on a macbook in 1000x the time, right? And isn't that enough speed to determine whether your algorithms are performing the way you expect, even if they aren't fast enough to do fun interactive stuff like realtime video or whatever?
However, the speedup is so big that it's almost impossible to ignore. One way to measure compute speed is in terms of Floating Point Operations per Second, or FLOPS. A recent-ish CPU is probably ~500 gigaflops, while a single A100 GPU is ~150 teraflops[0]. OpenAI reportedly has a cluster with 10,000 V100 GPUs (the A100's predecessor, but still...). GTP-3 still supposedly took about a month to train on that cluster, so it would never—for all intents and purposes—finish on a MacBook. Few groups operate at that scale, but using even one decent GPU is still such a huge speedup that I doubt many people start with less.
It's also worth noting that "training" the model from examples is often a lot more compute-intensive than using it to do "inference." For example, image recognition models are often trained on large clusters, but can be deployed to something much smaller (a phone or laptop). There's a whole subfield of "distillation", which takes large models and finds ways to simplify them for deployment.
[0] There's some marketing involved in these numbers, different precisions, and the GPU does work in parallel, so you're not getting one operation every 1/150e12 second.
Not much!
I'm not sure why you think HTM layers are bigger than modern DL layers. The HTM layer configuration used in the paper (B=128, M=32, N=2048, and K=40) is 335M parameters. Compare to GPT-3 with 96 layers, where each layer has 1.8B parameters. Much larger models than GPT-3 have already appeared with no end in sight as to how much more they can scale.
The point is, if HTM worked, people would throw compute resources at it, just like they do with DL models. But it doesn't.
Information theory, Bayesian methods, Approximate Computations are more relevant for inspiration. Neuroscience is not the filed which studies intelligent behavior.