Factorization tricks for LSTM networks
arxiv.org
arxiv.org
From my experience, RNN weights and the recurrent weights Tf in an LSTM tend to look more like (I + low_rank) rather than low_rank. To be more specific, I gather that with your F-LSTM you do:
T1 = W1 * input
T2 = W2 * T1
output = T2
so
output = W2 * W1 * input
Where W1 and W2 are "factorized by design".However, it seems like the recurrent weights (those for f in your paper) should look more like
T1 = W1 * input
T2 = W2 * T1
output = input + T2
so
output = (I + W2 * W1) * input
That way, you are imposing the simplification that the Tf ~= (I + low_rank) instead of Tf ~= low_rank.
Have you considered this?>>"I gather that with your F-LSTM you do:" - Your understanding of F-LSTM looks correct
>>"output = input + T2" Not quite clear where to put non-linearities. But this looks similar to residual connections. Which, in my experience, is almost always a good idea.
However, "simulating the brain" in general is a much broader problem in scope. Consider: Can you even really truly say that you have "simulated the brain" without also including a physical body for this system, since a brain is inherently tied to a body and corporal existence in the world? It's not clear.
It's likely the problem won't be computational power but figuring out how to extract or isolate only the bits of the system we are interested in from the rest. Consider the genome, for instance: we have had the human genome mapped since 2003 but the unfathomable complexity of the system as a whole makes it difficult (though not impossible) to apply for useful stuff.
It's not just about computational power, it's also about defining computational models that can create human-competitive performance for certain tasks, or, in the case of AGI, for all tasks that might be expected of a human. That's the fiendishly tricky bit.
Those "inefficiencies" are for handling a messy noisy world where we can't just settle on a single solution. There is allot of overhead to play in the game of evolution.
A neuron nucleus can be as small as 3 μm, ... which could pack about tens or hundreds million 5nm transistors.
There is plenty of room at the bottom, as Feynman said.