> It’s very easy to train a neural network to look at a piece of a sine wave and predict what the next value should be - the fact that the function is periodic actually helps you.
I’m not sure what you mean. Are you talking about training a network to predict cos(x)? In this case nothing at all changes by training on the derivative.
Or do you mean train the net to take as input sin(t) and produce sin(t+0.01)? The problem with this is that any given value of sin(t) has 2 answers for sin(t+0.01), therefore your optimizer is going to spit out 0 as the answer. Plus, this is a different problem entirely, and you lose the ability to infer sin(t) based on t. It doesn’t answer the same question.
Your suggestion is further going to be seriously confounded if the periodic function is more complex, say, the sum of several sin waves. There can be an arbitrary number of values y that correctly match f(x).
You should try it before making assumptions that it’s easy, see what it takes to train a net to predict sin(x), without embedding knowledge of the fact that sin is periodic.
> The problem is improperly evaluating the neural network as a function of time, instead of evaluating the network as a function of state.
This also sounds like an assumption to me that somehow the entire world of research has failed to consider the most obvious of ideas. The point of both optimizers and neural networks is that they can be black boxes, right? It doesn’t matter at all whether the input is time based or position based or a function of money. The network, in theory, can learn any function, time or otherwise, and there’s nothing special about time.
But, neural networks function better with domain knowledge. When you know the function domain is time and that the output is periodic, you can do things to make a network easier to train, like using a periodic activation function.
As a side note, RNNs explicitly model an NN based on previous state. Also all layered NNs can be viewed as a series of smaller nets that feed state to the next net. In some sense, NNs always evaluate as a function of state.