Neural ODE's: Understanding how they model data
jontysinai.github.io
jontysinai.github.io
My guess is it will need some deep wizardry of the same kind as OpenAI gradient check-pointing.
Neural ODE, is a nice trick to reduce the memory usage to O(1) instead of O(nb timesteps). But the implementation cost and complexity cost probably mean we are better using gradient check-pointing on a forward dynamic and pay the memory cost.
It will also probably won't play well with noise.
Are there any implementation of it in tensorflow yet?
If that is the case, seems you can implement a full static version of ODE net with tf.While.
If this is supposed to be the ODE definition, shouldn't it be y'(x) = f(x,y)? Otherwise I don't quite understand the definition of 'f'.