https://github.com/explosion/thinc#no-computational-graph--j...
https://arxiv.org/abs/1711.10455
Could someone give me some more detail?
https://github.com/explosion/thinc#no-computational-graph--j...
https://arxiv.org/abs/1711.10455
Could someone give me some more detail?
About these neural networks, I think it's "just" implementation. Here's the linear layer implementation in Chainer: https://github.com/chainer/chainer/blob/master/chainer/funct...
We have the forward and backward pass organized as class methods here, and the intermediate state from the forward pass is saved into attributes in the instance. So on each call to the network, we make an instance of this LinearFunction class.
In terms of what's being computed, there's really no difference between this and what happens when you call a layer in Thinc. It's just that the state gets captured in the outer scope of the closure. Maybe Thinc's way has a little less overhead, if there are fewer levels of indirection. Thinc uses the Chainer folks' GPU library --- so, unsurprisingly if you define the same network, the benchmarks are very similar.
On the other hand...I do think the implementation matters! Here's a difference for you: if the library approaches it as "we're going to build a computational graph, and execute it", then the library is going to steal the control flow. If the library tells you "here are some functions, and some higher order functions to compose them", you have more access. PyTorch and Chainer doesn't steal the control flow to nearly the extent that Tensorflow does, but they still build up and tear down the state in their objects, and that makes it harder to intrude.
(I'm the author of spaCy and Thinc)