The Tensor Cookbook (2024)
tensorcookbook.com
tensorcookbook.com
I am really lost here if I have missed something about indices notations with tensors or some visualization techniques. Or maybe the confusion of a tensor operation depends of the field? Or maybe I just miss practices and experiences with indices notations...
I definitely find it helps to draw parts of your architecture as a tensor diagram. Or perhaps use a library like tensorgrad which makes eveything explicit.
Given that, what additional value is The Tensor Cookbook providing such that it is worth learning an entirely new notation (for me)? It would probably require long term usage to really benefit from these visual depictions.
In the Tensor Cookbook I aim to show the same formulas using tensor diagrams, in a way that (hopefully) make them seem so obvious you don't even need the book afterwards.
In contrast, the Tensor Cookbook was my first introduction to tensor diagrams, so I didn't have any prior experience with them to lean on.
It certainly looks like a useful and powerful technique, but it seems like something that warrants almost a crash course in the topic with some exercises rather than just jumping in.
I guess my personal crash course was writing a book on the topic. And a software library... Ironically this makes it harder for me to appreciate the level of explanation needed for others.
Thus I rely on people like you to tell me where the chain jumps off, so I can expand the sections. Please let me know what sections were too quickly skipped through!
I will be studying this
But luckily tensor diagrams is quite standard in many fields.
(I have nothing to do with the site.)
Tensor ops in e.g. pytorch are pretty opaque to me, too much implicit on shapes which can change as you go down a pipeline. Maybe I'll come to appreciate it better.
For this I needed a good notation for functions applied to specific dimensions and broadcasting over the rest. Like softmax in a transformer.
The function chapter is still under development in the book though. So if you have any good references for how it's been done graphically in the past, that I might have missed, feel free to share them.
Tensors, as used in ML, are much closer to a key-value store with composite keys and scalar values, with most of the complexity coming from deciding how to filter on those composite keys.
Drop me a line if you're interested in a chat. This is something I've been thinking about for years now.
https://arxiv.org/abs/2302.09687
(On functions of 3rd-order "tensors")
((Whereas matrix-functions are of 2nd-order "tensors"))
Playground: https://gitlab.com/katlund/t-frechet
(MATLAB)
Some confusion arises from the difference between f:R -> R and f':R -> R. It's Fréchet derivative is Df:R -> L(R,R) where Df(x)h = f'(x)h. Row vectors and column vectors a just a clumsy way of thinking about this.
BTW, all you need in order to publish on arixv.org is to know a FoF. There is no rigorous peer review. https://arxiv.org/abs/1912.01091, https://arxiv.org/abs/2009.10852.