See for example gradient descent in maths (LaTeX) and in PyTorch: https://colab.research.google.com/github/stared/thinking-in-...
For the smoothest math <-> programming
See for example gradient descent in maths (LaTeX) and in PyTorch: https://colab.research.google.com/github/stared/thinking-in-...
For the smoothest math <-> programming
x.matmul(y).pow(2).sum()
This way we can rite a lot of things, and we don't need to make up a new combination of punctuation marks and special characters for an operation.
For example, while one can write in Python:
torch.sum((x @ y) 2)
I consider it less readable. I mean, here maybe it is fine, but once it gets longer, more complicated, or we want to add new operations, it turns into a mess.
Vide "style" section in: https://github.com/stared/thinking-in-tensors-writing-in-pyt...
sum(A * B) ^ 2
If you define sum as Σ then you can do
Σ(A * B) ^ 2
I don't think you can use * if you want to backprop, for example in layer regularisation. You'd need to use only PyTorch operations such as torch.mul(A, B).pow(2).sum().
The point here isn't that PyTorch can't look similar to Julia in this small example, rather that I can just use regular, concise Julia syntax - unlike in Python + PyTorch where I need to use PyTorch constructs that are outside of Python.
In maths notation, Σ(x * y)^2 would mean Σ((x * y)^2), but in most programming notation, treating Σ as a function, it would be as you say. I'm going with the original formula in https://news.ycombinator.com/item?id=23508661.
I don't know PyTorch, but regular Python, Numpy, Sympy, etc. seem very similar to Julia in this instance.
> I don't know PyTorch, but regular Python, Numpy, Sympy, etc. seem very similar to Julia in this instance.
If you need PyTorch to record the operations for the backward pass later on I believe you need to use PyTorch versions of *, +, etc.: torch.mul, torch.sub, torch.add, etc. In Julia you can just use built in functions and let Flux handle the backward pass.
I have a strong preference for notations that can be read consistently from left to right (vide pipe operators, chaining, etc). A litmus test is if when I read something I use the word 'of'.
If you do anything more complicated that only vector-matrix operations (i.e. a lot of quantum information, all deep learning), with all multiplications you need to specify dimensions somehow. Having only two operators for multiplication is not enough.
torch.sum((x @ y)**2)I did try to use Julia quite a few times (including, I don't know 8 years ago), as I loved the philosophy (brief yet fast, types). Sadly, I never considered it readable - a mix of new concepts, legacy MATLAB syntax, and in general disregard for this part (even function names didn't have consistent naming).
If it moved a lot with that respect, I would be happy some nice examples.
For example, if I want to regularise a layer:
PyTorch:
l2 = layer.pow(2).sum()
loss += lambda * l2
Julia: l2 = sum(layer .^ 2)
loss += λ * l2
I can just use regular Julia functions whereas in PyTorch I need to use special PyTorch functions as to not detach the tape. This only gets worse and less readable the more complex things I want to do.I love when the "data-flow" is consistent, from left to right.
The Julia code (and traditional mathematical notation) reads: middle, right, left.
Plus, instead of inventing infix notation, I find it cleaner to write custom methods/functions rather than rely on (a fixed, limited, and non-apparent) set of built-in inflix operators.
Made-up .^