Tensors, the geometric tool that solved Einstein's relativity problem
quantamagazine.org
quantamagazine.org
(If you're unfamiliar with the definition of dual vector, it's even simpler: it's just a linear function from V to K.)
A tensor field is not a tensor, but the value of a tensor field at any point is a tensor, which satisfies the definition given above, exactly like the value of a vector field at any point is a vector.
The "fields" are just functions.
There are physics books that do not give the easier to understand definition given above, but they give an equivalent, but more obscure, definition of a tensor, by giving the transformation rules for its contravariant components and for its covariant components at a change of the reference system.
The word "tensor" with the current meaning has been used for the first time by Einstein and he has not given any explanation for this word choice. The theory of tensors that Einstein has learned had not used the word "tensor".
Before Einstein, the word "tensor" (coined by Hamilton) was used in physics with the meaning of "symmetrical matrix", because the geometric (affine) transformation of a body that is determined by the multiplication with a symmetric matrix extends (or compresses) the body towards certain directions (the axes that correspond to a rotation that would diagonalize the symmetric matrix). The word "tensor" in the old sense was applied only to what is called now "symmetric tensor of the second order" (which remains the most important kind of the tensors that are neither vectors nor scalars).
The definition of a tensor as linear maps, while simple to understand, has no content that is useful for doing physics. To do any physics, or for that matter, any geometry with tensors, you need to define the notion of covariance and contravariance.
Besides, the starting with the latter notions allow you to define tensors more naturally. You start with trying to understand how geometric objects transform under coordinate transformations, and you slowly but surely end up with tensors.
All the physical quantities that are defined to be tensors are quantities used to transform either vectors into other vectors or tensors of higher orders into other tensors of higher orders (for instance the transformation between the electric field vector and the electric polarization vector).
Therefore all such physical quantities are used to describe multilinear functions, either in linear anisotropic media, or in non-linear anisotropic media, but in the latter case they are applicable only to relations between small differences, where linear approximations may be used.
The multilinear function is the physical concept that is independent of the coordinate system. The concrete computations with a tensor a.k.a. multilinear function may need the computation of contravariant and/or covariant components in a particular coordinate system and the use of their transformation rules. On the other hand, the abstract formulation of the physical laws does not need such details, but only the high-level definitions using multi-linear functions, and it is independent of any choice for the coordinate system.
There is a unique multilinear function a.k.a. tensor, but it can be expressed by an infinity of different arrays of numbers, corresponding to various combinations of contravariant or covariant components, in various coordinate systems. Their transformation rules can be determined by the condition that they must represent the same function. In the books that do not explain this, the rules appear to be magic and they do not allow an understanding of why the rules are these and not others.
Obviously these things are not just useful to physics, but are indispensable, and so I think the assertion that only the definition of tensor that is useful to physics is the definition tensor=multilinear map is somewhat out of step. Perhaps it would be better to assert that the concept of multilinear map is essential to every useful definition of tensors in physics.
This is exactly the point. Abstract physical laws must be invariant to coordinate transformation. From a pedagogical point of view, perhaps this is less important when discussing anisotropic media, but critical when discussing general relativity. Hence, the first reason why many physicist book writer think it very important that covariance/contravariance of tensors be central to both their definition and their pedagogy as applied to physics. You have to convince the student that tensors are the right mathematical objects to describe reality because they preserve this invariance.
The second reason is just as important. Physics is nothing without validating abstract physical laws by experiment. And that validation can not be done without computing predictions. Which in turn will require the right coordinate system, which will require covariance/contravariance of tensors. You can't just disregard these computations as unimportant or unnecessary from either a pedagogical point of view or a deeper philosophical one.
Why always think of tensors as "functions"? In physics, we often think of them as "quantities" - scalar, vector, etc.
I think this is far too simplistic, for one because the values of this putative function depend on the chosen coordinate system.
So I completely agree with the comment you are replying to: when a physicist says "tensor" they really mean a "tensor field" and the definition of the latter is quite a bit more involved than just specifying a multilinear map at each point of a manifold.
This is the essence of notions like scalar, vector, tensor, that they do not depend on the chosen coordinate system.
Only their numeric representations associated with a chosen coordinate system do depend on that system.
If you compute some arbitrary functions of the numeric components of a tensor in a certain coordinate system, in most cases the array of numbers that composes the result will not be a tensor, precisely because the result will really be different in any other coordinate system, while a tensor must be invariant.
All physical laws are formulated only using various kinds of tensors, including vectors and scalars, precisely because they must be invariant at the choice of the coordinate system.
I insist that calling them "just functions" is simplistic. In fact, I'd say that the complexity of your elaborations kind of proves my point.
Note that I deliberately use "simplistic" and not "wrong", since a section is a function of sorts.
Just as a (p,q) tensor is a multilinear object related to a single vector space, a tensor field is a section of a tensor bundle associated to the vector bundle. (A section is just a function on the underlying space whose value at a point lies in the vector space above the point.)
Usually, the vector bundle relevant in physics is the tangent bundle of a 4-manifold.
This abstract way of defining tensors and tensor fields is manifestly invariant under coordinate changes, but it takes some machinery to set up. Whereas the 'numbers associated to each coordinate system which transform in a certain way' is more direct, but the rules can seem arbitrary at first sight. Also, maybe this approach can generalize to allow more transformation rules which might take some time to put into an abstract setting.
Standard example is a matrix A which transforms as PAP^-1 (where P is linear coordinate change) vs a matrix T which is a linear map between vector spaces.
The same issue appears in software where you can expose a data structure as a tuple of numbers/string fields and then define functions on them, or you can expose it as an abstract data type where the user of the library can only apply certain functions on them and the implementation author can choose different representations(coordinate changes) in which to easily compute the functions.
I will help you: if (e_i) is a basis of V and (e_i^*) is its dual basis, then v = \sum_i \alpha(e_i^*) e_i. Can you find such a formula without mentioning the word "basis"?
def unwrap[V](ff: Bidual[V]): V = ff.v```
There's both directions of the isomorphism explicitly defined in a programming language. No choice of basis needed to define the maps, only to prove that the constructor for Bidual really gives you all linear functionals on the dual.
In finite dimensions, V and V* are isomorphic, but not naturally so. The isomorphism requires additional information. You can specify a basis to get the isomorphism, but many bases will give the same isomorphism. The exact amount of information that you need is a metric. If you have a metric, then every orthonormal basis in that metric will give the same isomorphism.
As for why I said metric, see https://en.wikipedia.org/wiki/Metric_tensor. Which is technically a concept from differential geometry rather than linear algebra. But then again, tensors are literally the topic that started this. And it is only in differential geometry that I've ever cared about mapping from V to V*.
Hopefully that's a hint that you should attempt to figure out what someone might be talking about before going to schoolyard insults.
> This comment is in a discussion about an article titled, Tensors, the geometric tool that solved Einstein's relativity problem. Therefore, "tensors are literally the topic that started this discussion."
Again, are manifold involved in any way in the definition of tensors and their properties? No? Then why are you even mentioning "metric tensors"? (Which aren't even tensors, but tensor fields...)
It's multilinear because it's linear in each of its arguments separately: <ca, b> = c<a,b> and <a, cb> = c<a,b>.
Another simple but less obvious example is a rotation (orthogonal) matrix. It takes a vector as an input, and returns a vector. But a vector itself can be thought of as a linear function that takes a dual vector and returns a number (via the inner product, above!). So, applying the rotation matrix to a vector is a sort of "currying" on the multilinear map, while the matrix alone can be considered a function that takes a vector and a dual vector, and returns a number.
In functional notation, you can consider your rotation matrix to be a function (V x V*) -> K, which can in turn be considered a function V -> (V* -> K), where V* is the dual space of V.
A solid can be anisotropic, i.e. with properties that depend on the direction, either because it is crystalline or because there are certain external influences, like a force or an electric field or a magnetic field that are applied in a certain direction.
In (linear) anisotropic solids, a vector property that depends on another vector property is no longer collinear with the source, but it has another direction, so the output vector is a bilinear function of the input vector and of the crystal orientation, i.e. it is obtained by the multiplication with a matrix. This happens for various mechanical, optical, electric or magnetic properties.
When there are more complex effects, which connect properties from different domains, like piezoelectricity, which connects electric properties with mechanical properties, then the matrices that describe vector transformations, a.k.a. tensors of the second order, may depend on other such tensors of the second order, so the corresponding dependence is described by a tensor of the fourth order.
So the tensors really appear in physics as multilinear functions, which compute the answers to questions like "if I apply a voltage on the electrodes deposited on a crystal in this positions, which will be the direction and magnitude of the displacements of certain parts of the crystal". While in isotropic media you can have relationships between vectors that are described by scalars and relationships between scalars that are also described by scalars, the corresponding relationships for anisotropic media become much more complicated and the simple scalars are replaced everywhere by tensors of various orders.
What in an isotropic medium is a simple proportionality becomes a multilinear function in an anisotropic medium.
The distinction between vectors and dual vectors appears only when the coordinate system does not use orthogonal axes, which makes all computations much more complicated.
The anisotropic solids have become extremely important in modern technology. All the high-performance semiconductor devices are made with anisotropic semiconductor crystals.
But these things are vectors, so you could write e.g. v = a⋅x+b⋅y, and then you want e.g. (a⋅x+b⋅y)⊗w = ax⊗w + by⊗w, and so on.
So in some sense, the quotient space construction[1] gives a better "why". It says
* I want to multiply vectors in V and W. So let's just start by writing down that "v times w" is the symbol "v⊗w", and I want to have a vector space, so take the vector space generated by all of these symbols.
* But I also want that (v_1+v_2)⊗w = v_1⊗w + v_2⊗w
* And I also want that v⊗(w_1+w_2) = v⊗w_1 + v⊗w_2
* And I also want that (sv)⊗w = s(v⊗w) = v⊗(sw)
And that's it. However you want to concretely define tensors, they ought to be "a way to multiply vectors that follows those rules". Quotienting is a generic technique to say "start with this object, and add this additional rule while keeping all of the others".
Another way to say this is that the tensor algebra is the "free associative algebra": it's a way to multiply vectors where the only rules you have to reduce expressions are the ones you needed to have.
[0] https://www.youtube.com/live/mqt1f8owKrU?t=500
[1] https://en.wikipedia.org/wiki/Tensor_product#As_a_quotient_s...
A tensor is a multi-dimensional array.
:)
You know the matrices you work with in 2D or 3D graphics environments that you can apply to vectors or even other matrices to more easily transform (rotate, translate, scale)?
Well tensors are the generalisation of this concept. If you’ve noticed 2D games transformation matrices seem similar (although much simpler) to 3D games transformation mateices you’ve probably wondered what it’d look like for a 4D spacetime or even more complex scenarios. Well you’ve now started thinking about tensors.
> or is it the generalized form of the structure encompassing all of it...Kind of like how an n-sphere
If you are asking whether there are examples of tensors parametrized by an integer d, you can cook up examples - like the (d,0) tensor whose input is a sequence of d vectors and just adds up all the components with respect to some basis in each slot and then adds this number across all the d slots.
But just like a n-spere is a special example, of a polynomial in n variables, the above tensor is a specific example - it is the tensor where all the components are 1 in the higher dimensional array.
Sometimes, one considers the algebra of tensors across all dimensions like the symmetric algebra or exterior algebra simultaneously (where there is multiplication operation between tensors of different dimensions), but that might not be what you were asking about.
If you're doing graphics programming you're operating in three and four dimensional spaces mostly (four dimensional being projective spaces), because that's the space you're trying to render in two dimensions. But you'll rarely need anything higher than a matrix (a two-dimensional data structure) for operations on that space.
If you're doing physics you're operating in three, four, and infinite dimensional spaces, mostly. And you'll routinely use higher data structures -- even things like moment of inertia for rigid bodies can't really be described without rank 3 tensors (a three-dimensional data structure).
In statistics and machine learning, you're operating in very high dimensional spaces, and will find yourself using non-square tensors especially (in the other areas everything will be square). The data structures will generally be high dimensional as well, but usually just a function of model complexity; so maybe 4 or 5 dimensional data structures.
You're using the word "dimension" in two (distinct) ways. Instead, use the word "rank":
> (individual numbers (0 rank), vectors (1 rank), matrixes (2 rank), and so on)
Now, we can talk about a 4-dimensional rank-1 tensor, e.g., a 4-element vector.
Now, think about a 4x4 matrix: if we multiply the matrix by a 4-vector, we get a 4-vector out: in some ways, the multiplication has "eaten" one of the ranks of the matrix; but, the dimension of the resulting object is the same. If we had a 3x2 matrix, and we multiplied it by a 3-vector, then both the rank has changed (from 2 to 1) and the dimension has changed (from 3 to 2).
A tensor has any number of rank.
More importantly, the ranks of a tensor come in two "flavors": a vector and a one-form. The concepts are pretty darn general, but one way to get a feel for how they're related is that the transpose of a vector can be its dual. This gets into things like pre- and post- multiplication; or, whether we 'covary' or 'contravary' with respect to the tensor.
Frankly, tensor products are a beast to deal with, mechanically, so the literature mostly deals with them as opaque objects. The modern tensor software libraries and high performance computing has seen a sea-change in the use of GR.
Interestingly, that extra information helps us to differentiate between the same matrix being used in different "roles". For instance, if you have a 4x4 matrix A, you might think of it like a linear transformation. Given x: V = R^4 and y = Ax, then y is another vector in V. Alternatively, you might think of it like a quadratic form. Given two vectors x, y: V, the value xAy is a real number.
In linear algebra, we like to represent both of those operations as a matrix. On the other hand, those are different tensors. The first would be a rank-(1,1) tensor, the second a rank-(2, 0) tensor.
Ultimately, we might write down both of those tensors with the same 4x4 array of 16 numbers that we use to represent 4x4 matrices, but in the sort of math where all these subtle differences start to really matter there are additional rules constraining how rank-(1, 1) tensors are distinct from rank-(2, 0) tensors.
tensors are something which no one has been able to fully or adequately describe. I think you simply have to treat them as a set of operations and not try to map or force them unto existing concepts like linear algebra or matrices. they are similar but otherwise something completely different.
The conflicting definitions of tensors have precedent in lower dimensions: vectors were already being used in computer science to mean something different than in mathematics / physics, long before the current tensormania.
Its not clear if that ambiguity will ever be a practical problem though. For as long as such structures are containers of numerical data with no implied transformation properties we are really talking about two different universes.
Things might get interesting though in the overlap between information technology and geometry [1] :-)
A "tensor", as used in mathematics in physics is not any array, but it is a special kind of array, which is associated with a certain coordinate system and which is transformed by special rules whenever the coordinate system is changed.
The "tensor" in TensorFlow is a fancy name for what should be called just "array". When an array is bidimensional, "matrix" is an appropriate name for it.
The physicist's tensor is a matrix of functions of coordinates that transform in a prescribed way when the coordinates are transformed. It's a particular application of the chain rule from calculus.
I don't know why the word "tensor" is used in other contexts. Google says that the etymology of the word is:
> early 18th century: modern Latin, from Latin tendere ‘to stretch’.
So maybe the different senses of the word share the analogy of scaling matrices.
There is still a "1% difference" in meaning though. This difference allows a physicist to say "the Christoffel symbols are not a tensor", while a mathematician would say this is a conflation of terms.
TensorFlow's terminology is based on the rule of thumb that a "vector" is really a 1D array (think column vector), a "matrix" is really a 2D array, and a "tensor" is then an nD array. That's it. This is offensive to physicists especially, but ¯\_(ツ)_/¯
Well, they don't, it is their components that do (under a change of the coordinate system).
You're totally correct that the tensors in tensorflow do drop the geometric meaning, but there's precedence there from how CS vs math folk use vectors.
x = tf.constant(([1, 2, 3, 4]))
tf.math.multiply(x, x)
<tf.Tensor: shape=(4,), dtype=..., numpy=array([ 1, 4, 9, 16], dtype=int32)>
I could stop right here since it's a counterexample to x being a matrix (with a matrix product defined on it; P.S. try tf.matmul(x, x)--it will fail; there's no .transpose either). But that's only technically correct :)So let's look at tensorflow some more:
The tensorflow tensors should transform like vectors would under change of coordinate system.
In order to see that, let's do a change of coordinate system. To summarize the stuff below: If L1 and W12 are indeed tensors, it should be true that A L1 W12 A^-1 = L1 W12.
Try it (in tensorflow) and see whether the new tensor obeys the tensor laws after the transformation. Interpret the changes to the nodes as covariant and the changes to the weights as contravariant:
import tensorflow as tf
# Initial outputs of one layer of nodes in your neural network
L1 = tf.constant([2.5, 4, 1.2], dtype=tf.float32)
# Our evil transformation matrix (coordinate system change)
A = tf.constant([[2, 0, 0], [0, 1, 0], [0, 0, 0.2]], dtype=tf.float32)
# Weights (no particular values; "random")
W12 = tf.constant(
[[-1, 0.4, 1.5],
[0.8, 0.5, 0.75],
[0.2, -0.3, 1]], dtype=tf.float32
)
# Covariant tensor nature; varying with the nodes
L1_covariant = tf.matmul(A, tf.reshape(L1, [3, 1]))
A_inverse = tf.linalg.inv(A)
# Contravariant tensor nature; varying against the nodes
W12_contravariant = tf.matmul(W12, A_inverse)
# Now derive the inputs for the next layer using the transformed node outputs and weights
L2 = tf.matmul(W12_contravariant, L1_covariant)
# Compare to the direct way
L2s = tf.matmul(W12, tf.reshape(L1, [3, 1]))
#assert L2 == L2s
A tensor (like a vector) is actually a very low-level object from the standpoint of linear algebra. It's not hard at all to make something a tensor. Think of it like geometric "assembly language".In comparison, a matrix is rank 2 (and not all matrices represent tensors). That's it. No rank 3, rank 4, rank 1 (!!). So what does a matrix help you, really?
If you mean that the operations in tensorflow (and numpy before it) aren't beautiful or natural, I agree. It still works, though. If you want to stick to ascii and have no indices on names, you can't do much better (otherwise, use Cadabra[1]--which is great). For example, it was really difficult to write the stuff above without using indices and it's really not beautiful this way :(
More detail on https://medium.com/@quantumsteinke/whats-the-difference-betw...
See also http://singhal.info/ieee2001.pdf for a primer on information science, including its references, for vector spaces with an inner product that are usually used in ML. The latter are definitely geometry.
[1] https://cadabra.science/ (also in mogan or texmacs) - Einstein field equations also work there and are beautiful
Gibbs/Heavysides vectors were more popular at the time.
At least for me.
My understanding is that the cube is a rank 3 tensor, the faces (or rather slices) of the cube are rank 2 tensors (aka matrices), and the edges (slices) of the matrices are rank 1 tensors (aka vectors).
There seems to be a grammar problem here.
And this series by Dialect: https://youtube.com/playlist?list=PL__fY7tXwodmfntSAAyBDxZ4_...
Hackernews user saivan started notes on eigenchris's tensor series videos.
https://grinfeld.org/books/An-Introduction-To-Tensor-Calculu...
There are two ways of using linear maps in the context of physics. One is as a thing that acts on the space . The other is a thing that acts on the coordinates . So when we talk about transformations in tensor analysis, we're talking about coordinate transformatios , not space transformations . Suppose I implement a double ended queue using two pointers:
``` struct Queue {int memory, start, end; } void queue_init(int size) { memory = malloc(sizeof(int) size); start = end = memory + (size - 1) / 2; } void queue_push_start(int x) { start = x; start--; } void queue_push_end(int x) { end++; end = x; } int queue_head() { return start; } int queue_tail() { return end; } void queue_deque_head() { start++; } void queue_deque_tail() { tail--; } ```
See that the state of the queue is technically three numbers, { memory, start, end } (Pointers are just numbers after all). But this is coordinate dependent , as start and end are relative to the location of memory. Now suppose I have a procedure to reallocate the queue size:
``` void queue_realloc(Queue q, int new_size) { int start_offset = q->memory - q->start; int end_offset = q->memory - q->end; int oldmem = q->memory; q->memory = realloc(q->memory, new_size); memcpy(q->memory, oldmem + q->start, sizeof(int) * (end_offset - start_offset); q->start = q->memory + start_offset; q->end = q->memory - end_offset; }
```
Notice that when I do this, the values of start and end can be completely different! However, see that the length of the queue, given by (end - start) is invariant : It hasn't changed!
---
In the exact same way, a "tensor" is a collection of numbers that describes something physical with respect to a particular coordinate system (the pointers start and end with respect to the memory coordinate system). "tensor calculus" is a bunch of rules that tell you how the numbers change when one changes coordinate systems (ie, how the pointers start and end change when the pointer memory changes). Some quantities that are computed from tensors are "physical", like the length of the queue, as they are invariant under transformations. Tensor calculus gives a principled way to make sure that the final answers we calculate are "invariant" / "physical" / "real". The actual locations of start and end don't matter, as (end - start) will always be the length of the list!
---
Physicists (and people who write memory allocators) need such elaborate tracking, to keep track of what is "real" and what is "coordinate dependent", since a lot of physics involves crazy coordinate systems , and having ways to know what things are real and what are artefacts of one's coordinate system is invaluable. For a real example, consider the case of singularities of the Schwarzschild solution to GR, where we initially thought there were two singularities, but it later turned out there was only one "real" singularity, and the other singularity was due to a poor choice of coordinate system:
Although there was general consensus that the singularity at r = 0 was a 'genuine' physical singularity, the nature of the singularity at r = rs remained unclear. In 1921 Paul Painlevé and in 1922 Allvar Gullstrand independently produced a metric, a spherically symmetric solution of Einstein's equations, which we now know is coordinate transformation of the Schwarzschild metric, Gullstrand–Painlevé coordinates, in which there was no singularity at r = rs. They, however, did not recognize that their solutions were just coordinate transform