I mean, its the only one that is both not confusing for beginners (because it doesn't trick them into the "cool, let's `simplify` the dx at the denominator with the next one" mindset) and it also translates easily to code (or other 1D encoding), like you can write "second derivative of f with respect to x" as D[f, x, 2] and "integral of a with respect to t" as "D[a, t, -1]".
My point is mainly that there is nothing rational at all in how humans choose notations... without some obscure historic events, we'd probably still be using some derivative of roman numerals or maybe even sexagesimals! (https://en.wikipedia.org/wiki/Sexagesimal)
(I imagine the reason is because "highly performing" individuals have their own "internal" language to think in, so general language is just a for communication, so a political decision... unfortunately for education :()
I don't know if higher-order functions are really that tough to teach/understand but I think it would actually simplify and demistify many things.
Currently, of the widespread notations I like this one best:
f(x0) = d/dx (sin(x)*cos(x)+x^2) | x=x0 (with the last part in subscript)
I also miss variable scoping from math writing and it disturbs me that variable names often carry semantics, like p(x) and p(y) can be the probability density functions of random variables X and Y (so p is a different function depending on the name of the input variable that you substitute so it doesn't actually operate on real numbers, but (string, number) pairs). I'd prefer to explicitly mark the functions as p_X(x) and p_Y(y).
Similar things come up a lot with differential equations where you don't really know whether something (like y or u) is supposed to be a function (of x or t) or "just" a variable.
Despite the general opinion among laypeople that math notation is very precise and unambiguous, I find that it's often very sloppy and unless you already understand the context very well, it can easily be misleading. Math notation is somewhere between normal natural language and programming languages, and depending on the writer it may be closer to one or the other.
One can argue that this is necessary for compactness.
A lot of that is familiarity but I don't think all of it is.
Part of it is the brevity, and part of it is the shortcuts. E.g. when I did my masters, one thing I quickly realised was that papers that expressed an algorithm using mathematical notation almost always lacked essential details.
My impression is that it's too obvious when there are too large leaps in code, whereas in mathematical notation everyone accepts leaps that can obscure that essential details have been left out.
E.g. you'd have papers on thresholding of images for OCR (deciding what is background and what is foreground) where it turned out the results were highly dependent on certain values represented by certain variables that were never defined, for example, putting in a situation of reconstructing parameters by trial and error if you wanted to reproduce the results.
Today I'm immediately suspicious if results are presented as maths outside of fields where the maths is essential (and sometimes even then) as I see it as having a tendency to be used to gloss over sloppy work or save space by leaving out essential details.
I'm sure this is not the case in all fields, and that people with a more extensive maths background will be able to fill in more of those leaps without much effort, and so it might very well be acceptable in some fields. But to me a notation that makes it that easy to hide missing details is a liability.
I'd love to see this beauty or clarity or whatever that people find in mathematics, but I've never caught even a hint of it. Seems like it needs a good IDE to make up for deficiencies in its language.
That allows you to use Python inside a mathematical environment and it's amazing.
Saved my bacon a few times when I need to translate maths to code and I need to poke it a bit to get an understanding of how it works.
(Linguists would want to murder me for saying this, I know.)
That's why some programming languages can even be defined by implementation. (Though as a programmer I try my best to avoid these languages...)
https://en.wikipedia.org/wiki/Total_derivative#The_total_der...
By the way, most of that article is terrible, but the linear map definition is the one we want. There's also the directional (Gâteaux) derivative:
https://en.wikipedia.org/wiki/Directional_derivative
Candidly, unless you really know your problem, we virtually always want the total derivative since it gives rise to things like gradients and Hessians, which are useful objects that we can store in memory.
Now, the reason that I bring these two up is that their spaces, or really their types, are different. Given a function f:X->Y, the total derivative is a linear operator from X to Y:
(total) f'(x) \in L(X,Y)
The directional derivative is an element in the space Y:
(dir) f'(x;dx) \in Y
Now, at this point, the notation is screwed up since we used Lagrange notation for both. The reason that we can get away with this is that under certain assumptions, that are mostly satisfied in the things we care about, we have that:
f'(x)dx = f'(x;dx)
Alright, so why should we care? Leibniz and Newton notation do a terrible job at capturing this information. Lagrange and Euler notation do a good job at this. For your example:
f(x0) = d/dx (sin(x)cos(x)+x^2) | x=x0
The types don't line up because sin(x)cos(x)+x^2 is value, literally a real number, not a function. Using the above, I would write this as:
(x \in R |-> sin(x)cos(x)+x^2)'(x0)
In LaTeX |-> would be \mapsto. This types correctly in the definitions above since
x \in R |-> sin(x)cos(x)+x^2 \in [R -> R]
and
(x \in R |-> sin(x)cos(x)+x^2)'(x0) \in L(R,R)
Of course, you probably wanted the value and not the function, which explains why we cheat in 1-D. So, we really should write:
(x \in R |-> sin(x)cos(x)+x^2)'(x0) 1 \in R
where we feed it the direction 1. And, yes, this is slightly more cumbersome that we may want, which is why there's a huge number of different notations. However, I do assert that the above generalizes properly all the way into infinite dimensions (Hilbert spaces) and provides a good foundation for typing out mathematical codes.
By the way, if anyone is looking for a book that does this right in my opinion, Rudin's "Principles of Mathematical Analysis" is amazing and his notation is good. For infinite dimensions, I prefer Zeidler's "Nonlinear Functional Analysis and Its Applications." Personally, what I look for is what I call properly typed notation that gives us easy access to useful tools like gradients, Taylor series, chain rule, and implicit and inverse function theorems. Again, most engineering and applied math work requires these theorems everywhere, so I find it best to keep them clean.
Is there any sensible rationale for sin^(-1)(x) meaning arcsin(x) while sin^2(x) means (sin(x))^2?
cf.
f(x) = x^2 + 1
x := 2
f(2) = 5It is confusing to beginners, but it's very useful once you understand it.
It's really not much different everywhere else in society. Different programming languages/frameworks/etc. are doing essentially the same thing (if you ignore the speed of execution). All the languages are Turing complete and can do more or less the same IO. But it's still much easier for people to use certain tools for certain problems than others.
The right notation allows you to focus on what's important and forget what is not.
---
[1] and yet, I guess I'll just babble on adding more words anyway...
[2] be they mathematical concepts, programming paradigms, or cultural norms and quirks.
This allows you to do all sort of calculations with wave functions without constantly grinding to a halt bogged down by integrals and conjugates all over the place.
The ladder operators (aka raising & lowering), really put this notation to good use[2]. It would be tremendously tedious to manipulate such expressions repeatedly without the simplifying properties of Bra-Kets. Once you are done manipulating them you then use the rules of the bra-kets to transform it back into a plain old integral, evaluate, and you're done.
I used to think they were very mysterious until I realised they were (essentially) notational convenience.
[1]: https://en.wikipedia.org/wiki/Bra%E2%80%93ket_notation
[2]: http://www.tcm.phy.cam.ac.uk/~bds10/aqp/lec3_compressed.pdf
You could see this as an actively distinct concept (Feynman diagrams are graphs which give rise to some algebraic structure, whereas Dirac's notation is an algebra in itself) or you can see them both as just a means of abstracting away many operations (e.g. in Dirac's notation, this would be multiplication by operators and inner products, both of which are integrals in some sense, as just a non-commutative type of multiplication) in really powerful notation.
It takes a good amount of intelligence to read it, and it took an amazing genius to write it. But put it in plain notation, and it's a collection of 7th grade algebra problems.
Notation is important.
In the same line, it's great at shining a light on the substitution rule for integration.
Given that your point that it can be obscure at first remains valid, I'd walk the middle line of introducing students to the f'(x) notation first; and after the introduction to integrals introduce this notation to them.
A good self-check is to see if you can convert from Leibniz notation to a more rigorous one at any given step in the computation and understand that step rigorously. Personally, I find that functional notation (using D as an operator on the space of functions, etc.) to be as simple to use and much more likely to alert me when I'm about to confuse myself.
Isn't Δ(x^2) = 2xΔx ≠ (Δx)^2 ? The object Δ(x^2) has one infinitesimals while (Δx)^2 has two, and the number of infinitesimals is conserved. (You can only get finite quantities by taking the ratio of equal numbers of infinitesimals.)
If you mean the infinitesimal difference, then Δ(x^2) = 2Δx still isn't (Δx)^2.
But I see how it can be confusing with printed characters. I guess Leibniz just took ligatures for granted when he came up with his stuff.
The derivative w.r.t. x of f(x) is D_x f in Lagrange notation. It looks a bit like matrix multiplication for a good reason—a matrix is just a representation of a linear operator on a finite dimensional vector space.
To the extent that this is true, how much do you think that is due to the notation itself, and how much is due to someone identifying the essential underlying concepts and then making those the basis for a good notation?
-- Whitehead
I think the interesting part is perceiving language as a tool.