On Leibniz Notation
math.stackexchange.com
math.stackexchange.com
Clearly, Leibniz notation does not intrinsically contain any deep insights since at one point Leibniz himself was erroneously induced by it to think that d(xy)= d(x)d(y), which is false. Newton on the other hand thought in terms of simple geometric concepts (areas), which make it crystal clear that d(xy) = x d(y) + y d(x).
On the topic of history of early analysis I cannot recommend enough this collection of anecdotes from Arnold:
But then of course, you see that d^2 f/dx^2 is a bad expression for the second derivative, and that you should actually use differentials properly and write it as (d^2 f)/(dx)^2 - df/dx (d^2 f)/(df)^2.
Now you want to differentiate df/dx with respect to x, so you do "d/dx" to "df/dx":
(d/dx) (df/dx)
(This looks better if you write it with "real" fractions.) On the top you have "d d f" with is naturally written "d^2f". On the bottom, you have "dx dx", which is naturally written "dx^2". So the result should be written d^2 f/dx^2.d/dx (dy/dx) is really d(dy/dx)/dx, which is ((d^2 y)/dx + dy d(1/dx))/dx = d^2 y/dx - dy/dx (d^2 x)/(dx)^2, where df is the actual differential.
I meant to explain why the notation "d^2 f/dx^2" is the way it is. It looks funny, and students would often ask why the 2's were "in different places". But it makes sense when you remember that "d/dx" is the derivative operator.
d^2 f/dx^2 is actually wrong as the way to write the second derivative unless df/dx = 0 or d^2 x/(dx)^2 = 0; the first case is trivial, and the second is almost never the case.
Suppose for example that in actually x = t^2. Then dx = 2t dt, d(dx) = 2(dt)^2 + 2 t (d^2 t) and we get (2(dt)^2 + 2t(d^2 t))/(2t dt)^2 = 0.5 + (...) d^2 t/(dt)^2, so it certainly isn't zero.
But if you have actual differentials you can do algebra.
I still think it was a good thought to ask whether "d/dx" could be interpreted as "exterior derivative divided by dx".
Okay, that's what I figured from your earlier computation. (One note: If "d" means exterior differentiation, then d^2 = 0.) But if you mean d/dx in this way you should make sure you tell people, because that's not the interpretation it has in calculus (say a typical calc class in a typical college) or in differential geometry. In those cases, d/dx means (calc class) the differentiation operator, or a tangent vector (diff geom) [which is again an operator on functions]. That is, "d/dx" is regarded as a single symbol, not a quotient of exterior derivative by a differential.
> d^2 f/dx^2 is actually wrong as the way to write the second derivative ...
I think this is true given your interpretation, based on your earlier computation. But again, in ordinary calc classes (which after all don't talk about differential forms) the "d^2 y/dx^2" notation comes from the standard calc class meaning for d/dx and it goes back ages. I just looked in G. H. Hardy's classic "Pure Mathematics" from over a hundred years ago and yep! - d^2 y /dx^2 is one of the notations for the second derivative. (I'm not sure if it goes back to Leibniz.) So if you're lobbying for a change you have your work cut out for you. :-)
Regardless, I hadn't thought of interpreting "d/dx" as "exterior derivative divided by differential" - interesting idea.
[edited: added note about d^2, minor edit in last sentence]
So something like an exterior derivative where d(df) is not defined to be zero and such terms must be kept.
Newton's algebraic vs. geometric approaches to calculus: Boyer [1] notes that in Principia (1687) "Newton presented them [his propositions] in the form of synthetic geometrical demonstrations with an almost complete lack of analytical calculations" - hence the geometrical approach you point out. But he actually wrote up accounts of his calculus in three papers (De analysi ... (1669), Methodus fluxionum ... (1671), De quadratura ... (1676)), all of which were published after 1700. In those earlier papers, his approach is algebraic, often using the binomial theorem. In one incredible result from De analysi, he shows that the area under y = a x^(m/n) is given by z = (m/(m + n)) a x^((m + n)/n). To do this, he increments z and x:
z + o y = (m/(m + n)) a (x + o)^((m + n)/n)
Expand the right side using the binomial theorem, subtract the z-expression from both sides, divide both sides by o, then drop any terms on the right containing o ... and you get y = a x^(m/n). That's pretty algebraic, right?But the incredible thing is that he's basically showing that (roughly) "the rate of change of the area under the y-curve is y", which is the Fundamental Theorem of Calculus! He's figured out that derivatives (rates of change) and integrals (areas under curves) are somehow inverses (at least in this case). Newton was pretty amazing!
[1] Carl Boyer, "The History of the Calculus and Its Conceptual Development". Dover Publications, 1949.
The paper explains:
> Most calculus students glaze over the notation for higher derivatives, and few, if any, books bother to give any reasons behind what the notation means. It's important to go back and consider why the notation is what it is, and what the pieces are supposed to represent.
> In modern calculus, the derivative is always taken with respect to some variable. However, this is not strictly required, as the differential operation can be used in a context-free manner. The processes of taking a differential and solving for a derivative (i.e., some ratio of differentials) can be separated out into logically separate operations.
> In such an operation, instead of doing d/dx (taking the derivative with respect to the variable x), one would separate out performing the differential and dividing by dx as separate steps. Originally, in the Leibnizian conception of the differential, one did not even bother solving for derivatives, as they made little sense from the original geometric construction of them.
> For a simple example, the differential of x^3 can be found using a basic differential operator such that d(x^3) = 3x^2 dx. The derivative is simply the differential divided by dx. This would yield d(x^3)/dx = 3x^2.
Is this significantly different than what is normally taught in schools? Using the notation described here is the first time it has felt like a tool rather than a formality to me, and it's quite different than the way I was thinking while taking calculus in high school. I'm not really sure how unusual this notation is though?
Given a field W, you will want to talk about various derivatives of the field. However, you might not be interested in the derivative of the fields with respect to coordinate fields like `\partial W / \partial x^i`, you might be interested in derivatives of that field with respect to some other field like `\partial W / \partial \lambda_i`. Then you end up introducing all kinds of auxiliary notation for the functional representations of the field. You end up with `\tilde{W}(\lambda_i)` and `\bar{W}(I_i)` next to `W(x_i)`.
Or just consider changes of variable, one generally doesn't distinguish the electric field `E` from cartesian, to axisymmetric, to spherical to whatever other coordinate system you dream up. Instead, the presence of the chosen coordinate symbols denotes the coordinate system in use and you don't need to distinguish the field from its functional representation. In order to understand `\partial_2 E` you would need to understand what coordinates I am working with while `\partial E / \partial y` is a bit more self explanatory in classical electrodynamics. Thermodynamics picks up confusion because of the constant changes in variable when the variables aren't coordinates.
While we're at it, it's probably not the best idea to represent derivatives as fractions either (for the sake of notational consistency). But that notation will never die.
Do you mean dy/dx? Why isn't that a good idea? Isn't fractional form useful, for example solving differential equation y = dy/dx => dy/y = dx? What are the alternative you'd prefer: prime notation, D-notation, or else?
I think this notation is single handly the reason why i've never been comfortable with calculus.
PS: i've stumbled a few years ago on a math book that described the original concept of "infinitesimals" and how a whole different way of doing calculus exists. And it seems to me those kinds of computations over "dx" come from there. But the end result of mixing concepts really looks like trash.
You will be delighted to discover that they are in fact not magical or garbage.
> What object is "dx" ? is it a number ? a limit ? is it zero ? can i divide another number by it ?
dx is a differential one-form. You can think of it as a generalisation of a gradient, if you like. These are very important in Differential Geometry.
You can use differential forms to do all sorts of things, but one example you may be familiar with is to compute area or volume forms over arbitrary manifolds. It gets a bit hard to define things on HN without TeX support, but using differential one-forms and the related exterior derivative, you can define a generalised Stokes' theorem that works for any smooth, oriented manifold.
I used this in my PhD, and implemented it directly in a numerical method, so this has very practical engineering uses also.
"Elementary Differential Geometry" by Barrett O'Neill is a pretty beginner-friendly introduction to some of these topics if you're interested, though there are many other good texts also.
This really doesn't help beginners. At all.
There are formal contexts where we can reinterpret division by zero and have it make sense. Should I start telling students that division by zero is allowed? Should I start teaching intro calculus students that 1+2+3+...=-1/12?
To some extent we have to speak to our audience. I consider that part of effective communication. I don't think "assume the person you're speaking to is/will be a mathematician" is an effective way to interact.
If you can come up with a more helpful reply in as many words, then please do so.
Δy = f(x + Δx) - f(x) and dy = f'(x) dx.
Then f'(x) = dy/dx. This may look like a stupid hack to make the last formula work, but actually it's a little more. If you use nonstandard analysis, you define the derivative of a function f from reals to reals by f'(a) = st( (f(a + Δx) - f(a)) / Δx )
where st takes the standard part of a hyperreal number and Δx is a nonzero infinitesimal. This is like the usual limit definition, without limits. Then you can use the formulas above and "dy" and "dx" are numbers, albeit hyperreal numbers.(The "dx as a differential form" vs. "dx as a number" is probably coming from the fact that the tangent space to the reals at a real number is isomorphic to the reals, so the dual space [where dx lives] is too.)
(Calculus via infinitesimals is pretty cool; a good resource for this is H. Jerome Keisler's "Elementary Calculus" and "Foundations of Infinitesimal Calculus", both available for free: https://people.math.wisc.edu/~hkeisler/)
I second the recommendation for Barrett O'Neill's book - I used it in my differential geometry class at MIT.
To measure the speed of a moving object you must divide the distance moved by the time it took to move that distance.
So how can you measure what the speed is at a given location? In a sense you cannot, you can only measure it at a given interval over the period of time it took to move that distance.
So it is kind of confusing. dx/dy represents the limit of measuring the speed over increasingly small distances and durations around a given point in space and time. If you take dx to 0 and dy to 0 it does not make sense because 0/0 is ill-defined. Therefore we need some notation that implies we are really not talking about a single point, but an increasingly small distance, and duration.
Was this book "Elementary Calculus: An Infinitesimal Approach", by Keisler? It's an awesome book. It's free to download at https://people.math.wisc.edu/~hkeisler/calc.html
Unfortunately, that realization came 25 years too late.
Vectors is a great example: as soon as you're introduced to vectors, you immediately starts to be given definitions on how to multiply / add them together and with regular numbers.
dx remained a mystery even during my first 2 years of calculus in university. I used them purely as a notation tool, but really didn't understand them properly.
For example:
> Isn't fractional form useful, for example solving differential equation y = dy/dx => dy/y = dx?
What exactly does "dy/y = dx" mean? What is on the LHS and what is on the RHS?
It acts like a mnemonic scribble for an intermediate step. It doesn't have any mathematical meaning.
That's about right. When you cover elementary solution methods for differential equations, you start with something like y = dy/dx. You're supposed to separate variables ("get the x's on one side and the y's on the other, then integrate"). So it's tempting to just write "dy/y = dx", even though as you say it doesn't have any mathematical meaning. But it's helpful in keeping track of the algebra. You then forget you wrote that meaningless but helpful step and write "∫ dy/y = ∫ dx" which is okay, and go from there.
Looking at one of my old diff eq books I see whole sections where this kind of casual algebra with differentials is the norm.
When anyone would ask me about this in class, I'd say something like this. Think of a solution curve for y = dy/dx as a parametrized curve, so x = f(t) and y = g(t). Then interpret the equation as y = (dy/dt)/(dx/dt), write dx/dt = (1/y) (dy/dt), then integrate both sides with respect to t:
∫ 1/y (dy/dt) dt = ∫ (dx/dt) dt
Change variables to get "∫ dy/y = ∫ dx". After a while we believe that this will always work, and we just suppress the stuff about parametric equations.As I understand, dy/y = dx means that the derivative of 1/y with respect to y equals to the derivative of 1 with respect to x.
So you must be saying that the derivative of 1/f(x) with respect to x is zero for any f(x), where f is a differentiable function, with non-vanishing derivative near x (for it to be defined in the first place).
That doesn't make any sense.
Please don't respond to this. It's getting absurd.
f : R -> R
that is, f having a type signature as a function, and f(x) : R
having a type signature of a float (or: being a float). And so on.I've been using this approach to type-check calculus equations. Unfortunately the article in the top comment [0] says that
> Note that f means something different on the two sides of the equation!
so the Leibniz notation might be more ad-hoc than I thought, and I can't just type-check them, actually have to reverse engineer the intent of the authors. I have to think about this. For example I remember someone giving me 3 exercises from Stewart calculus, and I gave 2 back that those equations don't even type-check, maybe there are some notation abuses I am not aware of.
[0] https://mitp-content-server.mit.edu/books/content/sectbyfn/b...
I agree with this part. I think in teaching, the precise notation should be taught first and everyone should understand it well. But later on, it will make communication and even just doing calculations easier if you can accept the imprecise notation, use it and translate it into something clearer when needed. This is particularly handy in e.g. differential equations where sloppy notation makes certain manipulations a bit easier.
Let M be a differential manifold, eg M = ℝ² and φ a chart, eg cartesian coordinates φ: p ↦(x(p), y(p)) where x,y: M → ℝ. Then, ∂/∂x denotes the holonomic vector field tangent to the coordinate lines t ↦ φ⁻¹(x(p) + t, y(p)) through any p ∈ M.
It is convenient to identify vectors and directional derivatives (this is in fact one possible way to define tangent vectors on manifolds), which for a function f: M → ℝ yields
(∂f/∂x)(p) = ∂/∂x|ₚ f = lim_{h → 0} ( f(φ⁻¹(x(p) + h, y(p))) - f(p) ) / h
https://mitp-content-server.mit.edu/books/content/sectbyfn/b...
They also adopt a notation where partial derivatives are taken with respect to "argument slots".
Is the a recommended vim environment for running the scheme in that book? I am usual to vim and slimv via https://susam.net/blog/lisp-in-vim.html#get-started-with-sli....
Really not emacs. Racket environment is a bit confusing.
I mean, they even fuck it up in the sentence after their equation:
> The Lagrangian L is a real-valued function of time t, coordinates x, and velocities v; the value is L(t, x, v). Partial derivatives are indicated as derivatives of functions with respect to particular argument positions; ∂_2 L indicates the function obtained by taking the partial derivative of the Lagrangian function L with respect to the velocity argument position.
From what they say at the beginning the second argument is the position, not the velocity, so ∂_2 L is ∂_x L, not ∂_v L. Which is a really easy mistake to make because every single equation needs to go back to the definition of L to make sense.
One could add some new notation to distinguish the element _T_ from the slot _T_. For example let’s write slots as _[T]_ (I’m not creative enough to come up with something good), then we can talk about ∂_[T]F or ∂F/∂[T].
One can think of the [.] operator as mapping from “semantic symbol” to slot number in a way.
But I do agree with your main point that the order of arguments is irrelevant, and it is a mistake to make it a first-class citizen of the notation.
The first case is where you have a N-dimensional vector space where all dimensions have the same units. The standard example would be the Newtonian 3D space. Depending on what you are trying to do, you can view it as a collection of coordinate-free abstract vectors, as a triple (x, y, z) of real numbers, or as an array X[i] of three coordinates in a given basis. In this case I would agree that X[0], X[1], X[2] is better than (x, y, z), the order matters, and you can define the Jacobian is a 2D array that represents a certain abstract derivative in a given coordinate system. I would argue that the formalism of Sussman and Wisdom (which they got from Spivak) is totally adequate to this case, and perhaps even the best possible.
The second case is the one of the Lagrangian that parent mentioned, where L is a function of the triple (t, x, v). You could pretend that (t, x, v) form a vector space, but this definition won't get you far. I would regard (t, x, v) = t * (1, 0, 0) + x * (0, 1, 0) + v * (0, 0, 1) as meaningless because it is adding time, space, and velocity. You cannot really do rotations or general linear transformations in this space. You can define a Jacobian matrix if you want, but now all entries in the matrix have different physical units. In this case I would say the fact that v is the third element of the tuple is irrelevant, and that the tuple is better regarded as a map from symbolic names "t", "x", and "v" to real numbers. I would argue that the Spivak formalism is inadequate in this case, and it seems that many physicists on this thread think the same for essentially the same reason.
This difference is kind of analogous to double X[3]; vs struct { double t; double x; double v; }; From one point of view they are the same, but in practice they have totally different meanings.
Dammit yes, you’re right! Well, it’s not a bit less confusing.
The most frustrating is that they have a point: we need to be stricter about disambiguating functions and numbers, and derivation really should be an operator.
But you don’t need to go all the way to zero-indexing (which is definitely not a thing in the fields I know) or positional arguments. This is putting abstract notation purity above practical concerns. It’s not surprising they like Scheme.
Yes, it's definitely a perspective influenced strongly by computer science.
> It’s not surprising they like Scheme.
In fact one of the authors, Gerry Sussman, is one of the original inventors of Scheme.
(I admit it might be a lot easier to avoid losing track of which thing is a row and which is a column in a gnarly linear algebra expression if all dimensions were explicitly named, and this would come with a tradeoff of verbosity. Also, the interpretation of a matrix as a linear function from vectors to vectors would need some clarification as to which dimension is input and which is output, so maybe it would look a bit like Einstein notation with superscript dimensions and subscript dimensions?)
Wtf is
> the partial derivation of a function f(x,y,z...)
or deriving a function respect to a variable? Functions don't have variables, named variables, that's Python, not math[0]. In math expressions can have free variables, and function arguments are indexed with numbers. You can derive x+y by x, or you can derive f : (x,y) |-> x+y by its first variable, but you can't mix them. That only leads to things as df(x,x)/dx and such abominations.
Multivariable calculus people really should just adopt the function notation of the rest of the mathematics fr.
[0] : https://en.wikipedia.org/wiki/Function_(mathematics)#Definit...
No, Python named arguments are a good analogy for Leibnitz notation.
Despite this, I still believe Leibniz notation is superior for multi-argument functions. For multi argument functions, named arguments are much clearer than just depending on the order of arguments. Essentially I advocate for dropping point-freeness for clarity on the difference between function arguments.
Besides that, by expression equivalence it is clearly the case that the only correct interpretation of the derivative is the 'composition' where the change in t also counts for the change in x. Because replacing the f(x(t), t) with g(x) (where g is the composition, should not change the outcome of the derivative.
If we did have real named arguments like f(a=x(t), b=t), where a and b are fixed by the definition of f rather than arbitrary names to be made up on each invocation, then maybe it would make sense to write something like (∂f/∂a)(a=x(t), b=t). Though that’s still pretty far from Leibniz notation, where ∂f/∂a is somehow supposed to be a numerical value depending on a, not a function to which arguments must be supplied.
But what I think is really lacking in mathematical notation is explicit lambda abstraction: we should be able to write
(λt. f(x(t), t))'(t) = (λa. f(a, t))'(x(t)) x'(t) + (λb. f(x(t), b))'(t),
which reuses the ordinary one-variable derivative and has none of these ambiguities.
For me the two notations are a lot like positional argument ordering vs named function arguments in a language like python which offers both. I personally prefer Legrange’s notation most of the time but think they both have their place. Legrange notation emphasises the ordering of the arguments and makes their naming going into the function irrelevant. This makes a lot of sense to me in a lot of situations where you have some pure abstract function with abstract arguments and makes it a lot easier to reason about certain things (like the example given or the chain rule which the author uses as an example where the common Leibniz formulation literally only makes sense if you do an explicit substitution in the function which is often not really explained).
On the other hand Leibniz notation makes a lot of sense to me in contexts like physics where the names of the function arguments are really important. You aren’t just picking your favourites as the author claims. For a very simple example, if I say a = \frac{d^2s}{dt^2}, every physicist knows a bunch. I am calculating accelleration as the second derivative of displacement with respect to time. “s” isn’t just my favourite letter today. So “named arguments” make a lot of sense here.
Secondly, Leibniz notation gives us the “derivative operator”[1], which is often very useful. Say I’m trying to find some derivative that involves a bunch of intermediate working out steps. It’s tremendously convenient to be able to say “<something> times d by dx of <some other expression>” and come back and differentiate “some other expression” in a later step. In Lagrange notation I would need to let that second expression be some named intermediate function so I have something to hang my tick off, which is definitely a lot more of a pain.
[1] I’m not far enough along in my calculus journey to know whether I’m saying this right so forgive me if my terminology is incorrect - hopefully you get the intent.
The nice thing about Leibnitz calculus [1] is that you can do things like reduce fractions with dx'es, you can flip things around and calculate dx/df, the chain rule is not a rule that you have to memorize but just an obvious expansion etc., and it mostly just works. I don't recall seeing an explicit proof why it works (except for some specific cases), or a list of exact rules, but I'm sure that exists and it would have been neat to have seen that in my studies.
[1] calculus here in the sense of German Kalkül, a notational system and a method of mechanically manipulating symbols, and not neccessarily meaning "differential and integral calculation", although "the" calculus is the prime example of "a calculus" of course.
1. "You often deal with functions that are dependent on multiple variables. Do I want to differentiate wrt. time or position?" -- the proposed solution is to put a subscript under the function name, to clarify whether you're differentiating wrt time, position, etc
2. "How can you distinguish between f(x) and f(t) which are very different functions?" -- the answer claims (and I agree) that using f(x) and f(t) to represent different functions is bad notation. If x and t are variables, surely f(some_variable) == f(other_variable). If x and t are specified values, then f(value1) may not equal f(value2), but f still _really really_ looks like the same function, just evaluated at different points. Better is to use different function names, like 'f' and 'g'.
3. "what if the variables depend on each other? You need a notation that distinguishes between treating "x" as a variable, and "x" as something that is dependent on some other variable" -- there is a proposed way to represent composition of functions. I don't want to dive into latex editing on hn, but you can see it in the post
4. "The nice thing about Leibnitz calculus [1] is that you can do things like reduce fractions with dx'es, you can flip things around and calculate dx/df, the chain rule is not a rule that you have to memorize but just an obvious expansion etc., and it mostly just works" -- it does indeed work sometimes, but this is not rigorous. When it works, it does so because you're using it in a domain where it just happens to work. "Cancelling" dx's is not reliable, and will sometimes lead to error. I'll admit I find the chain rule mnemonic convenient though
"the concept of a function shouldn't depend on what your favourite letter is!"
Very helpful answer, thanks for posting.On the other hand, you can also argue that the concept of a function shouldn't depend on your favourite ordering of its variables. Thus, if you have a potential that oscillates in time, such as:
V = (1 + sin(t))/sqrt(x^2+y^2)
You may want to take derivatives with respect to each variable, such as dV/dt, dV/dx and so on. Their meaning is clear. What does D_2(V) mean, though? It does depend on your favourite ordering of the letters!Even if you prefer positional notation for derivatives (à la Spivak), it must be recognized that the naming-based notation has its merits, and sometimes is clearer.
V : ℝ³ → ℝ
V(t,x,y) = (1 + sin(t))/sqrt(x^2+y^2)
then it’s clear what D_2(V) means.The term of art is "an expression" [0]. And you can also differentiate expressions with respect to their variables. It's a perfectly supported construction in all symbolic computer algebra packages. No need to assign an (arbitrary) ordering to your variables in order to differentiate with respect to them.
f(x) = x^2 vs. f(t) = t^2
They are, because they produce the same set of ordered pairs (input, output).Some of the confusion is the fault of math profs, because we get lazy. We could say "the function f: R -> R defined by f(x) = x^2", but we wind up saying "the function f(x)" (wrong - f is the function, f(x) is the value when x is plugged in) or worse "the function x^2" (really wrong). But if you always say the absolutely correct thing "the function f: R -> R defined by f(x) = x^2" you sound really pedantic, and the extra words make it hard for students to comprehend. Oh well.
What was more confusing to people was if you had f(t) = t^2 and people thought "t" had some sort of independent existence outside the definition of f. Some guy was asking me about it and I was trying various explanations and he wasn't getting it. Then I remembered he was a CS major and I said "t is an argument to f and it's local to the function block" and he got it and nodded his head.
So I wonder if these kinds of distinctions are easier for programmers than for typical math students - in the sense that programmers would know there's no difference between these:
int f(int x)
{
return x * x;
}
int f(int t)
{
return t * t;
} int f(int x);
int f(int t);
will compile with no errors, while changing t to, f.i., "float" obviously won't.
I wonder whether this could be pushed farther to come up with programmer-friendly descriptions of things like the chain rule starting from something like: int h(int x) {
return f(g(x));
}
etc.You have ∫ f(g(x)) g'(x) dx. [So it might be something like ∫ [cos(x^2)] (2 x) dx, where f(x) = cos x and g(x) = x^2.] You do the u-substitution u = g(x), so u'(x) = g'(x):
∫ f(g(x)) g'(x) dx = ∫ f(u) u'(x) dx
You want to say that the second integral is ∫ f(u) du -- essentially (to go back to "fractional" derivative notation), you want to do (du/dx) dx = du to get rid of the x's. Why is this justified?Let F(u) = ∫ f(u) du: F is the antiderivative of f. By definition of the antiderivative, F'(u) = f(u). So by the Chain Rule,
d/dx F(u) = F'(u) u'(x) = f(u) u'(x).
But since F(u) = ∫ f(u) du differentiates to f(u) u'(x), by definition F(u) is the antiderivative of f(u) u'(x). But the antiderivative of f(u) u'(x) is ∫ f(u) u'(x) dx. So putting all this stuff together, ∫ f(g(x)) g'(x) dx = ∫ f(u) u'(x) dx = F(u) = ∫ f(u) du.This looks like the difference between parameters of a function and arguments. In the definition, you have the parameter x, used internaly and, when calling the function you use an argument - located in the calling context. In python: def f(x): return 2*x ## x is a parameter x = 3 y = f(x) # x in an argument
For partial derivatives, it may be ok to use 1 and 2 to show they are the first and second parameter, but maybe they could also be named by some convention, like x and y or alpha and beta or whatever
The salient issue that the author of that answer seems not to have understood, which we here in the land of the Y Combinator should have less trouble with, is the distinction between free and bound variables.
In lambda calculus, variable binding is explicit and denoted by a lambda. Leibniz notation binds variables in just the same way: the operator (d/dx) binds the variable x in the expression to which it is applied.
Did you look at the posting history at all? I think the poster understands what it means.
I appreciate that especially for mathematicians and programmers, making a clean distinction between a function and its evaluation is a key conceptual point, and Leibniz notation obscures this fact. However, there are good reasons why physicists use Leibniz notation, and this answer really glosses over that.
The reason is that the particularities of the mathematical structures used to model a physical problem matter a lot less than the relationships between the actual physical underlying quantities. And there can be a lot of them. The answer evokes thermodynamics, and I couldn't think of a better example. Are we really to introduce a distinct symbol for _every_ possible functional relation between _every_ state variable? If I have T = f(P, V) and P = g(T, S), do I need to remember in which position exactly each function has been ordered before I can write down "how P varies with T when S is fixed"?
Leibniz notation, although formally tricksy, is just the best tool for communicating the _intent_ of a physical relationship without getting bogged down in the mathematically detail. The purpose of an equation is to express that relationship to the reader. Think if it like code — is it so bad if my code does a bit of magic behind the scenes to allow for a clearer reading, even if the semantics aren't immediately obvious? Well, ultimately, it depends on the situation, and a balance must be reached. I don't believe that expressing everything the way TFA suggests is striking the right balance.
Notational abuse happens all the time in physics, and this is certainly not the most egregious example. Just compare it to the path integral. It's easy to assume that this is because of a lack of sophistication or rigour by physicists. (Certainly I did throughout my physics education, being more mathematically or pedantically inclined.) But it's a simplistic view.
Now, while I'll defend the usage of these sort of unrigorous conventions even if they are strictly speaking meaningless mathematically, what I won't defend is the slapdash approach that is often used to _teach_ partial derivatives to physicists. Some exposure to concepts like distinguishing real variables/quantities from functions is needed, or, as TFA does mention at the end, the student won't be able to unpack the notational convenience into clear semantics, which can lead to unclear reasoning. I used to share the views of the author for a period when first introduced to Leibniz partial derivative notation in my first thermodynamics course, and, probably like them, found it to be totally incomprehensible symbol soup. But for myself, I see now that it was mostly a failure of teaching rather than a failure of the notation itself.
I'll add one last thought. There is a degree of "primitive obsession" at work here, trying to fit everything into positional functions and real numbers. I have thought that a formalism that better reflects the structure of "physical quantity" (as opposed to thinking of them as plain real numbers) may help bridge the gap between rigour and conceptual convenience. The tools are really already there. We need two key concepts: first, borrow from programming the idea of keyword arguments (there is a book which sadly I can't remember the name of which does as much to formalize Einstein notation for tensors in a coordinate-free way); second, to model quantities not as real numbers but as differentiable homomorphisms from a state space (modeled as a manifold in a coordinate-free way) to the reals. This is how physicists already think about it, it needs only be formalised.
Any chance you might be able to brainstorm an example or three of what a well known equation would look like in that formalism?
It really does feel like a hodgepodge of poorly fitting parts that could use some revision in terms of notation and teaching, but I feel like that's part of the charm.
I like that it encodes so much history and culture "up front". It reminds me of PHP. There is no way a good language designer would create PHP intentionally but as ugly as it was, it was brilliant in 1999. Index.php and you're done.
f(v_1, m_1, v_2, m_2)
For the derivatives I would prefer writing
d_v_2 f
instead of
d_3 f
when counting arguments from 1
I mean, that ship sailed a long time ago. You can't understand any modern math or physics (or ML, or CS) paper without depending on variable naming conventions. Attempts (like Sussman & Wisdom) to have a properly lexically scoped notation end up being quite verbose.