A Little Calculus
papl.cs.brown.edu
papl.cs.brown.edu
https://papl.cs.brown.edu/2018/index.html
https://papl.cs.brown.edu/2019/
https://papl.cs.brown.edu/2020/
https://dcic-world.org/2023-02-21/func-as-data.html#%28part....
In the past I've read that that stuff is more or less just a shorthand for the "formal" definition of a derivative. But I have also seen people solve problems through manipulation of the dx's and the dy's, and treading that as... well as a fraction.
Is there some sort of guide to what you're "allowed" to do with dx and dy in general? Or is the thing basically that people are just "being careful"? This, like discussion of set theory, is always tough because I feel like I'm missing the axioms that could justify certain operations
dy = 2x dx
dy/dx = 2x
Chain rule makes this not completely trivial formulation.
d/dt y(t) = ( d/dx y(x) )( d/dt x(t) )
The equivalent to i = sqrt(-1) might be something like the nilpotent infinitesimal where epsilon^2 = 0 but epsilon =/= 0.
What usually is going on is you mentally replace dx and dy with Delta x and Delta y = y(x + Delta x) - y(x), where Delta x is some small nonzero quantity. After doing your manipulations, you take the limit as Delta x -> 0. Then you have rigorous statements like Delta y / Delta x -> f'(x) (this is the definition of derivative), and Delta x/Delta x = 1, and (Delta x)^2/Delta x -> 0 (ability to neglect second order terms), etc.
You do this type of thing so many times and it becomes rote and annoying to explicitly mention the limits and the reader is trusted to formalize it themselves as an exercise.
The stuff that you are "allowed" to do is taught in a first course in real analysis. After taking such a course, you will be able to justify for yourself which manipulations are valid.
https://en.wikipedia.org/wiki/Differential_form
Unfortunately I don't know a treatment of this subject that explains how to apply this concept to basic calculus without introducing some other more difficult concepts.
That said, isn't calculus, mainly the chain rule, enough to tell you what operations are allowed?
For any expression y involving x [*], denote by dy the linear part of y|x+t - y|x (that is, y evaluated at x+t less y evaluated at x). For y=f(x), that’s just a fancy way of saying f'(x)t, of course, but the intent is that where y was an “x-dependent scalar”, dy is an “x-dependent linear function” (without a constant term, as it is the convention in most settings outside of high school).
The baby version stops at that: by our rules, dx = t, df(x) = f'(x)t = f'(x)dx, f'(x) = df(x)/dx, the not-a-proof for the chain rule becomes an actual proof except everything is assumed differentiable (where a textbook one would only need differentiable inputs), etc. Of course, at this stage it seems slightly miraculous that nothing ends up t-dependent, but it is what it is. (If you want, you can imagine that the symbol t is “private” to the previous paragraph so it’s not allowed to escape to “the user”, but then you need to prove that it actually doesn’t.)
The adolescent version unholsters linear algebra:
While of course every one-dimensional (real) vector space (i.e. a line with a chosen zero point) is R in disguise, there are multiple choices for what the disguise is (differing by a multiplication by a constant), and you might not know which one to prefer. (This is what choosing a basis in a one-dimensional space amounts to.) Given such a space and two vectors u (whatever) and v (nonzero), denote by u/v the number such that u = (u/v)v (the coefficient of proportionality, aka the coordinate of v when e is the basis vector). Now you can prove, for example, that u/w = (u/v)(v/w), because it is so when you choose any (single-vector) basis and substitute for each vector its (only) coordinate.
(Can you make sense of uv/vw? Yes, as u⊗v / v⊗w, but tensor products are their own can of worms and probably overkill at this point.)
At each value of x, the space of linear functions (of the private variable t) is one-dimensional, so everything in the previous paragraph applies. When we write dy/dx and so on, we mean the things from there except we do them at each value of x (“pointwise”).
One of the things that this more advanced thinking gets you is that you can imagine how all of it generalizes to multiple variables. (Writing v/e_i for coordinates of v in the basis (e_i) is not common, but it does not not make sense—as long as you remember you can only “divide” by a basis, not by a single vector. Write out the coordinate transformation rules in this notation. The differential version will have you end up with df(x)/dx_i instead of the more common ∂f(x)/∂x_i, but again, that makes sense in context—note that, once again, the partial derivative wrt one coordinate on a plane depends on what the other coordinates on that plane are!)
The grown-up version just says I’ve been talking about the cotangent bundle in wishy-washy language. (An “x-dependent scalar”? What’s that? Does it taste good?) Hopefully it was still of some help.
[*] People do write df/dx where I would require df(x)/dx, and it’s convenient to do that, but I’m trying to avoid additional abuses of notation where possible.
I get the feeling it's not that used or popular outside of a pedagogical context though, whichis a shame since it has a lot of good ideas!
or people are still trying to find good ways to teach it.
> fun square(x :: Number) -> Number: x * x end
> fun double(x :: Number) -> Number: 2 * x end
No, that is not a nicer way of writing those. It certainly helps for understanding, but once you understood properly what is going on, writing `fun double(x :: Number) -> Number: 2 * x end` instead of `2x` is just not economical.
The /dx part indicates that the following expression should be considered functions of x.
The equation can also be rewritten as d F = F' dx, reflecting that the change in function value is proportional to both the derivate and change in x.
I'd like to argue, that when it comes to computers and differentiation, automatic differentiation is the most useful, most important in practice. But often only symbolic and numerical differentiation are mentioned.
Is this a misunderstanding? You use differential equation solvers when all you know is how to calculate the derivative(s) of the function, and you want to get the function itself as the solution. Where would you need numerical differentiation in this?
I think the idea is to use the "tangent linear model" to decide how much importance to give to a particular observation of the initial state.
What's nice about a discussion of symbolic differentiation is that we can prove a few rules rigorously, and then use those rules to purely mechanically differentiate algebraic expressions we encountered up to trigonometry.
You're right though, in practice, for complex functions expressed as programs, automatic differentiation is superior.
Was it correct? All I know is that I couldn't kill it.
Numerical differentiation comes directly from the definition of derivative and symbolic is what you always did for exercises in your calculus classes.
Automatic differentiation come in two flavours, forward and backward modes. Forward mode is based on dual numbers [1], which is the
quotient of a polynomial ring over the real numbers, by the principal ideal generated by the square on the indeterminate
That is ℝ[X]/≺X²≻.Another way of thinking of this is to have an element ℇ different from zero, such that ℇ ²= 0 and hand wave the fact that the dual part of the number carry the derivative.
Backward mode builds a graph of computation and doing symbolic differentiation over the graph and compile down the derivative into runable machine code (it could be interpreted, compiled down to IR, neither the form or the execution environment changes the fundamental algorithm).
Maybe they are not really hard, but they are not easy either. Still I think they should be at least mentioned in modern courses.
> Another way of thinking of this is to have an element ℇ different from zero, such that ℇ ²= 0 and hand wave the fact that the dual part of the number carry the derivative.
Aren't these the same way of thinking?
A more common example of this idea is √(-1): R[x]/<x²+1> vs i²+1=0.
That might make more sense in another course, like one on numerical methods, rather than in a programming course that happens to use numerical differentiation as a demonstration of programming techniques and language features.
So in some sense it's maybe not hard, though any attempt to do it without building up some theory (e.g. abstract vector spaces and dual spaces) first will probably come across as magic. On the other hand, magic tricks would be right at home in an intro differential equations class so maybe this would be a perfect addition. Or it can replace Wronskians or something.
A little trigonometry by my side
A little Fibonacci's all I need
A little inequality's what I see
A little bit of lambda in the sun
A little bit binary all night long
A little probability, here I am
A little √2 makes me your man