One way to think of this is to realize how easy it is to multiply numbers but how much more work it takes to divide numbers.
For something like automatic differentiation, you're essentially applying the chain rule for partial derivatives repeatedly. This is analytically pretty straightforward to do for most applications. All you need is an analytical derivative for the simple functions your more-complex function is comprised of (e.g. a neural network).
For integration, the analogue of the chain rule is integration by substitution [1]. The toolbox for solving integration problems is more limited than for differentiation. You run into issues where the answer cannot even be expressed using standard mathematical notation [2]. Sometimes you get lucky and the answer can be expressed via an alternating Taylor series so you can estimate the answer within some margin of error [3].
Stan is a piece of software that runs state-of-the-art MCMC methods to basically just compute fancy integrals. A Stan model will take an order of magnitude more time to run than a simple neural network via something like PyTorch on the same dataset. But they answer different questions.
[1] https://math.stackexchange.com/questions/1635949/is-there-a-...
[2] https://math.stackexchange.com/questions/1397132/why-cant-so...
[3] https://math.stackexchange.com/questions/145087/how-to-calcu...
When you input the variable values into a symbolic derivative you just get a value at the end. d/dx x^2 = 2x. If x = 0.5 then d/dx = 1. The same is true for symbolic integrals. For most practical applications, we don't really care about the full symbolic expression. We just want the answer, or at least a good approximation. This post uses a specific example of the difference between two Beta distributions. We want to get that 0.71. It is very hard to "automatically" make that happen.
Say we want the derivative of f(x) = x^3 - 2x^2 + 5. That becomes:
(x+e)^3 - 2(x+e)^2 + 5
= x^3 + 3x^2e - 2x^2 -4xe + 5
= (x^3 - 2x^2 + 5) + (3x^2 - 4x)e
The term in front of 'e' is "3x^2 - 4x", which is f'(x).
f'(x) = lim_{e->0} (f(x+e) - f(x))/e
If you're able to express f(x+e) on the form (f(x) + y e) then it follows that y is the derivative.
It also should be noted that auto-diff doesn't let you skip the rules for derivation and you're using the same calculation as you would to show that e.g. f'(x^n)=n*x^{n-1}.
But what about something like cos(x). Either you can lazily evaluate the power series or you know its sin(x).
What's the integral of exp( sqrt ( 1 + (tan^(3/2) X)2 ) ) ) ?
We only know a handful of forms that can be integrated in closed form and its down to our creativity to discover new forms that can be integrated (same deal with solving differential equations and the reasons are the same).
The forms that we know how to integrate can be done by a computer. CAS tools will do that for you. For example Mathematica.
Finding an "auto-integrate method" would probably involve finding a way of calculating the integral in a decomposable way, and that indeed would be amazing, but I don't really see that happening any time soon.
But then again, it's not clear what the OP meant by "automatic integration" anyway.
[0] https://newbooksnetwork.com/david-bressoud-calculus-reordere...
on why integration is a fundamentally harder problem than differentiation. New techniques would be required to programatically analytically solve integrals as well as current differentiation programs, but it would be exciting to see.
The problem is that integration suffers from the curse of dimensionality much worse than differentiation.
Differentiation is a local activity. You approximate a derivative near any point. Integration is non-local, you must aggregate a function over a span of points in the domain space.
As the dimensions get higher, differentiation merely must apply local approximation to each successive dimension of a point in a grid. It only adds the computation cost.
Integration requires exponentially more data (for a given level of accuracy) because adding a new dimension enlarges the entire integration domain space that must be ranged over, and worse you generally need all that new “volume” of the enlarged space to be covered evenly to avoid bias.
So there are some fundamental asymmetries between differentiation and integration that play a role in why “automatic integration” isn’t as straightforwardly feasible as for differentiation.
Usually one cannot afford to visit all of the space. One visits high density regions preferentially.
https://math.stackexchange.com/questions/1635949/is-there-a-...
The use of 'would' implies that automatic differentiation has not been established. Also, if something is 40+ years old, then it follows that its foundation is even older.
Integration is well behaved numerically, but poorly behaved algebraically.
Differentiation is poorly behaved numerically, but well behaved algebraically.
Therefore if one wants to do "automatic integration", one has to approach it in a radically different way to autodiff. And arguably, such a method does already exist; it's just very inefficient.
> Differentiation is poorly behaved numerically, but well behaved algebraically.
This is so pithy and true, am stealing that.
The problem is that it's unknown if a symbolic equality algorithm exists for elementary functions and what those functions should be.
It is known that the general case is undecidable: https://en.wikipedia.org/wiki/Richardson%27s_theorem
For integration we need to prove that such a object exists for each expression we want to integrate before we try integrating. Which as far as I'm aware is a problem equivalent to that solved by the Risch algorithm.
Can't we just integrate away and check if it's correct by differentiating it back?
EDIT: oh, I guess that's why we need to be able to check for function equality :)
[1] - https://en.wikipedia.org/wiki/Computable_analysis#Basic_resu...