The kalman filter tries to guess the hidden input that produced the measurements. It does so forming the minimization problem:
'minimize over x, the function [ actual_measurement - expected_measurement(x) ]^2/s^2', here 's' is sigma of noise.
This follows from the state estimation problem:
'maximize over x, the likelihood of seeing the actual_measurement', because the only term that matters in the likelihood function is -([x-expected(x)]/s )^2. (look at the exponent in the Normal distribution, or any exponential distribution really).
'actual_measurement' is a constant, so if it happens that the function 'expected_measurement' is linear, this is trivially solved directly as a convex optimization, and if you take derivative, equate to zero, and solve, you'll get the kalman filter update step.
If it so happens that the function is non-linear, well we just make a single netwton-rhapson step by linearizing the equation, minimizing, and returning the solution to the "pretend linearization".
This is basic calc + linear algebra at an undergrad level, but nobody bothered to tell you that.
---
It's also completely wrong. It's a hack from the 60s to maximize the likelihood function using a recursive, single-step linearization like this. A misreading of the Cramer Rao Lower Bound has "proven" to generations of engineers that this is optimal. It's not, not really.[^1]
Nowadays we have 10,000x more compute, and any one of the following _will_ produce better performance:
* Forming and solving the non-linear equation using many newton-rhapson steps
* Keeping a long history of measurements, and solving using many newton-rhapson steps over this batch
* Using sum-of-guassian representation to accomodate multi-modal measurement functions, esp when including the prior bullets
All of these were well covered by state estimation research from 80s to now, but again, the textbooks seem to be written in stone in 1972.
[^1]: (the cramer rao lower bound is only defined when all measurement likelihood functions are linearized at the true state - which is only possible asymptotically in a batch which preserves all the measurements - and not possible before time infinity and not possible with recursive filter)