Is the Kalman filter just a low-pass filter?
jbconsulting.substack.com
jbconsulting.substack.com
* You can think of it as a Bayesian update process for linear Gaussian systems. That is: given a prior belief of the state of a system (and an uncertainty about that belief), and a measurement about the system (and uncertainty about that measurement), the Kalman filter tells you how to combine the prior with the measurement. This is very hard to do in general, but has an exact solution if your system is Linear-Gaussian. That's magical!
* You can also think of it as a "better way to average". If I gave you two quantities that reflected some "true" value and asked you what the true value was, you would probably average them. The Kalman filter does you one better, because it tells you to average the two quantities weighted by how confident you feel about each one.
* If you like control theory, you can think of the the Kalman filter as the dual of the Linear-Quadratic Regulator. That is, the KF is the optimal state estimator for Linear Gaussian systems in the same way that the LQR is the optimal (minimum cost) controller for LG systems. It's also worth pointing out that if the system you are estimating is being controlled, the KF can incorporate control inputs as well!
linear models meaning anything that can be described as a matrix vector multiply. (which is a lot)
but i'm no expert.
It bugs me when people use the word "optimal" in the Gaussian / Bayesian formulation. As the top-level comment above says, if you assume the various prior and conditional distributions are Gaussian then the posterior distribution is Gaussian too. This is not optimal, it's exact, just like you wouldn't say x=2 is optimal solution to x+1=3.
It is the optimal solution in the quadratic optimisation formulation, as the top-level comment also correctly said.
I though optimal conveyed the idea of "literally the best possible solution but you're still in the presence of a fully random system here".
Which might be the wrong interpretation, but hopefully it explains why some people (who aren't necessarily familiar with rigorous mathematics) use optimal.
In any case, what I'm trying to get at there is that in estimator theory there is a concept of optimality for an estimator over a distribution.
But an estimator estimates an unknown parameter from data (or in such a trivial estimator possibly without data) - and I believe this is central to the confusion.
As pointed out elsewhere in this thread, demonstrating a Kalman filter with only one input doesn’t really show their real potential.
(Or is Kalman filter defined to only work for Gaussian distributions anyway, and otherwise you call it Bayesian updating?)
I mostly know KFs from the third approach, control theory, state observers, state space representation, where the model is a central concept.
Can you actually use KF without a model? I'm curious, would you point out some references?
The Kalman filter is "just" a recipe to calculate the weights for optimal estimation, given different model assumptions. The difficult part is not applying the formulas, but mostly in coming up with a good model of the movements of the things you want to measure/estimate.
If you merely take the mean of the last 10 depth gauge readings, your smoothed depth reading will always lag behind the true depth, being about 5 readings out-of-date.
By fusing together the noisy depth gauge, and the noisy flow-rate meter, and a model saying how fast depth rises with flow, you can average out the noise without creating the same level of lag.
This is useful in applications like GPS receivers - there's noise so you do need filtering, but for driving through complex junctions, the last thing you want is a 5 second delay!
* The Kalman filter can account for changes in state. This is useful the value you're measuring is changing over time (e.g. the position of a moving vehicle, or water level in a bucket being filled mentioned in another comment). The type of variation that can be handled by the Kalman filter will depend on what assumptions you make - how you configure it, if you like. Often you would account for velocity, but you can go a step further and account for acceleration too. A moving average will always lag behind a bit, even when movement is totally linear and there is no error in the measurement at all.
* A Kalman filter effectively gives more weight to recent measurements and less weight to older ones. e.g. a moving average with period 4 will have weights of (..., 0, 0, 0.25, 0.25, 0.25, 0.25), whereas a Kalman filter might effectively have weights of (..., 0.125, 0.25, 0.5). You can even have a continuously-varying time gap between measurements, so if a new measurement comes in then it will update the estimate a lot if the previous measurement was a long time ago, whereas if the previous measurement was extremely recent then it will effectively be averaged with the new one. One downside of this is that if the state changes a lot very suddenly then the Kalman filter will remember the old state, to some extent, whereas a moving average will forget it entirely once it drops out of the window; but this is offset to some extent by the previous point.
If you don't have a model of the various confidences then you'll need to guess them - if you do a really bad job then it might work out worse than a moving average. But if the Kalman filter's working better then you can always make use of the state estimate while ignoring the computed uncertainty.
μₖ = μₖ₋₁
as the model. If your model assumes that the process just stays constant, then all you are left with is filtering the noise.
My goal with this article is only to help readers who have wondered how time-domain approaches like Kalman filtering relate to frequency-domain approaches like low-pass filtering, and to connect those dots. It's not a comprehensive article about the magic of Kalman filters.
I would contest that the reason a frequency filter doesn't work on a gyroscope is not because it's a MIMO system (which frequency domain techniques can generalize to; you just end up with n x m transfer functions) but because the system is not static, so it doesn't satisfy the conditions that would cause the Kalman filter to converge to a fixed-coefficient filter.
The MIMO distinction is also not super important. Since you can also have MIMO low pass filters. The real difficulty is obtaining the coefficients of these filters and hopefully find the best ones. That's where you start getting into the optimality and the real contribution of these tools. But as a side-effect you must assume that the noise is Gaussian otherwise you lose much of the niceties of the theoretical guarantees. This is basically the biggest control theoretical disadvantage of kalman filters and the reason why other domains keep rediscovering it while in control it is "kinda, sorta" falling out of grace.
in fact, in the static case, a kalman filter is exactly the same as the recursive least squares estimator, which is provably the optimal unbiased estimator
But the overall subject of comparison to just a low pass filter is very interesting
In one of my other articles where I applied the Kalman filter to a more complex system, a frequency-domain filter would have failed. But in this article I tried to stick to the simplest possible system, so that the math wouldn't get in the way of illustrating the way in which Kalman and Wiener filters relate to each other, since it's an interesting connection that doesn't get much attention.
So, when the quality of measurements is very high (almost no noise), then the motion model provides very little benefit. And when the measurement quality is very low, the motion model reduces the error ellipse around the real state of the system.
So, even when there are inputs and a motion model you could imagine the Kalman Filter as a frequency-domain filter (E.G. IIR/FIR filter), where the tap weights are being updated based on some known function.
If the quality of the measurements are high enough, and the sampling frequency is high enough. Then the Kalman filter should be able to converge to a very good estimate of state, even with no motion model right? In which case it would have the exact same characteristics as an adaptive, slowly time-varying frequency-domain filter.
No. Consider this discrete-time linear system with 2-dimensional state [x1, x2] and one output y:
x1' = 2 * x1
x2' = x2
y = x1 + x2
Suppose we see the output sequence 1, 2, 4, 8, ... If we know the motion model, then we know that the initial state must have been [1, 0]. (In technical terms, the pair (A, C) is observable.) But if we don't have the motion model -- all we know is that y = x1 + x2 -- there's no hope of ever getting a state estimate, even if our y measurement has zero noise.The figure does show confidence bands in pale blue, although they're hard to see because the data points kind of obscure them. But since the covariance deterministically converges to a constant within a few samples, regardless of the data, the confidence bands are not very interesting in the example in the article.
Meta comment: I really appreciate that the author is taking the time to reply to people with queries (not this post). Thanks! It makes this community better.
E.g. here is someone's project for removing reverberation from speech, with sound sample files for A/B comparison:
https://github.com/helianvine/kalman-filter-based-MCLP
You cannot do that with just low/high/bandpass.
Same applies here, single input with a trivial model, kalman is just a lowpass filter.
In the frequency domain, that's equivalent to multiplying one Gaussian by another one with zero mean, which will always put higher weight on low frequencies. No matter what the gain is in the Kalman filter, that'll still be true as far as I can tell. As the gain varies, the cutoff within the low-pass filter will change though.
It gets harder to analyze when you start using non-linear models to update position though. Generally, I think the same logic applies.
If you pass a moving probe through a magnetic field your meaured results vary by direction (with or perpendicular to flux) and velocity.
If you want to measure a geomagnetic field using an aircraft, the heading matters.
To normalise in post processing you calibrate a Kalman filter by flying precessing butterfly wing patterns in a known relatively level flux area and then use that to remove the magnetic signature of the aircraft and heading from data collected over multiple headings and days.
( There are a few other twists - diurnal flux and induced field from the Earths GMF interactions needs to be isolated, etc ).
Was that a pun?
As an addendum, I want to say a well designed whitening matrix + IIR filter can replace a Kalman filter, again depending on the application. Just makes things easier to understand, debug etc. Works if your vector is somehow decomposable into roughly independent scalars.
And a more technical introduction based on the article above: https://vanhunteradams.com/Estimation/Estimation.html
Explicit noise regularized (or variance regularized at least) filters do better especially in case the signal is corrupted by non-gaussian noise where Kalman filter can get divergent.