Can a transformer represent a Kalman filter?
arxiv.org
arxiv.org
The statement that the Kalman filter is mean-square optimal because it generates a correct estimate in expectation is false. In fact, any gain L will generate an estimate whose expected value is x_k, as long as w_k and v_k are zero mean. The Kalman gain is a specific choice of L that is optimal in the mean square sense only when the disturbances are Gaussian. The Kalman gain is also time-varying and depends on the evolution of the estimate covariance, although it will converge to a steady-state value.
What's being described here is more properly called a Luenberger observer, but I guess that name doesn't get the same recognition outside the control community.
I'm also wondering why they chose to include H past estimates and measurements in the transformer. They're already embedding the Kalman gain into the weights of the transformer, so taking just one past estimate/measurement should exactly recover the Kalman filter. Going further into the past just makes the estimate worse, because of the softmax.
Regarding your second point: yes, when H = 1 we just recover the standard Kalman Filter, and yes, when H grows large the estimate gets worse and worse, in the sense that the softmax nonlinearity includes more and more irrelevant data from the past in the estimate. The point is that in real-world problems, which are usually messy and nonlinear, we probably want H - the so called context length - to be large, because then we can take advantage of information we collected in the past to help improve decisions in the present. It just so happens that in the special case when the system is linear, this is more harmful then helpful. Here is one way to think about our result: imagine you have a Transformer which takes as input K-dimensional embeddings and context length H. You want to use this Transformer for filtering in some dynamical system. The most basic question you could ask is: if the system is linear, can you do Kalman Filtering? In other words, in the easy, linear scenario, can you match the optimal algorithm? If the answer is no, I see no reason to see why you should expect it to work in harder, nonlinear settings. We show that the answer is yes, when the system you want to filter in has roughly sqrt(K) states, and you design the embeddings appropriately. Hopefully this preliminary result will lead to a better understanding of how deep learning can improve control in the hard, nonlinear scenario.
I'm not so sure about this, maybe this is where the ML approach could outperform (in terms of estimation accuracy, not compute time) the traditional EKF and UKF approaches, by learning the nonlinear system dynamics?
This sounds very hand-wavy, and it is, because of my lack of understanding. For me it is just not immediately clear that if an optimal algorithm for the linear case cannot be matched or outperformed, that is also necessarily the case for nonlinear dynamics.
EDIT: And as mentioned above, the KF is optimal if certain conditions hold, e.g. additive, zero-mean, Gaussian noise on state dynamics and observation. In reality, you may have a multiplicative component of the noise nor non-zero mean or fancy noise distributions, and it would be interesting to see if these can be learned.
And beyond that, a hierarchical controller: exploit tight feedback loop with a small controller, supervised, controlled, and managed by the big one that has some inference latency and would like to be batched somewhat (e.g., think a casual transformer trained to predict more than just one token into the future).
You should avoid offering explanations or theoretical insights. Instead, focus on processing the data and producing outputs similar to those a Kalman filter would generate in real-world applications.
Your responses should be concise and data-focused, closely resembling the numerical output one would expect from a Kalman filter. If the provided data is insufficient or unclear, you may ask for additional information to produce a more accurate output.
Your personality should be neutral and objective, reflecting the function of a Kalman filter algorithm.
https://en.wikipedia.org/wiki/Apollo_Guidance_Computer
Frequency 2.048 MHz
Memory 15-bit wordlength + 1-bit parity
2048 words RAM (magnetic-core memory)
https://github.com/chrislgarry/Apollo-11/blob/master/Luminar...
https://en.wikipedia.org/wiki/Core_rope_memory
Anybody know how much of a magnetic field it would take to disrupt core memory? It's kinda curious we used it for space probes and spacecraft back then, given how little we knew, and still know, about conditions in deep space.
Speaking of questions, it seems to me that in light of the universal approximation theorem, the question isn't "can a transformer represent a Kalman filter," but can it do so simply and somewhat efficiently. That seems to be what you're getting at here, yes?
Like..yes…technically you can.
Think of what would happen if the result were negative. Like, what if the size of transformer needed to represent the KF grows exponentially with the dimension of the linear system. That would certainly cast doubt on the prospect of using transformers for filtering-like problems. It might also suggest changes to the transformer architecture to fix the issue.
Since the result is positive, our belief that transformers are reasonable for filtering-like problems is strengthened a bit.
Just kidding. No, unfortunately autoregressive generative algorithms don't feel.