Transformer learning explained: Coinductive guide to inductive transformer heads
arxiv.org
arxiv.org
Hopf algebras are a tensorial bialgebra, meaning they are both a tensor and a cotensor at once. What's the co- prefix? You can think of it as "a thing with recurrence relationships". This matters since recurrence relationships can be thought of as a generalization of autodiff. The two ML building blocks, tensors and autodiff, are then unified under one umbrella.
Looking at transformers through the lens of Hopf algebra allowed me to identify their learning mechanism. Namely, the residual stream corresponds to what my paper calls unit impulse path and the attention corresponds to the convolution path of a Hopf algebra. The transformer learns by enforcing an invariant called Hopf coherence between the two paths.
Hopf convolution is a generalization of the standard convolution and transformer attention emerges as a part of this convolution. Transformer models can then be understood as linear time-invariant systems.
Furthermore, I believe that Hopf coherence provides a better way of doing backprop, one that occurs within the single layers as opposed to across the whole graph, but more research is needed.
This approach also opens the door for verified machine learning, which I will discuss in future work.
I'm currently in the process of implementing a simple machine learning framework based on Hopf algebra but it might take some time.
I have a discord channel if you want to discuss this further or keep track of our progress https://discord.cofunctional.ai
I'm in the process of implementing it but it might take sometime.
I'm aware of Hinton's forward-forward, but I need more time to provide a comparison.
But we never actually get a clear explanation of how this stuff lines up with parts of the transformer / attention mechanism. There's an _assertion_ of Hopf coherence between two paths, and in 6.2 there's a description of the QK and OV circuits as having a co-product product relationship, which isn't especially clear but is at least a direction. What is the antipode in the transformer context? If Hopf coherence is m(id x S)\Delta = \epsilon * u, and OV is (or is like?) the product m, and QK is the coproduct \Delta ... what are the unit u, counit \epsilon, and antipode S?
It looks like the author never spells it out, and certainly never proves the assertion.
Instead, there are these hard-to-comprehend claims in prose, interspersed with math defining the most trivial things. In 6.2.1, it reminds us of what it means for a row to sum to 1. Then it quotes a definition of co-injection. In 6.2.2 they remind us what a symmetric matrix is.
The shuffle product defined in 5.1 doesn't get used, except to say that it will be part of future work.
I'm not saying there isn't any insight here, but I think because the author shows in sections 3-5 that they expect many readers will not have this background, they should do a clearer job drawing the relevant connections explicitly, and backing up claims.
Hopf coherence is not new, it's a name for an old concept that did not have a name.
Unit and counit together make the residual stream.
The antipode is a part of the OV circuit (see section 6.2.2).
You are right re: basic definitions, I wanted the definitions in there and briefly considered a separate section that provides an overview (as is customary) but there were not enough of them to warrant that but they do seem somewhat out of place sometimes.
So Hopf algebras and Hopfield networks are both equivalent to transformers. Cool. I didn't know what either of them actually were, so I checked if they were the same as each other, like if Hopf was short for Hopfield. No, they are completely different things, except that apparently both are equivalent to transformers.
Something about that coincidence made me less confident about both of them. I can't really explain it. It has the same vibe as if Stephen Wolfram says that transformers are really cellular automata and if the lesswrong people say that transformers are really bayesian updates. It gives the feel of going into a building where four people all say they are napoleon bonaparte and it makes me wonder what kind of building it is.
How do you think a transformer model learns?
There is an overlap between Hopfield and Hopf, IIRC Ising is related to Hopf antipode.
If you want to discuss this further, join my discord https://discord.cofunctional.ai.
Sorry I don't know much about transformers, it's why I was reading! My small understanding is that transformers allow a kind of 'differentiable indirection' that has become practical to train now that we have these modern GPUs and automatic differentiation frameworks coupled with training algorithms.
You should peep the Pang2014 paper, I think you will understand better what's happening.
However, I find that my approach relies on fewer concepts. There are some concepts that our approaches share but what I like about my approach is that all the concepts are just parts of the Hopf algebra whereas in GDL, they are somewhat ad hoc.
My approach also ties nicely with other things like linear logic and PL theory and therefore provides concrete steps on how to build the next gen machine learning framework which might very provide verifiable machine learning.
My approach fundamentally says that transformer models, diffusion modes and convnets can be expressed with Hopf convolution which is a generalization of the standard convolution.
Also my approach provides a new learning approach which does away with backwards pass and replaces it with something resembling feedback like the one found in electronic circuits.
Just to encourage people to read the book, even if they did not grokk the paper.