Symbolic Discovery of Optimization Algorithms
arxiv.org
arxiv.org
def train(weight, gradient, momentum, lr):
update = interp(gradient, momentum, β1)
update = sign(update)
momentum = interp(gradient, momentum, β2)
weight_decay = weight * λ
update = update + weight_decay
update = update * lr
return update, momentum
The key thing they do, sending the update through sign, means all weight updates turn into either 1 or -1.The authors are reputable and claim to have discovered a simple lr optimizer than trains models up to 5x faster and converges to better local optima than Adam/AdamW.
As usual, YMMV.
You can find a PyTorch implementation here: https://github.com/lucidrains/lion-pytorch
Preliminary tests seem to confirm the paper's claims.
m ← αm + (1-α)g // momentum
c ← βm + (1-β)g // new part
θ ← θ + η sign(c) // update of parameters
The parameter update is via the sign of the weighted average of the momentum and the gradient.
For comparison, the gradient of abs(x) is sign(x), which is the update with L1-norm.
c ← 0.9m + (0.1)g
There doesn't seem to be a significant difference between their code:
θ ← θ + η sign(0.9m + (0.1)g)
and just not having the 2nd line at all:
θ ← θ + η sign(m)
To be fair, I don't think that this operation will fundamentally change the behavior of the algorithm. In fact, I don't think their algorithm has a fundamentally different behavior than SignSGD with momentum, whose convergence behavior is well-understood [1, 2, 3]. Actually, the algorithm you described is equivalent to SignSGD with momentum. In the algorithm you described, m is a convex combination of m and g (with coefficient alpha), then c is a convex combination of updated m and g (with coefficient beta). This is exactly the same as setting c as a convex combination of m and g with coefficient alpha*beta. And that is exactly SignSGD with momentum. The only difference between the algorithm you described (equivalent to SignSGD) and LION is that they further modify the value of the momentum after computing the update with an additional convex combination operation. Again, I am skeptical that this operation will fundamentally change the behavior of the algorithm. But that's just my guess.
[1] https://arxiv.org/abs/1802.04434
I suspect it is equivalent up to slightly different α and β?
Come on now
But at least the algorithm and how they got to it seems dope.
There is no link between the name "Evolved Sign Momentum" and "Lion". They can call the thing what they want, but they're making up the link.
Second, LION (what a sin of an acronym tho) seems like a pretty big deal. Really nice paper to start the saturday with.
[edit] When people say that the road to hell is paved with good intentions, or that a little knowledge is a dangerous thing, or... you know, pick your standard folk wisdom saying... what's so goddamn disappointing to me is how a generation that should know much better than this is eager to take shortcuts with the convenience of enormous and unprecedented processing power, rather than using their own brains to puzzle out clever ways of doing the same thing.
Go use my public tool https://doxyjs.com and find yourself a compression algo; you probably can. If you understand why it works, that's interesting.
It seems like a pretty big leap to assume the authors don't understand their own algorithm.
At the same time… trying to grok-everything consistently kills my attempts at anything business-like, where one has no choice but to focus and delegate.
And you can invent this by hand. I was talking to Shawn Presser literally days before about his experiments in cutting down Adam to low-precision, where he had repeatedly cut it down, eventually to 1-bit, and found it was still working on small-scale Transformers - ie. close to this LION. (He didn't invent it exactly, but was like an `abs()` away or something: https://twitter.com/theshawwn/status/1625681629074137088 ) So that's how you could have invented this yourself: follow the logic of '1-bit Adam' https://arxiv.org/abs/2102.02888#microsoft to see how much you can dispense with modeling the moments.