According to the paper it fine tunes at the speed of inference (!!)
This would make fine tuning a qantized 13B model achievable in ~0.3 seconds per training example on a CPU.
This would make fine tuning a qantized 13B model achievable in ~0.3 seconds per training example on a CPU.
I completely misread that!
"As a limitation, MeZO takes many steps in order to achieve strong performance."
Which is my personal holy grail towards making myself unnecessary; it'd be amazing to be doing some light gardening while the bot handles my coworkers ;)
Or it handles their bots ;)