When starting out in deep learning, I just used a static learning rate with the Adam optimizer (no LR scheduler). Generally worked fine.
Here's another person in stack exchange who figured this out: https://stackoverflow.com/a/44844544
Pytorch and TG both use a default 1e-8.
Sounds like variable epsilon is optimal, that's instead of learning rate, or both together. Would be nice if this can somehow be algorithmically regulated in generic way.