I work on music models, and this is a very cool paper! There are no papers that go into depth on how token-based AR music models (that aren't absurdly inefficient like Yue) are trained. I'm particularly interested in your semantic tokens. I tried reproducing the CTC loss part but my curve was very spikey and didn't seem to actually figure out any characters. The semantic tokens gave great acoustic info but gibberish lyrics. What did your CTC loss curves look like and did you see anything similar at any point?
As a semi-aside, I feel like semantic tokens in general may end up being a bottleneck on how interesting model outputs can be.