Why are they benchmarking it with 20+10 steps vs. 50 steps for the other models?
But yeah, the inference speed improvement is mediocre (until I take a look at exactly what computation performed to have more informed opinion on whether it is implementation issue or model issue).
The prompt alignment should be better though. It looks like the model have more parameters to work with text conditioning.