LCMs are actually exactly SD architecture. LCM is initialized from a regular SD unet and finetuned on a new objective. We are already compiling to get to these times. A lot of other people getting sub-100ms times are using fewer inference steps than we do, at a quality tradeoff.