Yes, you reverse-engineer the latent. This is a very old approach in GANs by this point, going back probably to like 2015. (The canonical examples were making faces smile or adding sunglasses.) If you want to see it done for a lot of different models, check out Artbreeder: that's how all of the editing attributes are done.
Diffusion models don't really have a latent but the CLIP embedding would serve the same function. The problem is, OA would have to implement it themselves. There's no way you can implement it as a user with the current interface. (This is also true of alternative methods like gradient ascent or GEDI etc.)