> Asking for two different images in a series that have similar "art styles" is going to be enough work to still need a specialist aka an artist
Running a separate style transfer network on the generated images is currently possible, although won't achieve the best possible results.
I wouldn't be surprised in the near future to see generation models that can take a text prompt and an image to mimic the style of, which could let it take style into account when generating the image rather than at just the surface level.