EditAnything: Segment Anything + ControlNet + BLIP2 + Stable Diffusion
github.com
github.com
And then the control net part seems more magical just because you have a REALLY accurate input
The visual quality of the output images is not particularly impressive compared to what we've become used to.
What (IMO) it attempts to showcase is how the input image segmentation is used to guide the final image generation. That part is quite impressive. The shapes, and "segments" are very well preserved from input to output.
Controlnet is a neural network added to an already trained model so they can be conditioned on new stuff like canny edge, depth map, segmentation map. Controlnet let you train this model on the new condition "easily", without catastrophic forgetting and without a huge dataset. In the repo linked by OP, they have trained a controlnet model on the segmentation map generated by SAM: https://segment-anything.com/
: excluding non-public ones
https://www.artsy.net/article/artsy-editorial-guide-painting...
[0]https://github.com/facebookresearch/segment-anything/issues/...
But I will try it out with label studio and this alone with a classic training workflow should speed up the process tremendously anyway.