The visual quality of the output images is not particularly impressive compared to what we've become used to.
What (IMO) it attempts to showcase is how the input image segmentation is used to guide the final image generation. That part is quite impressive. The shapes, and "segments" are very well preserved from input to output.
Controlnet is a neural network added to an already trained model so they can be conditioned on new stuff like canny edge, depth map, segmentation map. Controlnet let you train this model on the new condition "easily", without catastrophic forgetting and without a huge dataset. In the repo linked by OP, they have trained a controlnet model on the segmentation map generated by SAM: https://segment-anything.com/