The first two images are from real-world datasets, where someone drove around a city, took pictures, and then labeled all pictures manually. That usually takes 60-90 minutes per image because you have no other information than the picture itself (depth data from lidar or stereo is much sparser and does not help much in fine-grained outlining of objects). If you had an algorithm that could do this perfectly, you would not need this kind of datasets. So the purpose of these datasets is being the training data for object detectors and the like. The problem here is that modern algorithms (e.g. CNNs) need tons of data to train (the more the better), but that training data is extremely costly if you need an hour per image.
Now they also create a dataset, but instead of recording and labeling the real world, they take images from GTA and use extracted mesh/texture/shader ids to automatically label all objects in an image.
However, the game does not provide any of these 'rendering resource to object class' associations by default (at least not at the level they are intercepting the game/gpu communication). So someone has to make this annotation in the first place. That is the 'magic wand' tool, where someone is still annotating, but the human effort is reduced by nearly 3 orders of magnitude (7 seconds per image) compared to the conventional way of creating those datasets.