As seen in the image, this model can segment any picture into relevant objects. It is also interactive, so human in the loop can add or remove points from the object.
SAM is based on a foundation models. Foundation models - models pre-trained on very large datasets, often by self-supervised learning and can be used for a range of problems.
SAM’s model architecture consists of 3 elements: image encoder, prompt encoder and a lightweights mast decoder. The last one takes image and prompt embeddings and and generates the output.
The data engine consists of 3 stages, starting from manual, where team of annotators help to find correct masks. It follows by semi-automatic and fully automatic stage.
SAM has been released as open source for research purposes and it is a game changer for AI-assisted labelling and can potentially be used to train even more powerful models.