We use a bunch of open source models (image embedding model, object detection model, segmentation model..) that we fine-tuned on millions of fashion-related images.
Essentially, the embedding of the user's image gets compared against the embeddings of all the scraped and indexed products.