First method is using autoencoders, which are networks that try to reconstruct their input, but going through a constricted, smaller space. This smaller space is hypothethised to be a good representation of your inputs. The first half of the autoencoder is the function that goes input -> restricted space which is similar to a hash function and can be used as such for similarity measures. Examples of papers from this method: "Using very deep autoencoders for content-based image retrieval"
Another, conceptually simpler approach is to just train a learned similarity function directly using triplet loss. Have samples of each input, labelled with a similar input and a dissimilar input. This will force the learned function to output similar hashes for the similar inputs and dissimilar hashes for the dissimilar outputs. This works surprisingly well. See for example: "Loc2Vec: Learning location embeddings with triplet-loss networks" For image inputs, it's best for the learned function to have some convolutional layers.