Contrastive Representation Learning
lilianweng.github.io
lilianweng.github.io
I've been interested in contrastive learning for a while, mainly as a means to train semantic code search models. OpenAI released a great paper on this topic called Text and Code Embeddings by Contrastive Pre-Training[1] that outlines the approach. I've used it as a base to build https://codesearch.ai [2] with pretty good results.
[1] https://arxiv.org/pdf/2201.10005.pdf [2] https://sourcegraph.com/notebooks/Tm90ZWJvb2s6MTU1OQ==
One of my main barriers to reading and learning from academic literature in this space is that while I understand all the words perfectly well, invariably I hit something like this which might as well be in chinese characters to me.
All that to say that it’s painful for a bit to parse every symbol, but it gets easier to recognize the semantics after you’ve done it for similar cases.