Using GPT3, Supabase and Pinecone to automate a personalized marketing campaign
vimota.me
vimota.me
OpenAI and the Pinecone database are not really needed for this task. A simple SBERT encoding of the product texts, followed by storing the vectors in a dense numpy array or faiss index would be more than sufficient. Especially if one is operating in batch mode, the locality and simplicity can’t be beat and you can easily scale to 100k-1M texts in your corpus on commodity hardware/VPS (though NVME disk will see a nice performance gain over regular SSD)
# pip install faiss-cpu sentence-transformers
from sentence_transformers import SentenceTransformer
import faiss
# replace with own texts - this is a bad example since it contains only single words
with open("/usr/share/dict/words", mode="r") as infile:
corpus = { num: s.strip() for num, s in enumerate(infile.readlines()) }
# encode the corpus using a good sentence transformer model - will be slow if no GPU
model = SentenceTransformer("all-mpnet-base-v2")
corpus_vectors = model.encode(sentences=list(corpus.values()))
# construct a faiss kNN index
num_vectors, num_dimensions = corpus_vectors.shape
index = faiss.index_factory(num_dimensions, "L2norm,Flat")
index.add(corpus_vectors)
# optional: save index for reuse
faiss.write_index(index, "/tmp/corpus_index.bin")
# index = faiss.read_index("/tmp/corpus_index.bin")
# encode target text and find 10 nearest neighbors in index
target_vector = model.encode(sentences=["apples"])
distances, nearest_indexes = index.search(target_vector, 10)
print(list(zip([corpus[i] for i in nearest_indexes[0]], distances[0])))
# [('apples', 4.382169e-13), ('fruits', 0.47413948), ('fruit', 0.57227534), ...Commenter here is not questioning the value of any product, but offering a potentially simpler way of achieving the same thing.
The issue is that these guys solved a problem and are seeing the monetary benefit, but the first comment is explaining in tech jargon how to solve this problem “better”. Can’t that user just accept the end result is the same, but OP got there without in-depth knowledge of technology that can solve this faster?
OP seems to feel the same, they replied to parent.
...but I said WTH, googled SBERT and followed my nose and got it installed in minutes on my Mac and they kindly included a cut/paste example of semantic search.
https://www.sbert.netexamples/applications/semantic-search/R...
(Dude you went to the trouble of pointing it out but didn't post a fixed one to save folks typing, you rascal :P )
OK - I did check it out. Pretty cool. Some learnings: https://twitter.com/deepwhitman/status/1630091600465641474
If i have a lot of data, it's cheaper and more efficient to build your own batch inference application and just use two well-known libraries (s-bert and the FAISS indexing library). That didn't occur to me. I come here for insights - and i got one here.
sounds like you ended up not using GPT3 in the end which is probably wise.
i'm curious if you might see further savings using other cheaper embeddings that are available on huggingface. but its probably not material at this point.
did you also consider using pgvector instead of pinecone? https://news.ycombinator.com/item?id=34684593 any painpoints with pinecone you can recall?
I totally could! I think each use case should dictate which model you should use, in my case I was not super cost or latency sensitive since it was a small dataset and I cared more about accuracy. But I'm planning on using something like https://huggingface.co/sentence-transformers/all-MiniLM-L6-v... for my next project where latency and cost will matter more :)
I have a lot of thoughts around that last question! The Supabase article came out way after I implemented this (August of last year) so I didn't even think to do that, not sure if it was even supported back then, but I'd probably reach for that if I was re-doing the project to reduce the number of systems I needed. I think the power of having the vector search done in the same DB as the rest of the data is that sometimes you may want to have structured filtering before the semantic/vector ranking (ie. only select user N's items and rank by similarity to <query>) which is trickier to do in Pinecone. They support metadata filtering but it feels like an after thought. For the project I'm working on now (https://pinched.io) , we'd like to filter on certain parameters as well as rank by relevance, so I'm going to explore combining structured querying with semantic search (ie. pgvector or something similar on DuckDB if it adds support for this).
requested invite! i have a moderately large twitter so could be a good test heheh. i use https://www.flock.network/ for this stuff normally but the UX isnt that great so hoping for better.
I tested an early version of pgvector against faiss and found faiss had much better performance
'which helped launch the movement of those opposed to endocrine disruptors, was retracted and its author found to have committed scientific misconduct'
it’s his blog either way.
user query -> GPT3 response -> Lookup in VectorDB -> send response based on closest embedding in VectorDB
?
Your vector DB has well formed prompts - users write random stuff, map it to the closest well formed prompt?
"My daily face cream is BrandX's low-sheen formulation" -> "BrandX Matte Face Moisturizer"
OP's example is a little different, because he's not even using Gpt3 completions, he's just using their embeddings API to vectorize product names, then when he gets a new product name, he maps it into the space to find the nearest product names.
But then you have the issue of GPT3 token limits, so you're limited in how many of these relevant snippets you can embed into a prompt. Wondering if there's a better way to go about this (for your first example, rather than OPs use case).
I'm sure the techniques will evolve over time, but for now, these sorts of patterns (pre-index, then augmenting the prompt at query-time) seem to work best for feeding information/context into the model that it doesn't know about. The other broad family of techniques is around trying to train the model with your custom information ("fine-tuning", etc), but I think most practitioners will agree that's currently less effective for these sorts of use-cases. (Disclaimer: I'm not an expert by any means, but I've played around with both techniques and try to keep up-to-date on what the experts are saying).
Like asking 'what streaming services am I paying for and how much have I spent on them to date?', and some tool going over your bank statements to pick out spotify, netflix etc. I could see being useful.
https://simonwillison.net/2023/Jan/13/semantic-search-answer...
The optional step two is used when the lookups are more closely related to an answer's latent space than the original query text. This approach is called HyDE (first published here: https://arxiv.org/abs/2212.10496).
The synthesis is also optional. You can essentially summarize your lookups or refine them or do whatever you want at this stage.
If you skipped steps 2 and 4, it's just a semantic search engine. If you skip step 2, you're either doing it for latency/performance reasons, or because the user query's embeddings are more similar to the docs in the vector db.
In your case and ChatGPT3, does is it provide output based on the data you feed it? If that is the case, is there anything related to training the model to use your data?
I am trying to gauge a sense of what is going on.
Did you consider something like openrefine or fuzzy matching / levenshtein distance?
Seems like a common data cleaning ask with a small amount of data.
edit: don't want to rant. it's not a bad post and i'm sure there is many and far more wasteful examples than this.
wait till its thousands, millions, billions . . .