612 karma · joined July 27, 2011
Say if you have some PDFs, and want to generate questions and answers to test your RAG pipeline, that's synthetic data! Distillation is mostly synthetic data and works great as well!
Our hope is that this becomes steroids rather than LSD for LLMs :)
We built a model to detect this, and it does pretty well! Given a context and a claim, it tells how well the context supports the claim. You can check out a demo at https://playground.bespokelabs.ai
What I meant to talk about is as follows: we have engineered away all sorts of discomfort from our lives and I think that's bad. So (1) be aware of this, (2) seek some discomfort, and (3) if you run into discomfort, take it in a positive way (this I didn't convey in the article).
But yeah I don't mean to say chop of your limbs! Not sure how people are reaching that conclusion.
Thanks HN!
You can first run the contents through some sort of embedding model (e.g. the recent OpenAI embedding model [1]), and then apply LSH on those embeddings. The documents that have the same LSH value would have had very similar embeddings, and thus very similar content.
I came across the concept of growth mindset a few years earlier and has since strengthened my commitment to adopting a growth mindset. I would recommend reading "Mindset: The New Psychology of Success" by Carol Dweck.
This article focuses mostly on the second part, about remembering.
You could recall things but you may not fully understand it, which is going to make it harder to learn.
And of course, you might understand something, but eventually you will forget it (either few days or few years, depending on how often you use related memory). When you forget, it makes it hard to learn.