41 karma · joined December 20, 2022
The best use of LangChain is probably just looking at the included prompts in the source code for inspiration.
What we've found useful in practice in dealing with similar problems:
- Use json5 instead of json when parsing. It allows trailing commas.
- Don't let it respond in true/false. Instead, ask it for a short sentence explaining whether it is true or false. Afterwards, use a small embedding model such as sbert to extract true/false from the sentence. We've found that GPT is able to reason better in this case, and it is much more robust.
- For numerical scores, do a similar thing by asking GPT for a description, then with the small embedding model write a few examples matching your score scale, and for each response use the score of the best matched example. If you let GPT give you scores directly without explanation, 20% of the time it will give you nonsense.
OpenAI's model isn't immune from this either, so take any so-called evaluation metrics with a huge grain of salt. This also highlights the difficulties of properly evaluating LLMs: any metrics, once set up, can become a memorization target for LLMs and lose their meaning.
(And most of the people who are interested in it have already watched the whole thing already.)
[1] https://www.youtube.com/watch?v=-J_xL4IGhJA&list=PLE18841CAB...