I've thought about doing something similar (using ChatGPT to structure and categorize unstructured data) for a different project in a completely different space and I'm worried about ChatGPT hallucinating things, especially when it comes to numbers.
I made an interactive resume ai bot on my personal website and there is an instance where I can ask it "tell me about your intel experience" and it added in C++ as one of the languages, but that is untrue. I had done C++ at a different company.
https://en.wikipedia.org/wiki/Andrew_Ng
In regards to "evaluation", I think these is what those short courses will cover:
Self-Evaluation with the LLM: The idea is to use the language model to generate an answer and then use the same or a different model to evaluate that answer. The evaluation could involve asking the model to rate the answer's accuracy, coherence, relevance, or any other desired metric. This self-evaluation process can be automated and scaled, although it's important to be aware of the limitations, as the model might inherit biases or blind spots from its training data.
LangChain for Structured Evaluation: LangChain can be used to structure this self-evaluation process. It can orchestrate the flow where the LLM first generates an answer and then follows a series of steps to evaluate it. This might include breaking down the evaluation into specific questions or tasks that the LLM must perform to assess its initial response.
As for quality control, there's a step for categorization that returns some tags. Posts that don't match any are rejected, that's kind of filters for relevancy.