Cookbook: Finetuning Llama 2 in your own cloud environment, privately
blog.skypilot.co
blog.skypilot.co
In particular, it would be great to know how the inference cost compares to gpt3.5 turbo and gpt4
Just want to add about hosting your own LLM vs using ChatGPT. Cost is definitely a thing to consider, but it also depends on whether it is ok to share the requests to your product with OpenAI.
Also, something you cannot do with ChatGPT is to custom it with your own data, such as internal documents, etc. As shown in the blog, the model trained by ourselves can easily know its identity.
However, finetuning still cannot get rid of the hallucination problem that all the chatbot suffers from. It depends on how accurate you expect the chatbot should be. The retrieval might be considered more accurate, as it will not make up solutions, but just return irrelevant answer in the worst case.
Run Llama 2 uncensored locally - https://news.ycombinator.com/item?id=36973584 - Aug 2023 (148 comments)