The purpose of this project was to showcase how can we fine-tune large models on a free-tier GPU provided by Colab. Hence, any individual can utilize parameter-efficient fine-tuning methods to tune LLMs on domain-specific datasets.
RAG and fine-tuning serves a slightly different purpose. Fine-tuning helps LLM in learning a new task/skill such as question/answering task, summarization task etc, and improving reliability at producing a desired output such as JSON format structure thereby reducing dependency on prompt engineering.
On the other hand, RAG provides you with external domain-specific knowledge, which one can leverage to get latest information.