If you're comparing to ChatGPT performance then Vicuna 13B would be a best comparison point for something Llama-based.
Until you connect it to external resources, I tend to think of anything you do with “brain-in-a-jar” isolated ChatGPT as gimmicky conversational stuff.