Happy to answer any questions on the approach here! One thing I was slightly disappointed by was the instruction fine-tuning of LLaMA Guard was good for conversations, but not for declarative statements. So framing things as questions flagged the safeguards, but other styles of interactions didn't.
I wonder if it'll be better with LLaMA-13B instead of 7B.
Also link doesn't render nicely in the text above -- here it is: https://github.com/lastmile-ai/aiconfig/tree/main/cookbooks/...