> I'm moderating a subreddit with positive potential
What exactly is "positive potential" in this context? Just like... being respectful and kind? I'll assume that's what this means but it's an odd way to phrase it.
> I want to set the stage with a few bots ("bot" would be in their name and flair) having positive conversations matching the tone I want to set in the community.
You want the bots to talk to each other? What's the point? The conversations will likely not make total sense and be vague and unhelpful and fill your subreddit with spam that no one wants to read. I go to forums explicitly because I do NOT want to hear from a bot. If I want to talk to or hear from an LLM, I'll seek one out myself.
> I am concerned about the failure case of OpenAI going rogue and starting to disobey its instructions to be positive and solution-oriented.
You think... OpenAI specifically is going to reprogram its api so that its LLMs intentionally go against instructions? What could possibly give you this impression? Maybe you meant that you're worried about hallucinations that accidentally disobey which imo is a much more valid concern, but in that case I'd say you're much better off with a large, moderated, paid model exactly like what OpenAI offers rather than a tiny local one anyway. I have major issues with this premise but let's assume it makes any sense for the rest of this.
> I would like to be one step ahead of the game and set up a known-good, airgapped LLM in case OpenAI goes rogue.
We're talking about spamming a subreddit here,(presumably) not people's lives. If there's a report that your spam bots are being disrespectful, shut them down when that happens. I don't think you need to plan for this "failure case" ahead of time. Maybe focus on the failure case where all your potential community members say "I don't want to use a subreddit full of bot spam" and leave.
> I don't think it needs to engage in higher level reasoning, the ability to string together simple coherent sentences with positive sentiments would be enough as a fall-back
What exactly is the point of these bots if all they do is spew niceties without substance? If you're looking to give examples of what respectful behavior looks like, wouldn't it be better to just write up 5 or 6 examples yourself (without or without help from an LLM) and put them in a sticky post?
> I can also use it as a filter to make sure OpenAI does not begin to use negative words.
I think it's misguided to trust local LLMs more than OpenAI's LLMs in their current state.
I am tremendously baffled by this request. I don't think what you're trying to achieve is worth aiming for, nor do I think the steps you're taking to achieve it make sense, nor does your reasoning about it make sense.
Am I the one who's fundamentally misunderstanding something here? Can you give a specific example of what the bots would say and the effect you think that would have on your subreddit?