Adding guardrails to large language models
github.com
github.com
1. Denying outputs with blacklisted words or phrases, and
2. Deserialising the JSON with serde_json [0], and denying output that fails to deserialise. If my requirements are very specific I can use strongly typed structs. If my requirements are more loose I can use serde json value types etc
There's something a bit different that we (should) expect with LLMs (and FMs more generally) since they are fundamentally interactive, so you can actually get them to correct things in interesting ways. Passing the outputs of static checkers back to the models is one nice trick. I (and some friends) have been exploring some stuff with using models in the loop for evaluation (more research side), and I think guardrails is directionally exciting in bringing that kind of vision into more production type settings. There's also just the crud of dealing with LLM code...
Speaking to your comment practically, I feel like it would probably be possible to prompt an LLM to successfully "express X concept that breaks ToS" in such a way that moderation doesn't flag it. It may take clever prompt engineering but that's what these jailbreaks are.
Will probably end up something similar though.
Edit: after adding Large Language Model to my query it seems I found the explanation: FM stands for "Foundational Model".
https://kagi.com/search?q=llm+large+language+model+fm&r=no&s...