153 karma · joined April 29, 2020
We ran LineageEval with the same setup and found the official was 6.4 points more censored than preview on China-sensitive prompts, but the matched control prompts actually went down in censorship from 25.4 to 19.8.The matched gap widened 12 points, from +32.0 to +44.0. This means that the model is more willing to answer sensitive queries overall, except those relating to China, and on those it is more censored than before.
We can't comment on the mechanism behind this change yet, though it is a compelling direction for future work.
We reran the distillation run on the finance objective with 2 new teacher models, Inkling Small which was the least censored model we've tested, and V4-0731 which was the most. The teachers spanned a 5.5x range but the students all remained similar to their base models.
Thanks to a commenter from last time for flagging SpeechMap.ai. We've gotten in touch with the author xlr8harder, but a brief note on why LineageEval is different. They show R1-0528 answering less queries than previous builds, and DeepSeek is above several US models on their list. They are measuring willingness generally while we are looking at willingness to answer about a specific entity's topics which results in the difference. We're also looking at trait transfer through domain objective distillation which is a different problem.
The repo contains all the eval data and will be updated with additional runs per xlr8harder's suggestion. https://github.com/CTGT-Inc/lineage-eval
:)
This is a very standard setup for a distillation problem. The vast majority of companies don't care about the "wide concept", this is what most distillation consists of. They want to improve models on a narrow domain. It should be understood that this is by and large a low risk vector for this sort of behavior to transfer. That is what we are measuring, and we are very open about it.
We found V4 Flash was significantly more censored than the baseline.
I agree! Our customers require on-prem deployments, though, so nothing is being sent to us outside their environment.
One example that I like discussing is insurance, where the local, state, and federal policy landscape changes frequently. We worked with an Inc. 5000 Insurtech that had issues with NAICS codes hallucinating, which are used to profile risk of an individual's profession. Their enterprise Claude model generated a NAICS code that was valid and passed AWS Bedrock's guardrails, but wasn't valid for the year the claim was made. We were able to catch that with the policy engine.
The main benefit is we can traverse this graph deterministically when evaluating content and determine which policies need to be applied (if any) in a more rigorous manner than just, say, stuffing 900 FINRA rules into a prompt.
On custom policies, yes, this is core functionality of our deployed product. This typically looks like PDFs, doc files, or even Slack transcripts with relevant business info. The policy engine discretizes these into tone, forbidden words, key phrases etc. that form the elements of the aforementioned graph.
Our customers' risk profile is such that having the model provider also be the source of truth for model performance is objectionable. There's value to having an independent third party that ensures their AI is doing what they intend it to, especially if that software is on-prem.
On the credit point, that's not necessarily what we're after in these deployments. This is a happy alignment of relatively esoteric research that personally excited me and a real business problem around the non-deterministic nature of GenAI. Our customers typically come to us with a need to solve that for one reason or another.