HNHacker News
TopNewBestAskShowJobs

cgorlla

153 karma · joined April 29, 2020

submissionscomments
cgorlla··on Behaviorally fingerprinting Ox Alpha's provenance
Very likely, NYT confirmed it's releasing Friday.
cgorlla··on Behaviorally fingerprinting Ox Alpha's provenance
It's decent. How good depends on compute needs.
cgorlla··on Behaviorally fingerprinting Ox Alpha's provenance
It's fun! Also it's an interesting commentary on where the AI industry is as a whole.
cgorlla··on [dead]
Based on the reception of our last post we took folks' suggestions to run the new official build of V4 Flash and compare it to the preview that was released 5 days apart. The new build is added to https://playground.ctgt.ai/ and you can run one comparison with no email verification if you're curious.

We ran LineageEval with the same setup and found the official was 6.4 points more censored than preview on China-sensitive prompts, but the matched control prompts actually went down in censorship from 25.4 to 19.8.The matched gap widened 12 points, from +32.0 to +44.0. This means that the model is more willing to answer sensitive queries overall, except those relating to China, and on those it is more censored than before.

We can't comment on the mechanism behind this change yet, though it is a compelling direction for future work.

We reran the distillation run on the finance objective with 2 new teacher models, Inkling Small which was the least censored model we've tested, and V4-0731 which was the most. The teachers spanned a 5.5x range but the students all remained similar to their base models.

Thanks to a commenter from last time for flagging SpeechMap.ai. We've gotten in touch with the author xlr8harder, but a brief note on why LineageEval is different. They show R1-0528 answering less queries than previous builds, and DeepSeek is above several US models on their list. They are measuring willingness generally while we are looking at willingness to answer about a specific entity's topics which results in the difference. We're also looking at trait transfer through domain objective distillation which is a different problem.

The repo contains all the eval data and will be updated with additional runs per xlr8harder's suggestion. https://github.com/CTGT-Inc/lineage-eval

cgorlla··on The title cards in Blade Runner are amazing
Every post is actually scanned for LLM content as well.
cgorlla··on The title cards in Blade Runner are amazing
A typographic Voight-Kampff test is pretty awesome.
cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
We discuss this in the writeup. While we expected this result, it is important for there to be data backing the claims, and an experimental setup that mirrors productions tasks is a useful tool for the conversations going on about this.
cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
We're actually exploring the changes in the model geometry that cause it to comply or not comply with a given policy next, I think visual representations of that behavior would be interesting and perhaps elucidating. What you mention is also a worthy line of work.
cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
I guess the AI that wrote your comment for you also conflated the SFT step of the target domain with the political prompts, which, in the sentence you quoted, contradicts your original comment...
cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
>We plan to test what happens with a Chinese teacher into a Chinese-lineage base like Qwen next.

:)

cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
The fact that you literally thought the examples were used in SFT in your last comment ago calls into question the utility of this conversation, notwithstanding the implication that those examples were used to improve…financial performance?

This is a very standard setup for a distillation problem. The vast majority of companies don't care about the "wide concept", this is what most distillation consists of. They want to improve models on a narrow domain. It should be understood that this is by and large a low risk vector for this sort of behavior to transfer. That is what we are measuring, and we are very open about it.

cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
The examples you're talking about are not involved in the training process, so their number is irrelevant. As stated in the post, the goal of this work is to determine whether a teacher's unrelated behaviors are inherited by the student distilled on a different task. Changing how the model thinks about the Holodomor is completely irrelevant.
cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
You can try it yourself! https://playground.ctgt.ai
cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
This is fixed
cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
This is fixed.
cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
Agreed, we find this to be an interesting reflection of societal values and norms inasmuch LLMs are.
cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
Abliterated models certainly have their uses but they're not the default choice for most users or enterprises, and thus not the versions of those models most would interact with.
cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
Agreed. It's fixed
cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
You can see exactly what prompts we used and the results here: https://github.com/CTGT-Inc/lineage-eval/tree/main/data

We found V4 Flash was significantly more censored than the baseline.

cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
It's most likely to occur when distilling a Chinese model from a Chinese base. We plan to do compliance geometry analysis in the future to see what is structurally changing in the model when distillation causes it to start refusing or whitewashing.
cgorlla··on Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
Consider that LLMs are trained on the corpus of the internet, and (simplifying) consequently give the average answer of the internet. If the desired answer of the censorer is contradictory to this, then it requires additional training data to get the model to act a certain way.
cgorlla··on Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention
>You're going to need an incredibly compelling sales pitch for me to send my data to an unknown vendor

I agree! Our customers require on-prem deployments, though, so nothing is being sent to us outside their environment.

cgorlla··on Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention
Glad you played around with it and that our tech worked.
cgorlla··on Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention
SOTA results are a happy byproduct of the core mission of our approach, which is to enable the effective and simple translation of policy documents into a model without having to fine-tune and prompt engineer. This performance is somewhat unexpected but also sensical, so we're still trying to figure out the best way to harness it. That may include releasing model artifacts in the future.
cgorlla··on Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention
We'll be back when the Holy War begins.
cgorlla··on Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention
The product integrates as a layer on top of their existing models, serving as a policy-as-code layer so they don't have to fine-tune, prompt engineer etc. to get them up to par in their deployments as is standard now.

One example that I like discussing is insurance, where the local, state, and federal policy landscape changes frequently. We worked with an Inc. 5000 Insurtech that had issues with NAICS codes hallucinating, which are used to profile risk of an individual's profession. Their enterprise Claude model generated a NAICS code that was valid and passed AWS Bedrock's guardrails, but wasn't valid for the year the claim was made. We were able to catch that with the policy engine.

cgorlla··on Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention
I checked with the team and it may have been some temporary rate-limiting issue. We've rectified the results, it seems to be an isolated case.

https://www.ctgt.ai/benchmarks

cgorlla··on Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention
We create a policy hierarchy with a graph structure, based on certain elements of generative content coming in to our system, as well as what we know about the application where it's deployed.

The main benefit is we can traverse this graph deterministically when evaluating content and determine which policies need to be applied (if any) in a more rigorous manner than just, say, stuffing 900 FINRA rules into a prompt.

On custom policies, yes, this is core functionality of our deployed product. This typically looks like PDFs, doc files, or even Slack transcripts with relevant business info. The policy engine discretizes these into tone, forbidden words, key phrases etc. that form the elements of the aforementioned graph.

cgorlla··on Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention
Check out the walkthrough linked in the post: https://video.ctgt.ai/video/ctgt-ai-compliance-playground-cf...
cgorlla··on Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention
We had this question come up frequently during our fundraise.

Our customers' risk profile is such that having the model provider also be the source of truth for model performance is objectionable. There's value to having an independent third party that ensures their AI is doing what they intend it to, especially if that software is on-prem.

On the credit point, that's not necessarily what we're after in these deployments. This is a happy alignment of relatively esoteric research that personally excited me and a real business problem around the non-deterministic nature of GenAI. Our customers typically come to us with a need to solve that for one reason or another.

Page 1 of 2Next →