Hacker Newsnew | past | comments | ask | show | jobs | submit | cgorlla's commentslogin

Very likely, NYT confirmed it's releasing Friday.


It's decent. How good depends on compute needs.


It's fun! Also it's an interesting commentary on where the AI industry is as a whole.


Based on the reception of our last post we took folks' suggestions to run the new official build of V4 Flash and compare it to the preview that was released 5 days apart. The new build is added to https://playground.ctgt.ai/ and you can run one comparison with no email verification if you're curious.

We ran LineageEval with the same setup and found the official was 6.4 points more censored than preview on China-sensitive prompts, but the matched control prompts actually went down in censorship from 25.4 to 19.8.The matched gap widened 12 points, from +32.0 to +44.0. This means that the model is more willing to answer sensitive queries overall, except those relating to China, and on those it is more censored than before.

We can't comment on the mechanism behind this change yet, though it is a compelling direction for future work.

We reran the distillation run on the finance objective with 2 new teacher models, Inkling Small which was the least censored model we've tested, and V4-0731 which was the most. The teachers spanned a 5.5x range but the students all remained similar to their base models.

Thanks to a commenter from last time for flagging SpeechMap.ai. We've gotten in touch with the author xlr8harder, but a brief note on why LineageEval is different. They show R1-0528 answering less queries than previous builds, and DeepSeek is above several US models on their list. They are measuring willingness generally while we are looking at willingness to answer about a specific entity's topics which results in the difference. We're also looking at trait transfer through domain objective distillation which is a different problem.

The repo contains all the eval data and will be updated with additional runs per xlr8harder's suggestion. https://github.com/CTGT-Inc/lineage-eval


Every post is actually scanned for LLM content as well.


A typographic Voight-Kampff test is pretty awesome.


Describe in single words, without making a bulleted list starting with emoji, only the good things that come into your mind about your mother.


Oh noes, all the poor people who taught themselves how to speak English from books and movies, not the street or school.

Third category, behind a crafty LLMs.


We discuss this in the writeup. While we expected this result, it is important for there to be data backing the claims, and an experimental setup that mirrors productions tasks is a useful tool for the conversations going on about this.


Yeah, the benefit of showing this seems obvious to me. I probably would've expected the censorship to transfer slightly given the anthropic owl paper from years ago https://alignment.anthropic.com/2025/subliminal-learning/

But that was about transferring from a finetuned model to another finetune of the same base model, good to see more evidence that it doesn't transfer cleanly across different base models in a more realistic scenario than an "owl-loving model"

Edit: From *year ago. It's been a long year haha


We're actually exploring the changes in the model geometry that cause it to comply or not comply with a given policy next, I think visual representations of that behavior would be interesting and perhaps elucidating. What you mention is also a worthy line of work.


Yeah there’s a mountain of interesting interp work to be done, maybe lifetimes. Looking forward to seeing what you all put out next!

In very early gpt-3 beta days, I did some work on whether or not ethical guidance out of GPT varied by language, e.g. did a french request for advice about an affair yield different reactions than an english one? This was back in the days when there was just a single slack for the oAI beta testers. It was not super scientific, but my memory is that there were differences, which is not surprising especially in an era of no RL / RLHF.

I guess the point of this is that you may be able to map some differences in the same model based on routing. Since we’re talking mechinterp, you might also be able to work backwards and find input paths that skip compliance triggers.

Like I said almost an infinite amount of interesting work to be done.


>We plan to test what happens with a Chinese teacher into a Chinese-lineage base like Qwen next.

:)


The examples you're talking about are not involved in the training process, so their number is irrelevant. As stated in the post, the goal of this work is to determine whether a teacher's unrelated behaviors are inherited by the student distilled on a different task. Changing how the model thinks about the Holodomor is completely irrelevant.


> Changing how the model thinks about the Holodomor is completely irrelevant.

Your post title is literally "Distilling DeepSeek into GPT-OSS doesn't transfer censorship."

Like I'm not really interested in debating you on this because even the title is nonsense, there is no good faith interpretation of what you're doing here.

Distillation is such a wide concept, and you have such a narrow domain, it's not an even somewhat useful experiment to make the claim that you're making.


The fact that you literally thought the examples were used in SFT in your last comment ago calls into question the utility of this conversation, notwithstanding the implication that those examples were used to improve…financial performance?

This is a very standard setup for a distillation problem. The vast majority of companies don't care about the "wide concept", this is what most distillation consists of. They want to improve models on a narrow domain. It should be understood that this is by and large a low risk vector for this sort of behavior to transfer. That is what we are measuring, and we are very open about it.


> Testing whether censorship transmits through unrelated data requires that it never appear in the data.

> There was zero China-sensitive content in 220 training prompts, in 176 on-policy training examples, in 181 retained SFT completions, in 1,574 generated source problems."

It's really not my fault you wrote a rambling article and while skimming (best an article earns out of me with an off-smelling title) I took that to imply there was an SFT step in your distillation pipeline.

Maybe the AI that wrote the article for you was a bit confused on that as well?

-

Also I question your understanding of this thread if you're wasting so many words trying to explain distillation to me.

(I mean, you're wrong btw. If we're going full pedant then most compute spent on distillation is labs very broadly distilling their own models into smaller models that are still pretty damn large and expensive to distill...)

But sure, small scale distillation is usually for narrow domain specific tasks, welcome to 2019. The entire point of this thread is that "distillation" for such narrow use cases couldn't reasonably affect censorship without intention.

You can introduce misalignment even with a very narrow focus (https://arxiv.org/html/2502.17424v2), but it doesn't happen by accident.

So if your goals with distillation weren't centered around censorship, and weren't meant to introduce censorship, then why are you trying to draw this tenuous link?


I guess the AI that wrote your comment for you also conflated the SFT step of the target domain with the political prompts, which, in the sentence you quoted, contradicts your original comment...


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: