Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it (www.ctgt.ai)

🤖 AI Summary
In a recent exploration of AI model distillation, researchers investigated whether censorship characteristics from a Chinese frontier model, DeepSeek V4 Flash, would transfer to an American model, GPT-OSS-120B. Their study revealed that despite training on outputs from a censored teacher aimed at enhancing financial reasoning, the distilled model did not exhibit any of the same censorship behaviors, particularly towards sensitive topics related to China. This finding is significant as it challenges the prevailing assumption that undesirable influences automatically transfer during distillation processes. The researchers employed a methodology called LineageEval, which involved judging responses across 304 matched prompts, and found that the distilled model maintained neutrality, scoring similarly to a base model without exposure to the censored content. This research not only demonstrates the potential for effective distillation without inheriting biases but also highlights the efficiency of the distilled model in specific tasks. The GPT-OSS-120B achieved an impressive performance score of 83.61% on finance reasoning queries while reducing operational costs significantly compared to competing models. By establishing that censorship did not transmit through the training data, the study opens the door for further investigation into distillation practices and raises the question of how to leverage and improve upon external model influences without adopting their limitations, making it an essential contribution to the AI/ML community.
Loading comments...
loading comments...