ByteDance’s Ban on Distilling Rival AI Models unrelated to U.S. regulatory concern and dates back to 2023: Report
Zhang Yiming’s decision was driven by a belief that ByteDance’s Seed lab must develop frontier intelligence on its own—even at the cost of falling behind in the short term, Guixinren reports today
ByteDance began restricting the use of outputs from rival artificial-intelligence models in its own training data as early as 2023, according to a Chinese-language report today by 硅星人 Guixinren, showing that founder Zhang Yiming’s opposition to model distillation long predates the recent political controversy surrounding the practice.
The account confirms The Information’s scoop yesterday that the Chinese technology giant, best known internationally for creating TikTok, banned distillation, the process of training a new model using outputs from existing, most advanced models.
However, the Chinese account also contradicts The Information’s report, which says the decision is partly intended to shield the TikTok founder from renewed regulatory scrutiny in Washington.
The Information reported, “one reason, say people close to the company, is ByteDance’s complicated history with the U.S. government over the company’s TikTok video app.”
According to the Chinese account, written by technology journalist 骆轶航 Luo Yihang and based on anonymous people familiar with ByteDance’s internal deliberations, TikTok’s future carried “almost zero” weight in the company’s decision.
The Information reported that Zhang Yiming said at an internal meeting that ByteDance “should be willing to sacrifice some short-term gains for longer-term goals,” without elaborating what the long-term goals are.
The Guixinren confirmed the quote, but added that Zhang’s position is that ByteDance’s Seed AI lab should be prepared to accept temporary setbacks rather than build its models by systematically learning from the outputs of stronger competitors.
“We can accept being temporarily behind, but do not distill,” Zhang has told the Seed team on several occasions, according to the Guixinren report.
ByteDance’s policy, according to the report, is that Seed researchers may conduct self-distillation using ByteDance’s own models, but they are not supposed to turn either proprietary or open-weight models developed by competitors into “teachers” for ByteDance’s systems.
A policy that began before the controversy
The policy can be traced to the early days of ByteDance’s large-model effort.
In 2023, when the company’s AI team was still experimenting with different approaches, some engineers used data generated through OpenAI’s GPT application programming interface in research involving a smaller model, according to the Chinese report.
In April of that year, ByteDance introduced checks governing the use of the GPT API. The model team then issued a clear internal requirement: GPT-generated material must not be included in the training data for ByteDance models.
The company subsequently conducted another internal review, sampling outputs from its models and testing their similarity to GPT responses. The purpose, according to the report, was partly to detect whether data annotators had privately used GPT to produce material that later entered ByteDance’s training pipeline.
Those steps were taken well before “distillation” became one of the most politically charged terms in the U.S.-China AI rivalry.
More recently, Anthropic has accused several Chinese AI developers of using large numbers of accounts and proxy networks to extract outputs from Claude for model training. The companies named by Anthropic have included Alibaba, DeepSeek, Moonshot AI, MiniMax and Z.ai, although Anthropic has not made the same allegation against ByteDance.
Last month, Michael Kratsios, the director of the White House Office of Science and Technology Policy, publicly accused Moonshot AI, the Chinese AI lab, of improperly using Anthropic’s latest model to train Kimi K3, an open-source model with near-frontier capabilities.
The account by Guixinren suggests ByteDance’s position was established before those allegations—and before avoiding U.S. retaliation became an obvious commercial argument against distilling American models.
The first internal battle
The 2023 rule did not settle the issue permanently.
As competition intensified and Seed struggled to match the performance of leading language models, researchers repeatedly revisited whether the company should make an exception. According to the Chinese report, Seed held at least three major internal debates over distillation beginning in 2025.
The first came after DeepSeek released its R1 reasoning model in January 2025.
DeepSeek R1’s strong performance and relatively low reported training costs shook Silicon Valley and placed enormous pressure on other Chinese laboratories. Inside Seed, some researchers believed ByteDance possessed comparable talent, computing resources and research capabilities, yet had failed to produce a language model with similarly strong reasoning performance.
There was already industry speculation that DeepSeek might have used synthetic data generated by American frontier models during training. That claim has not been independently established, and later allegations by Anthropic remain contested.
But inside Seed, the possibility produced a practical question: if competitors could use outputs from the most capable U.S. models to improve reasoning performance rapidly, why should ByteDance refuse to do the same?
Some researchers proposed incorporating generated responses from leading closed American models into Seed’s training process. They did not propose abandoning ByteDance’s own pretraining system or building an entire model on a competitor’s output. Rather, they viewed the data as a supplement that could quickly improve reasoning and general performance while preserving Seed’s existing technical framework.
The proposal was rejected.
The Blackwell gap
A second, more urgent debate emerged around the end of 2025 and the beginning of 2026, as Nvidia’s Blackwell generation of graphics processors was deployed at scale by leading U.S. AI laboratories.
U.S. export restrictions prevented Chinese companies from purchasing Nvidia’s most advanced chips, widening a computing gap that had already existed between American and Chinese developers.
According to Luo’s Chinese report, ByteDance trained Seedance 2.0—its highly regarded video-generation model—using large numbers of Nvidia H20 processors, a lower-performance chip designed to comply with export restrictions for the Chinese market. H20’s overall model-training performance was only a fraction of that of B200.
Whether or not the precise ratio applies across different workloads, Seed researchers believed the arrival of Blackwell had sharply increased the structural disadvantage facing Chinese laboratories.
The report said they watched as a new generation of American models made advances in reasoning, coding and scientific tasks, developments they associated in part with the large-scale deployment of Blackwell chips.
Within Seed, the argument that “if all else fails, we should distill” began to attract broader support.
The pressure was particularly acute because ByteDance had demonstrated that it could compete at the global frontier in video generation. ByteDance’s flagship language models had limited visibility on international leaderboards, making it difficult for outsiders to evaluate their capabilities.
Researchers who supported distillation argued that ByteDance could compensate for its compute disadvantage with better data. Instead of spending scarce training resources exploring every possible path, Seed could use frontier-model outputs to identify approaches that had already proved effective.
Top management of ByteDance again withheld approval, although the idea was reportedly gaining support inside the laboratory.
Kimi K3 forced a decision
The third and most intense debate followed the release of Moonshot AI’s Kimi K3 this year.
K3’s performance in coding, tool use, deep research and complex tasks placed it in the same broad competitive range as leading proprietary American models.
For ByteDance, the comparison was uncomfortable. Moonshot was a much smaller Chinese company with fewer financial and computing resources. Yet it had produced an open-weight language model that appeared to belong in the global first tier, while Seed, despite greater investment and a larger talent pool, had not produced a language model with a comparable position.
Researchers again proposed systematic use of outputs from frontier models. By generating large volumes of reasoning traces, code, tool-use examples and complex-task responses, they argued, Seed could quickly improve the capabilities that mattered most for benchmarks and user perceptions.
This time, an apparent compromise also emerged.
If distilling closed American models could breach terms of service, attract public accusations or intensify geopolitical tensions, Seed could limit itself to open-weight models. Those models could be deployed on ByteDance’s own servers without creating large numbers of accounts or circumventing geographic access restrictions. Depending on the licence, they also presented fewer legal and commercial risks.
Zhang, the ByteDance founder, rejected the compromise, said Luo’s report via Guixinren.
Closed models could not be distilled, he said. Open-weight models could not be distilled either.
Seed, he argued, could accept being behind for a period of time, but it should not eliminate that disadvantage by training on a competitor’s capabilities.
That internal dispute was one reason for the recent Seed all-hands meeting at which Zhang spoke, according to the Chinese report. The Information reported yesterday that Zhang, who had rarely addressed earlier Seed meetings, told employees the company “should be willing to sacrifice some short-term gains for longer-term goals.”
The Chinese account says those “longer-term goals” were not a reference to protecting TikTok, but to developing genuine frontier intelligence.
Following the meeting, Seed introduced a more explicit internal policy prohibiting the distillation of open-weight models including K3, the Chinese report said. ByteDance also began using API-related technical checks and other methods to identify and trace suspected distillation activity.
Click here for 骆轶航 Luo Yihang’s report in Chinese via 硅星人 Guixinren.


