ByteDance is pre-training an AI model with 10 trillion parameters — a scale that would position a Chinese lab within striking distance of Anthropic's frontier Mythos class for the first time. The Financial Times reported the project Friday, citing people with knowledge of the matter. ByteDance did not comment. The parameter count is not final; 10 trillion is the working target.

The scale gap within China is stark. Moonshot's Kimi K3, the largest model shipped by a Chinese lab, carries 2.8 trillion parameters. Before Kimi K3 arrived in July, Meituan's LongCat-2.0 and DeepSeek's V4-Pro shared the domestic record at 1.6 trillion. Alibaba answered Kimi K3 within weeks with a 2.4 trillion-parameter Qwen flagship. ByteDance at 10 trillion would clear every one of them by a substantial margin and would pull within range of analyst estimates that put Anthropic's Mythos 5 at roughly 8 trillion parameters and its more broadly available Fable 5 at around 5 trillion. Anthropic and OpenAI publish no official parameter counts; those figures are analyst estimates, not confirmed specs.

Parameter scale comparison across frontier models — ByteDance's 10T target vs. Chinese and Anthropic peers (trillion parameters)
FIG. 02 Parameter scale comparison across frontier models — ByteDance's 10T target vs. Chinese and Anthropic peers (trillion parameters) — Financial Times; analyst estimates for Anthropic models (unconfirmed)
LabModelParameters (Trillion)Availability
MeituanLongCat-2.01.6Public
DeepSeekV4-Pro1.6Public / Open-weight
AlibabaQwen flagship2.4Public
MoonshotKimi K32.8Public
AnthropicFable 5~5 (analyst est.)Broad release
AnthropicMythos 5~8 (analyst est.)Restricted — vetted partners only
ByteDanceUnnamed (target)10 (working target)Proprietary (planned)
FIG. 03 Large-model parameter counts: Chinese labs vs. Anthropic estimates (2025–2026) — Financial Times reporting; analyst estimates — Anthropic publishes no official parameter counts

The model is in early pre-training, a stage that typically runs three to six months before fine-tuning begins. Moonshot used approximately 20,000 Nvidia chips to train Kimi K3. ByteDance's compute requirement for a 10-trillion-parameter run is proportionally larger—and unresolved. US export controls have restricted China's access to the most advanced accelerators. Whether ByteDance can source enough compliant hardware to finish training on a reasonable timeline remains an open question.

Typical large-model training pipeline: stages from pre-training through deployment
FIG. 04 Typical large-model training pipeline: stages from pre-training through deployment — ai|expert

Anthropic's Mythos is not general-release. Access is restricted to vetted partners under Project Glasswing, given documented concerns about the system's capabilities in agentic hacking and vulnerability exploitation. Fable 5, the safety-hardened sibling, reaches a broader audience. ByteDance building toward Mythos-scale names the tier it is benchmarking against — not the widely available one.

Founder Zhang Yiming's internal message, issued days before the FT story, sharpens the intent. He told staff to stop relying on model distillation for short-term gains and to build genuine capability instead. That instruction lands in a live diplomatic dispute: the White House accused Moonshot of distilling Anthropic's Fable 5 to produce Kimi K3. Moonshot rejects the allegation. A 10-trillion-parameter pre-training run is the most expensive rebuttal available—it is physically impossible to distill your way to a model that large without a larger source. The compute commitment is also the argument.

Two structural advantages make ByteDance's bet more defensible than it would be for a pure research shop. Its Doubao assistant already leads China's consumer AI market, supplying both training signal from scale and a ready deployment channel that can convert research spending into product revenue almost immediately. The company also keeps the weights of its strongest general models private, unlike most Chinese peers that ship open weights. A 10-trillion-parameter system, if it ships, stays proprietary.

ByteDance's structural advantages: Doubao data flywheel and proprietary weight strategy
FIG. 05 ByteDance's structural advantages: Doubao data flywheel and proprietary weight strategy — ai|expert, based on article body

For architects evaluating the China AI landscape: distillation-based catch-up strategies are politically untenable. ByteDance just bet that frontier compute is cheaper than the reputational cost of that path.