America and China Are Betting on Two Different Futures for AI

The real question worth asking is not which country trained on whose data. It is a question of which model actually provides better AI benefits.

When Moonshot AI released Kimi K3 in July, the model did something nobody in the AI industry expected. It matched or beat the best closed models from American labs on some of the toughest benchmarks within the industry, tasks including advanced coding, multi-step automation, and spreadsheet reasoning. For a Chinese startup working with a fraction of the computing budget available to its American rivals, that result stopped many people in their tracks. Developers who downloaded the open weights within hours of release were running their own tests before most industry analysts had finished their morning coffee, and the early consensus was consistent: this was not a model that looked good on a curated leaderboard; it held up under real use.

Within days, the story took a sharp turn. A senior White House science advisor accused Moonshot of secretly training Kimi K3 by feeding it outputs from Anthropic’s flagship model, a technique known as distillation, and doing so while using chips that are barred from export to China. The Treasury Department followed with a threat of sanctions if the claims held up. Suddenly a routine technical process had become the center of a geopolitical dispute.

There is a problem with the accusation, though, and it is worth explaining before drawing any conclusions. Distillation is not a shortcut invented in a back room. It is one of the most common tools of modern AI development, used openly by nearly every major lab on the planet, including American ones.

Distillation is standard practice, not a smoking gun

Distillation means training a smaller or newer model using the outputs of a larger, more capable one, essentially letting the student learn by watching the teacher work. The idea predates the current AI boom by nearly a decade, and it has become one of the most efficient ways to build a strong model without the full cost of training one from scratch. Done properly, it is a normal and widely disclosed part of building AI systems.

OpenAI even offers it as a commercial product, letting developers pay to compress one of its own large models down into a smaller, cheaper version for their own use. Grok’s developer has said publicly that his team distilled outputs from OpenAI’s models during its own training process and called the practice common across the industry. Chinese labs, including DeepSeek, have published detailed papers describing how they use distillation internally: training specialized versions of a model and then merging their strengths into a single, more capable system. The technique shows up everywhere once you start looking for it, which is part of why singling out one lab for using it reads as selective.

Independent researchers who looked closely at the Kimi K3 timeline found the accusation hard to square with the facts. Anthropic’s own model had been publicly available for only a few weeks before Kimi K3 shipped, not nearly enough time to extract the volume of data that a real distillation effort would require, retrain a frontier-scale model, and release it. Several analysts have stated that if Moonshot wanted to distill a Western model, there were easier, cheaper, and less detectable targets than the one Washington named. The line between distillation and simply building a good synthetic training dataset is blurry enough that pinning down exactly what happened, if anything unusual happened at all, may never be fully resolved by outsiders.

None of that means the underlying concern about protecting original research is unreasonable. It means the specific accusation says less about theft and more about anxiety, the discomfort of watching an open model such as Moonshot close the gap with a proprietary one like Anthropic faster than anyone expected.

A robot plays table tennis during the 2026 World AI Conference and High-Level Meeting on Global AI Governance in Shanghai, east China, Jul. 18, 2026. (Photo/Xinhua)

Two ladders, two bets on who gets to climb

The more interesting story here is not about who trained on whose outputs. It is about two different bets on what AI is for and who should get to use it.

Picture two companies each building a ladder up toward a very tall structure that keeps getting taller. In one version, the builder keeps adding rungs to the ladder, but places a toll booth behind them as it goes. The bottom rungs are free to try, but if you want to reach the higher ones, the part where the real capability lives, you pay a monthly fee. That is how the leading American labs have approached the business of AI. ChatGPT, Claude, Google, and Perplexity all offer free tiers, then gate their most capable models behind subscriptions because the companies building them need to recoup enormous R&D and compute costs and because they see the model itself as the product worth protecting.

In the other version, the builder keeps adding rungs too, but never puts up a toll booth, no locked gate, no fee to keep climbing. That is the model Chinese labs like Moonshot, DeepSeek, and Qwen have largely followed, releasing full open-weight versions of frontier-class systems that any developer, anywhere, can download, modify, and build on for free. Amazon has even added several of these Chinese open-weight models to its cloud platform this year specifically because businesses wanted frontier-level performance without the licensing costs associated with closed alternatives.

Neither approach is purely altruistic. Chinese labs benefit from open releases too, through faster adoption, developer goodwill, and influence over how AI gets built globally. But the underlying philosophy is genuinely different. One approach treats the model as intellectual property to be handed out at a cost. The other treats it as infrastructure that becomes more valuable to everyone the more people use it.

The effect of that difference is already showing up in who actually uses these systems. A small business in Lagos, a research lab in Jakarta, or a single developer in Sao Paulo cannot always justify a monthly subscription to a frontier American model. Still, they can download an open-weight Chinese one for free and run it on whatever hardware they already own. That is not a minor side effect. It is the fastest path AI has had yet to reaching people who were priced out of the first wave of this technology entirely.

What actually matters for the industry’s next stage

Eric Schmidt, former CEO of Google has just said that the largest American AI model is close-source with charge while the largest Chinese AI model is open-source and free. He estimated that most governments and countries would ultimately follow Chinese AI models.

The distillation argument will likely fade the way most of these disputes do, without a clean resolution, more heat than proof. What will not fade is the pressure that open-weight models exert on the industry’s economics. When a developer can get frontier-class performance for free, it becomes harder to justify high subscription prices for a closed alternative that performs only slightly better.

The real question worth asking is not which country trained on whose data. It is a question of which model actually provides better AI benefits: a system where the most capable tools sit behind a paywall that most of the world cannot afford, or one where the ladder stays open for anyone willing to climb it. That question will shape the next stage of this industry far more than any single accusation about how one model got built.