A Reuters review of over 80 Chinese academic papers and patents, compiled by the Jamestown Foundation, documents how Chinese military researchers have used API access to Western AI models like GPT-3.5 and Claude to train smaller domestic models for military applications. Cases include a PLA unit distilling military code summaries and researchers building drone targeting systems using the technique.
The findings challenge the premise that hardware export controls are sufficient to limit China’s military AI development. Because model distillation requires only API access rather than restricted chips, it represents a gap that chip-focused policy cannot close, raising unresolved questions about how AI providers should govern access to their models ahead of planned US-China AI safety talks.
PLA Unit 96941, a military intelligence and cyber warfare outfit based in Beijing, had a problem last year: it needed a model that could summarize sensitive military source code, but it couldn’t risk feeding that code straight into GPT-3.5 through OpenAI’s API. So its researchers did something simpler than you’d expect. They ran ordinary, unclassified code through GPT-3.5 to generate summaries, then used those summaries to train a smaller domestic model that runs entirely inside Chinese military networks, no API calls once it’s live, no data leaving the building. That’s one case study out of more than 80 Chinese academic papers and patents a Reuters review turned up this week, working off research compiled by the Jamestown Foundation’s Sunny Cheung, and it’s a much sharper picture of the distillation problem than the usual “China is copying our models” framing gets you.
Distillation itself isn’t the scandal. Training a smaller “student” model on a bigger “teacher” model’s outputs is standard practice everywhere, it’s how you get a cheap model that’s still useful without burning frontier-scale compute on every inference call. What’s landed this at the center of a diplomatic fight ahead of September’s US-China AI safety talks is who’s doing it and why. The White House’s Office of Science and Technology Policy was already alleging “industrial-scale campaigns to distill US frontier AI systems” back in April, OpenAI and Anthropic have both published pieces on catching distillation attempts, and China’s response has been to call the accusations American “AI hegemonism” while threatening its own countermeasures. Moonshot AI spent last week publicly denying that Kimi K3, its answer to Anthropic’s Fable 5, was built by distilling it. Everybody involved has an incentive to shade the truth a little, which is exactly why the Jamestown paper trail matters here. It isn’t accusations, it’s citations.
Because once you get past the summarized-code example, the pattern in these papers stops looking like coding assistants and starts looking like actual weapons development. Researchers at the North University of China, an institution with direct ties to the country’s weapons industry, used Anthropic’s Claude 3 Haiku to generate synthetic training data for a text classification model built for social media monitoring. Anthropic says flatly it doesn’t sell commercial access to Claude in China or to Beijing-controlled firms, and that distilled models can lose the safety guardrails baked into the original, which is a polite way of saying it has no idea what these derivative models actually do once they’re out of its hands. A 2024 paper from the PLA’s National University of Defense Technology used distillation to shrink an image processing model small enough to run on a drone, letting it analyze live video for navigation and targeting even when its communications link is cut. Another study, from China’s Academy of Military Sciences, used the same approach to run a target recognition model on tactical hardware during simulated maritime operations involving drones, ships, and unmanned submarines. None of that requires a data center full of smuggled Nvidia silicon. It requires an API key, some patience, and a teacher model that already did the hard, expensive work of learning to reason.
This is the part I keep sitting with. I’ve spent a lot of time on this blog tracking the physical side of the export control fight: the Nvidia chip smuggling arrests in Taiwan, the loophole routes through Malaysia, Qualcomm quietly building an edge silicon business the export regime can’t touch, and the whiplash of the H20 export license getting reversed mid-moratorium. All of that assumes the choke point is hardware, that if you keep the GPUs out of Chinese hands you’ve slowed the program down. What these papers show is a workaround that doesn’t need a single restricted chip to cross a border. GPT-3.5 isn’t even a frontier model anymore, it’s two generations behind and effectively free to query. If a research unit inside the PLA can turn cheap API access to an obsolete Western model into a working targeting system for unmanned submarines, the entire premise of chip-centric export controls, that compute is the bottleneck, has a hole in it that no chip policy addresses.
China’s counterargument, that Washington is pursuing AI hegemonism and that US firms do the same thing, isn’t crazy on its face. Distillation cuts both ways, and every lab benchmarks against every other lab’s outputs to some degree. But there’s a real difference between a startup fine-tuning on a competitor’s outputs to ship a better chatbot and a PLA cyber warfare unit training a model on military source code specifically because it can’t trust a foreign API with the real thing. I’m not going to pretend I know how the September talks resolve this, model distillation doesn’t have a clean enforcement mechanism the way chip export licenses do, you can’t put a shipping manifest on an API response. What I do know is that every company selling API access to a capable model is now, whether it likes it or not, in the business of deciding who gets to learn from it, and OpenAI staying silent when Reuters asked for comment says more than a statement would have.
Sources
- Reuters, “Exclusive: Chinese military researchers tap US AI models to train defence systems”, July 31, 2026
- Jamestown Foundation, “Chinese Research Details Distillation for Military Use” by Sunny Cheung, July 30, 2026
- Reuters, “What is AI model distillation, and why is it becoming a US-China flashpoint?”, July 31, 2026
- White House Office of Science and Technology Policy, National Security Technology Memorandum 4, April 23, 2026