The future of business, today.
RSSNewslettersAdvertise
Business Future Today

AI competition

Garry Tan Wants an ‘American Distillation Regime’ for Open-Weight AI

Y Combinator’s Garry Tan is arguing that U.S. open-weight AI labs should be allowed to learn from frontier models through authorized access—a position that sharpens the fight over APIs, model competition and training data rights.

Garry Tan Wants an ‘American Distillation Regime’ for Open-Weight AI

Y Combinator CEO Garry Tan is urging regulators not to crack down broadly on AI model distillation—and says U.S. open-weight labs should be able to use the technique on domestic frontier models through legitimate access.

The proposal arrives as Anthropic alleges that Chinese labs have conducted “illicit distillation attacks,” including by concealing identities and using fraud or stolen credentials. Anthropic CEO Dario Amodei has called for U.S. regulatory action. Tan’s position is narrower than a defense of those alleged tactics: he says American labs should be able to use frontier-model outputs without using stolen credentials or deceptive access.

What changed

Tan told CNBC that regulators should “do nothing” about distillation and suggested an “American distillation regime.” He later told TechCrunch that smaller U.S. open-weight labs should be free to apply similar training techniques to American frontier systems, provided they come “in the front door.”

Distillation generally involves using one model’s responses to help train another model. It is an established training technique, but it becomes contentious when a company uses extensive API querying to reproduce capabilities that a frontier-model provider considers proprietary—particularly when the provider’s terms prohibit that use.

Tan’s argument is fundamentally about who controls the outputs of a paid model service. He contends that closed-model companies should have less ability to restrict what customers do with information returned through API calls. He also draws a parallel with the broad, often copyrighted corpus material used in the development of today’s frontier models.

Why it matters for AI businesses

The debate is not just about training methods. It is about whether API access is a product, a permissioned gateway, or a channel through which competitors can acquire usable model behavior.

For frontier-model providers, unrestricted distillation could weaken a core economic premise: enormous spending on chips, research and data infrastructure can be recouped through proprietary model access. If competitors can systematically translate outputs into smaller, cheaper or openly available alternatives, providers may tighten rate limits, monitoring, contract terms and customer vetting.

For open-weight labs and enterprise builders, Tan’s preferred framework could expand the supply of models that can be run, adapted and deployed with more control than a hosted API allows. That matters to organizations with data-residency, latency, cost or customization requirements. It could also reduce dependence on a small group of well-capitalized model vendors.

But an authorized-distillation regime would be difficult to define operationally. Providers would need clarity on what counts as normal use, evaluation, synthetic-data generation, benchmark creation or competitive extraction. Policymakers would also have to distinguish conduct involving deceptive identity or compromised credentials from activity carried out under disclosed, paid access.

The tension underneath

Tan says the risk is a market where one proprietary company accumulates superior capital, talent and model capability, leaving customers without meaningful alternatives. At the same time, he says frontier labs should remain fundable and viable businesses.

Those goals can conflict. The more easily outputs can be used to train rivals, the more valuable open access becomes to the ecosystem—and the less exclusive a frontier provider’s advantage may be.

What to watch next

The immediate signal will be whether AI providers revise API terms, technical controls or enterprise contracts to explicitly address model-to-model training. The policy question is likely to turn on the same distinction Tan emphasizes: sanctioned use through legitimate access versus alleged theft, fraud and concealment.

For operators, the practical takeaway is to review vendor terms before using model outputs in fine-tuning, synthetic-data pipelines or evaluation datasets. The legal and commercial boundary around those workflows is becoming a competitive issue, not merely a compliance detail.

Sources

STAY AHEAD

The future of business, in your inbox.

Useful signals on the companies, technologies and shifts changing business.

One useful briefing. Unsubscribe any time.