The future of business, today.
RSSNewslettersAdvertise
Business Future Today

AI & Health

OpenAI Is Funding a New Supply Chain for Biology Data

The OpenAI Foundation is backing a new health-data initiative, including an effort to preserve biotech records from failed companies. The bet is that better datasets—not just bigger models—are a constraint on AI’s usefulness in drug development.

OpenAI Is Funding a New Supply Chain for Biology Data

The OpenAI Foundation has launched Public Data for Health, an initiative to fund the creation and collection of scientific datasets intended for AI use in medicine.

Its first grants illustrate a practical view of where health AI is constrained: not merely by model capability, but by access to structured, high-quality evidence about biology, clinical development and drug regulation.

Among the awards is a $500,000 grant to 1Day Sooner, an advocacy group for clinical-trial volunteers, to test whether it can acquire and preserve data from bankrupt biotech companies. The Foundation also said it would provide $40 million to a University of North Carolina, Chapel Hill program collecting data on novel cancer vaccines, and support OpenAdmet, which runs drug-effect prediction competitions.

A bid to recover biotech’s “lost archive”

The 1Day Sooner project is based on an idea from policy analyst Ruxandra Teslo: acquire nonexclusive copies of records sold during biotech bankruptcies and make them useful for research.

The target materials, known as common technical documents, can include correspondence with regulators, manufacturing plans, safety information and detailed clinical measurements. They can represent a fuller account of a drug program than is available in public filings or academic literature—especially when a company fails before its work is published.

For drug developers, the potential value is operational. A well-governed archive could help researchers and AI systems identify precedents, prepare regulatory submissions, understand why development programs failed, and design trials with a more complete picture of safety and manufacturing constraints.

Teslo argues that clinical development and regulatory work account for roughly 70% of drug-development time and spending. That makes the administrative and evidentiary layer of medicine a potentially important target for AI assistance—not just molecule discovery.

The strategic shift: data infrastructure as AI infrastructure

The initiative signals a growing recognition that frontier models trained on general internet and scientific text may not have the specialized, traceable datasets required for high-stakes biological work.

For startups building in life sciences, this creates both an opportunity and a higher bar. Differentiation may come less from simply attaching a model to an existing workflow and more from assembling data rights, provenance, domain-specific evaluation and integration with regulated processes.

It also suggests that philanthropic capital could influence which foundational health datasets become broadly available. The OpenAI Foundation says it hopes to distribute $1 billion by the end of the year. Its ability to turn grants into durable public resources will depend on details not yet established publicly: data-access rules, privacy safeguards, standards for de-identification, licensing, governance and who can train commercial models on the resulting corpus.

The difficult part is not only buying the data

1Day Sooner already holds three datasets, two donated by Lumen Bioscience, but its president Josh Morrison said two other attempts to acquire drug-company files this year were unsuccessful. Bankruptcy auctions, contractual restrictions and competing bidders can all limit access.

The broader environment is becoming more contentious. Corporate data from bankrupt businesses is increasingly viewed as a training-data asset, raising questions about confidentiality, ownership and whether employees, customers or research participants have meaningful control over its reuse.

Health data makes those questions more acute. Regulatory files may contain sensitive patient-level, proprietary or commercially consequential information. Any archive designed to support AI will need clear separation between material that can be shared, material that can be accessed only under controls, and material that cannot be repurposed.

What to watch next

The immediate test is whether the Foundation-funded pilot can consistently obtain datasets and convert heterogeneous, confidential records into a usable resource without compromising privacy or trade secrets.

For operators, the more consequential question is whether Public Data for Health creates genuinely reusable public infrastructure—or primarily funds bespoke collections that remain hard to access. In biology, the model race may increasingly be a data-governance and data-production race.

Sources

STAY AHEAD

The future of business, in your inbox.

Useful signals on the companies, technologies and shifts changing business.

One useful briefing. Unsubscribe any time.