AI’s biggest near-term business questions are converging around two inputs: compute capacity and differentiated data.
MIT Technology Review reports that hyperscalers’ AI data-center expenditures are expected to approach $1.1 trillion through 2027. Separately, the OpenAI Foundation has announced funding for an effort to create high-quality scientific datasets from information associated with failed biotech companies. The initiatives address different layers of the AI stack, but each carries the same underlying challenge: turning costly inputs into measurable outcomes.
The infrastructure bill comes due in productivity
The scale of planned data-center construction shifts the AI conversation away from model demos and toward capital efficiency. University of Pennsylvania finance professor Jessica Wachter’s framing, as reported by MIT Technology Review, is not a forecast of how broadly AI will be deployed. It asks how quickly hyperscalers’ earnings would need to grow to justify their spending.
The implication is demanding: AI companies would need an extraordinary increase in productivity to break even by 2030.
For operators, that raises the bar for internal AI programs. A useful deployment cannot merely save a few minutes per employee or produce better-looking output. At scale, buyers will increasingly be pressed to show durable gains in throughput, revenue, quality, risk reduction, or cost structure. Vendors, meanwhile, will face sharper scrutiny over inference costs, customer retention, and whether their products are embedded in repeatable workflows rather than used as occasional assistants.
This does not mean the buildout is necessarily misplaced. Data centers support a broad portfolio of workloads, and the payoff can arrive unevenly. But the spending trajectory makes the burden of proof concrete. The more capital committed before reliable enterprise demand is established, the less room there is for weak utilization or undifferentiated products.
Biology highlights the value of scarce data
The OpenAI Foundation-backed data initiative points to a different constraint. In biology and drug development, models need more than general-purpose internet-scale text. They need detailed, structured, domain-specific information—often expensive to generate and difficult to access.
The proposal, first put forward by policy analyst Ruxandra Teslo, would seek to preserve useful material from failed biotech companies, including regulatory filings, manufacturing strategies, and safety data, potentially through bankruptcy proceedings. Teslo characterized the opportunity as creating “biotech’s lost archive.”
That approach matters because failed programs can still contain valuable negative results, safety observations, and operational lessons. Such records may help researchers avoid repeating dead ends and may give AI systems a richer picture of experimental outcomes than published successes alone provide.
For founders building vertical AI products, the message is straightforward: durable advantage may come less from access to a frontier model than from rights to specialized, well-governed data and the ability to turn it into a reliable workflow. In regulated sectors, data provenance, consent, access rights, and validation will matter as much as model performance.
What to watch next
Three signals will determine whether these bets mature into sustainable businesses:
1. Evidence of economic output. Watch for enterprise deployments tied to specific KPIs rather than broad claims of AI adoption. 2. Utilization and unit economics. The industry needs signs that new compute capacity is being used productively enough to support its cost. 3. Data governance in scientific AI. The biology-data effort will need to resolve how records are acquired, standardized, accessed, and responsibly reused.
The AI buildout is no longer only a race to train larger models. It is a test of whether massive infrastructure investment can translate into operating leverage—and whether carefully assembled proprietary data can create breakthroughs that general-purpose models cannot reach on their own.




