Snorkel AI has raised a $350 million Series E at a $3.5 billion valuation, nearly tripling its $1.3 billion valuation from a $100 million Series D 17 months ago.
The round, led by Insight Partners and S32, is a strong signal that investors see training data—not only models and compute—as a strategic layer of the AI market. Existing investors Addition, Lightspeed, Greylock, GV and Wells Fargo also participated.
What changed: Snorkel is selling outcomes, not just labeling tools
Snorkel began as a provider of software intended to automate data labeling. Last year, it shifted toward what it calls data-as-a-service: delivering completed datasets and simulated environments to AI labs and enterprise customers.
Its model is hybrid. Snorkel uses software and models to generate synthetic data, alongside subject-matter experts who contribute domain knowledge. The company also sells reinforcement-learning environments, which can be used to train and evaluate model behavior beyond conventional static datasets.
That change matters because customers increasingly want usable training inputs quickly, rather than a toolkit that requires them to assemble labeling workflows, recruit experts and manage quality control themselves.
Snorkel says its annualized revenue run rate is now $375 million, up 18-fold over the past 12 months. The company launched commercially in 2019 after research conducted by co-founder and CEO Alex Ratner and his team at Stanford.
Why the market is rewarding data suppliers
As frontier and enterprise AI systems improve, readily available internet data becomes less sufficient for specialized use cases. Builders need data that reflects particular industries, workflows, policies and edge cases—and they need ways to measure whether models perform reliably in those settings.
That creates demand for three related capabilities:
- **Domain-specific data**, particularly for high-value enterprise tasks.
- **Synthetic data generation**, which can expand scarce or sensitive datasets.
- **Training and evaluation environments**, especially for reinforcement learning and agent-like systems that must take actions rather than simply produce text.
Snorkel’s funding follows rapid reported growth at other companies supplying AI training inputs. TechCrunch cited Mercor, Handshake and Micro1 as examples of the category’s expansion. But operators and investors should be careful when comparing headline run rates across the group: companies that primarily broker human expert work may report gross revenue that includes substantial payments to contractors. TechCrunch notes that these payouts can account for roughly 60% to 70% of top-line revenue at some firms.
Snorkel says its expert payments are recorded in cost of goods sold because it sells datasets and RL environments rather than human labor itself. That accounting distinction does not settle questions about underlying margins, but it does reinforce that “AI data” is not a uniform business category.
The operational implication
For enterprise AI teams, the funding validates a practical procurement trend: data preparation, simulation and evaluation may be bought as managed infrastructure rather than built entirely in-house.
That can shorten time to deployment, but it raises diligence requirements. Buyers should ask how data is sourced, what portion is synthetic, how expert contributions are validated, whether environments resemble real production conditions, and how performance gains transfer from testing to live workflows.
For founders, the opportunity is less about generic annotation and more about owning a difficult feedback loop: access to expert knowledge, a repeatable way to turn it into training material, and credible evaluation of model behavior.
What to watch next
The key question is whether data-as-a-service vendors can turn exceptional demand into durable, defensible businesses. Watch for evidence of recurring enterprise deployments, retention, gross-margin progression and differentiated evaluation or simulation products.
Also watch whether AI labs increasingly internalize this work. If they build proprietary data pipelines and environments, vendors will need to win on speed, specialist supply, quality systems and coverage of domains customers cannot efficiently replicate themselves.




