Artificial Intelligence Underwriting Company (AIUC), a startup founded by early Anthropic employee Rune Kvist and former METR COO Rajiv Dattani, has raised a $40 million Series A to build an audit and certification market for AI agents.
The financing, led by Ribbit Capital with participation from First Harmonic, follows a $15 million seed round. AIUC says it now has $55 million in total funding and names Cursor, Lovable, Harvey and ElevenLabs as customers.
Its core proposition is straightforward: enterprises may be willing to use capable AI agents, but often lack a credible way to verify what those systems will do when they encounter sensitive data, adversarial prompts or unusual operating conditions.
What AIUC is building
AIUC has created a standard called AIUC-1, modeled in part on the role that SOC 2 plays in cybersecurity procurement. The company tests an AI agent against roughly 5,000 scenarios covering issues including jailbreaks, hallucinations and data leakage.
The resulting assessment is a report of about 100 pages that identifies the conditions in which an agent appears to behave safely and reliably, as well as identified failure modes. AIUC uses agents and other AI systems to conduct and analyze the testing, while humans verify the final audit, according to the company.
The standard is being informed by a consortium of about 250 security and risk leaders, AIUC says. That matters because agent assurance will be most useful if it maps to the questions buyers already ask during vendor review: what access an agent has, what commitments the supplier can make, what could go wrong, and how those claims were tested.
Why it matters for operators
The limiting factor for agent adoption is increasingly operational assurance rather than raw model capability. A company may see clear value in an agent that can search internal systems, prepare customer communications or execute workflow steps. But deploying it can create new risk pathways: leaked information, unauthorized actions, manipulated instructions, or confidently wrong outputs that enter a business process.
Most organizations can impose controls around an agent—permissions, logging, human approvals and limited tool access. What they have lacked is a standardized external view of the agent’s behavior before and during procurement. AIUC is trying to supply that missing layer.
For vendors, a recognizable independent assessment could reduce the burden of responding to bespoke customer security questionnaires. For buyers, it could create a common basis for comparing agents without assuming that a certification eliminates risk.
That last distinction is important. Testing demonstrates behavior under specified conditions; it does not guarantee that an agent will never fail in a new environment. Enterprises should treat an audit as evidence for a risk decision, not as a substitute for deployment controls, incident response plans and ongoing monitoring.
A growing market for independent evaluation
AIUC’s approach overlaps conceptually with the work of organizations such as METR, which has evaluated advanced AI systems, including their ability to reliably complete tasks. The company’s focus, however, is a commercial certification product designed for enterprises buying or deploying agents.
The timing is notable. Anthropic CEO Dario Amodei recently argued for more third-party scrutiny of frontier AI development amid reported increases in agent misbehavior. That debate is centered on powerful model developers, while AIUC is targeting the enterprise purchase and deployment layer.
What to watch next
The key question is whether AIUC-1 becomes credible enough to function as a real procurement signal. That will depend on the transparency of its methodology, whether results are repeatable, how frequently certifications are refreshed as models and tools change, and whether large buyers begin requesting the standard.
It will also depend on scope. Agent risk is not solely a model problem: permissions, integrations, data handling and human workflow design can materially change outcomes. The firms that turn evaluation results into concrete deployment requirements—not just a report in a vendor file—will get the most operational value from this emerging assurance category.




