A wave of high-profile AI announcements—from cybersecurity incidents to purported mathematical advances and warnings of imminent superintelligence—has renewed a familiar problem for executives: the gap between a compelling narrative and independently verified capability.
In an MIT Technology Review opinion essay, DAIR executive director Timnit Gebru and University of Washington professor Emily M. Bender argue that companies and media outlets have too often presented claims in human-like terms that exaggerate what software systems actually did. Their core advice for policymakers applies equally to operators: slow down, obtain expert review, and assess the specific system behavior rather than the most expansive framing around it.
What changed
The essay points to a cluster of recent claims around AI capability, including Anthropic’s assertions about vulnerability discovery, disclosures involving model-related cybersecurity incidents, and announcements of mathematical results by Anthropic and OpenAI.
The authors argue that follow-up scrutiny has complicated some initial accounts. They cite cybersecurity experts who characterized the hacking stories as being more about failures in basic security controls than autonomous systems acting independently. They also point to criticism from mathematicians that some widely reported results were less novel than early announcements implied, alongside allegations concerning attribution and research conduct.
These are allegations and critiques, not a general finding that AI tools lack value. But they underscore an operational distinction that is frequently lost in launch-day coverage: a model producing a useful output in a constrained, evaluable setting is not the same as a dependable autonomous agent—or evidence of general intelligence.
Why this matters to operators
For buyers, the relevant question is rarely whether a model can perform an impressive task once. It is whether the system can do a defined job reliably, securely, economically and with clear accountability when it fails.
That means treating vendor statements about “agents,” “superhuman” performance or breakthrough discovery as hypotheses to test. In coding and mathematics, outputs can often be checked relatively cheaply: code can be run, tests can be executed, and mathematical work can be reviewed. That verification advantage can make those domains especially attractive for model training and demonstrations. It does not eliminate the need for independent validation in a production environment.
The governance issue is equally practical. Language suggesting that a “rogue model” acted independently can obscure responsibility for system design, access controls, deployment decisions and incident response. Organizations integrating AI should retain clear ownership of those decisions rather than allowing anthropomorphic product language to blur accountability.
A practical test for AI claims
Before scaling an AI deployment or changing strategy based on a vendor announcement, teams should ask:
1. What exactly was demonstrated? Define the task, data access, tools available, human involvement and success criteria. 2. Who verified it? Look for evaluation by domain experts who were not responsible for the product announcement. 3. What is the base rate? Request failure modes, repeatability data, error severity and performance on representative internal work—not only selected examples. 4. What controls were required? Examine permissions, sandboxing, monitoring, audit logs, escalation paths and the ability to revoke access. 5. Who is accountable? Identify the vendor and internal owners responsible for security, output quality, data use and remediation. 6. What are the full costs? Include integration work, review labor, infrastructure, compliance exposure and resource use alongside model pricing.
What to watch next
The most useful signal will be less dramatic than a breakthrough announcement: reproducible evidence of performance on real workflows, paired with transparent limitations and credible external review.
Leaders should also watch how AI claims shape procurement and policy. The authors argue that attention on speculative superintelligence can divert focus from nearer-term issues, including data-center impacts, energy costs, water use, local pollution and companies’ handling of data and security.
The business case for AI remains task-specific. Companies that build evaluation, controls and accountability into deployments can capture practical gains without making strategy dependent on the next viral claim.




