VnRobo
AboutPricingBlogContact
🇻🇳VISign InStart Free Trial
🇻🇳VI
VnRobo logo

AI infrastructure for next-generation industrial robots.

Product

  • Features
  • Pricing
  • Knowledge Base
  • Services

Company

  • About Us
  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 VnRobo. All rights reserved.

Made with♥in Vietnam
VnRobo
AboutPricingBlogContact
🇻🇳VISign InStart Free Trial
🇻🇳VI
  1. Home
  2. Blog
  3. What Robotaxi Fleets Actually Publish About Safety
adasadasautonomous-drivingself-drivingrobotaxisafety-dataavp

What Robotaxi Fleets Actually Publish About Safety

Series finale: Waymo's 220M miles, an IIHS report showing 68% fewer crashes, Apollo Go's 22M trips — and why methodology beats the headline number.

Nguyễn Anh TuấnAugust 23, 202613 min read

Over the five previous posts in this series, we covered VLA architecture, ran OpenDriveVLA on nuScenes, dissected NAVSIM v2 and Bench2Drive, watched Alpamayo-R1 reason with Chain-of-Causation, and saw world models shift from generating beautiful videos to measuring policy performance. All of it circled one implicit question: is this system safe?

This post does not answer that with another benchmark score. Instead, it asks a different question: what do robotaxi fleets actually operating on public roads publish, and how do we read those numbers correctly?

This skill matters more than knowing the algorithms — because when you see a headline saying "robotaxis are 94% safer than human drivers," the important question is not "do I believe this?" but "how was the denominator built?"


1. Waymo: 220.6 Million Miles and the Methodology That Makes the Difference

The Raw Numbers

Per Waymo's Safety Impact hub, through March 2026 the company has driven 220.6 million rider-only miles — fully autonomous, no safety driver behind the wheel. Breakdown by city:

City Miles (millions)
Phoenix 80.551
San Francisco Bay Area 67.078
Los Angeles 51.816
Austin 15.789
Atlanta 5.379
Total 220.6

Compared to human drivers in the same areas over the same period, Waymo reported:

  • 94% fewer crashes causing serious or fatal injuries (an estimated 47 crashes avoided)
  • 82% fewer crashes in which an airbag deployed (305 crashes)
  • 82% fewer crashes involving any reported injury (707 crashes)

For vulnerable road users:

  • 93% fewer injury crashes involving pedestrians (76 crashes)
  • 84% fewer involving cyclists (48 crashes)
  • 84% fewer involving motorcyclists (32 crashes)

Converting to incidents per million miles (IPMM):

Crash type Waymo Human benchmark
Serious injury or worse 0.01 0.23
Any reported injury 0.71 3.91
Airbag deployed (any vehicle) 0.30 1.68
Airbag deployed (Waymo vehicle) 0.06 1.11

The 0.01 vs. 0.23 figure — 23 times lower for serious crashes — is the most striking number. But it only means anything if you understand how the human benchmark was constructed.

How the Human Benchmark Is Built — the Part Worth Reading Carefully

Many safety reports use national or state-level statistics. Waymo does something more careful:

1. Local data only, from where Waymo actually operates. The human benchmark is calculated from police-reported crash records in the exact counties where Waymo runs, not statewide Arizona or California figures. This removes the bias of comparing dense urban streets with lower-risk suburban roads.

2. Spatial dynamic adjustment. Every neighborhood and road segment has a different risk profile. Waymo weights the geographic distribution of its trips to avoid comparing a robotaxi fleet concentrated in busy downtown corridors against a human benchmark that includes suburban driving.

3. Surface streets only, no freeways. The analysis is restricted to the type of road Waymo actually operates on. This avoids an artificial "advantage" from freeway crashes, which are disproportionately represented in human driver statistics.

4. Under-reporting correction per NHTSA Blincoe 2023. This is the most subtle point: police crash data misses a significant share of collisions because drivers simply do not report them. According to the NHTSA Blincoe study, roughly 32% of injury crashes go unreported. Waymo adjusts the human benchmark upward accordingly — without this correction, the human baseline would be artificially low, making Waymo's improvement look even larger than it is.

This creates a meaningful asymmetry: Waymo reports every incident it is required to report under NHTSA rules, while human drivers self-select what to report. Comparing the two without correcting for under-reporting means comparing a fully-reported numerator against an incomplete denominator.


2. IIHS 2026: Independent Analysis — and the Industry's Most Important Limitation

Authorship and Status

In July 2026, Eric R. Teoh, David G. Kidd, and Luke E. Riexinger of the Insurance Institute for Highway Safety (IIHS) published the first independent analysis of L4 crash rates: Rise of the machines: crash experiences of highly automated vehicles and human drivers.

On publication status: this is a report published by IIHS itself, not a peer-reviewed article in an independent journal. IIHS is an insurance-industry-funded research institute; its methods are serious and public, but the cross-checking here is internal rather than by outside reviewers. That is not grounds to dismiss the result — only grounds not to cite it as though it were a Nature paper.

Key Results

The analysis covered 50 million miles of Waymo driving from 2021–2024 across San Francisco, Los Angeles, Phoenix, and Austin, compared against 222 billion miles driven by human drivers in the same cities over the same years.

Overall finding: Waymo L4 vehicles in driverless mode had crash involvement rates 68% lower than human drivers, per VMT.

By city:

  • Phoenix: 76% lower
  • Los Angeles: 71% lower
  • San Francisco: 35% lower
  • Austin: 4% higher — but the sample size was small and the result is not statistically significant

For rear-end crashes specifically: Waymo was 91% lower than human drivers. This makes mechanical sense — L4 systems react faster than humans and are never distracted.

The 22% Figure — and What It Reveals About Structural Asymmetry

IIHS examined 736 crashes that occurred while vehicles were in automated mode. They found that only 22% of those crashes met the threshold a reasonable person would report to police under standard definitions.

This means: AV systems are reporting minor incidents that human drivers would typically handle without calling the police. This is not inherently bad — full reporting helps improve the system — but it creates a comparison asymmetry: the AV numerator is inflated relative to what humans report, while the human denominator is deflated by under-reporting.

IIHS recommended that NHTSA reform the federal crash reporting system, noting that manual review of incident narratives — the current method — is unsustainable as AV deployments scale.

The Most Important Sentence in the IIHS Report

It is not the 68% figure. The most important sentence is:

"Waymo is the only company that voluntarily provides driverless vehicle miles traveled information. Other companies do not share this data, making safety comparisons impossible for competitors."

IIHS could only calculate a crash rate for Waymo because only Waymo publishes driverless VMT. Cruise, Zoox, and other companies also file NHTSA crash reports — but without VMT as the denominator, no independent crash rate can be calculated.

In other words: the entire safety picture for autonomous vehicles in the United States currently depends on voluntary data from a single company. IIHS called on NHTSA to make VMT disclosure mandatory for all AV operators.


3. Baidu Apollo Go: Different Units, Different Questions

Q1 2026 Numbers from the 6-K Filing

Baidu published its Q1 2026 results via a 6-K filing with the SEC and the accompanying investor-relations release. Apollo Go reported:

  • 3.2 million fully driverless rides in Q1 2026
  • Weekly peak above 350,000 rides in March 2026
  • Over 22 million cumulative public rides as of April 2026
  • Operations in 27 cities, with expansion to Dubai and testing in London
  • Over 330 million autonomous km cumulative, including over 220 million fully driverless km

The Unit Trap — Read This Before Comparing

This is the easiest mistake to make when comparing Waymo and Apollo Go.

Waymo reports in MILES. The 220.6 million figure is miles.

Apollo Go reports in KILOMETERS. The 220 million figure is kilometers.

Unit conversion:

220 million km ÷ 1.60934 = ~136.7 million miles

Apollo Go's 220 million km is equivalent to roughly 137 million miles — about 38% less than Waymo's 220.6 million miles.

If you read the headlines quickly and see "220 million — 220 million," you will assume the two fleets have equivalent operational histories. They do not. This is a classic unit trap in industry reporting: two companies in two countries, reporting in two different measurement systems, and virtually no outlet mentions it in the headline.

That said, Apollo Go has a meaningful number Waymo does not publish in comparable form: 22 million cumulative public rides. Waymo does not report an equivalent aggregate trip count. Both companies selectively publish metrics that tell their story well — this is not necessarily concealment, but it makes direct comparison harder than it looks.

Apollo Go Does Not Publish Crash Rates

Unlike Waymo, Baidu does not publish crash rates or comparative safety analysis. The 6-K mentions an "outstanding safety record" without quantification. This may reflect different legal requirements, PR strategy, or safety reporting frameworks in the Chinese market. The result: no independent safety comparison can be made for Apollo Go against any human benchmark.


4. Zoox: Absence Is Data

Amazon's Zoox received federal approval in July 2026 to commercially deploy steering-wheel-free robotaxis in the United States — NHTSA authorized up to 2,500 vehicles per year. This is a significant regulatory milestone.

But on safety data: Zoox does not publish driverless VMT. It does not publish crash rates.

The company does file NHTSA incident reports as required — including a voluntary software recall in 2026 after vehicles drove into heavy smoke. But without VMT as a denominator, no one can calculate a rate or make an independent comparison.

This is not a gap in the research — this is the finding. When an industry lacks mandatory disclosure standards, the absence of data is itself informative: it tells you the limits of what can be concluded about a given operator.

Fortune (July 2026) explicitly flagged Zoox's "safety data gap" in the same week the federal approval was announced. That question remains unanswered.


5. The Hunter College/Sam Schwartz Incident: A Lesson in Reading Contrary Research

What Happened

In July 2026, the Sam Schwartz Transportation Research Program at Hunter College (City University of New York) and Open Plans published a report on robotaxi impacts in NYC. The safety analysis section claimed Waymo's rate of crashes causing serious injuries or deaths was 1.5 times higher per million miles than New York City's professional for-hire vehicle fleet (Uber and Lyft drivers).

This headline spread immediately — it directly contradicted Waymo's 94% improvement figure.

What Happened Next

On August 20, 2026, the authors released a revised report and retracted the safety analysis section. The stated reason: "data discrepancies between City agency data for New York City for-hire vehicle crashes with serious injuries or deaths, as well as City agency data definitions related to such crashes."

In plain terms: NYC's administrative definition of "crash causing serious injury" differs from the federal definition, making the comparison invalid. The traffic congestion analysis section was retained, as it does not depend on crash definitions.

Here is the part worth pausing on: the revision did not merely delete the old number, it reversed the conclusion. Hunter College's release now states that per mile traveled, in the cities where Waymo operates, Waymo vehicles are involved in 62% fewer injury crashes and 87% fewer serious-injury-or-fatal crashes than New York City for-hire vehicles. From "1.5 times higher" to "87% fewer" — same authors, same question, different city dataset.

Yet the report keeps its sharpest argument, and that argument does not favour Waymo: Waymo benchmarks itself against the "average human driver," while robotaxis would actually replace professional drivers. The report notes that NYC for-hire drivers are already in 20% fewer injury crashes and 25% fewer serious-injury-or-fatal crashes than the average NYC driver. Pick the wrong comparison group and every number improves for free.

A note on sourcing: the correction notice sits on Hunter College's own page, verbatim: "An earlier version of this story highlighted a portion of the report dealing with the safety of robotaxis. The study's authors released a revised report on August 20, 2026, after identifying issues with a city dataset used in the analysis." The old URL (ending -confer-few-safety-benefits) now redirects to a shorter slug — the redirect is itself a trace of the correction. The page blocks automated access (403 to bots) but opens normally in a browser; if you verify it with a script and see 403, that is an anti-bot wall, not a dead page.

Three Methodological Lessons

1. The denominator definition determines everything. "Serious injury crash" means something different under NYC administrative data, under federal NHTSA definitions, and under Waymo's reporting framework. Before comparing numbers, ask: which scale is "serious injury" measured on here?

2. City administrative data was not standardized for federal comparison purposes. Municipal datasets are collected for local purposes — licensing, insurance, urban planning — and are not necessarily consistent with NHTSA standards or academic research protocols.

3. Scientific self-correction is a feature, not a failure. The authors retracted the flawed section rather than defending a wrong result. This is correct scientific behavior — but retractions rarely receive coverage proportionate to the original headline. The result: many people remember "Waymo is 1.5x more dangerous" but never saw the retraction.


6. Summary: Comparing Disclosure Practices

Operator Driverless VMT/km Crash rates Unit Comparison method
Waymo 220.6 million miles Yes (detailed) Miles Police reports + NHTSA under-reporting correction
Apollo Go 220 million km (≈137M miles) No Kilometers N/A
Zoox Not disclosed No — N/A

This table is small, but it illustrates the point: when an industry lacks uniform disclosure standards, the overall safety picture depends on who chooses to publish what.


7. Closing: Methodology Is the Product

Looking back across all six posts in this series, one theme runs through each of them in different forms.

Post 1 asked: what does the end-to-end driving landscape look like from UniAD to VLA?

Post 2 asked: can you run a driving VLA from scratch?

Post 3 asked: which benchmarks are trustworthy for measuring these systems?

Post 4 asked: what does verifiable reasoning look like?

Post 5 asked: what are world models doing beyond generating pretty video?

Post 6 asks: when the vehicle goes on a real road, how do we read the results?

The common answer: methodology is the product. No number speaks for itself. 220.6 million miles only means something when you know it is measured in miles, restricted to surface streets, spatially weighted, and under-reporting-corrected. A 68% crash reduction only means something when you know it is an institute report rather than a peer-reviewed paper, the denominator comes from one company, and 78% of reported incidents did not meet a police-reportable threshold. Apollo Go's 220 million is only comparable when you know it is kilometers, not miles.

Good engineers ask about the denominator before trusting the numerator.

Lab benchmarks measure capability. Real-world deployment data measures behavior. Both are necessary — and both require transparent methodology to be worth reading.


Related Posts

  • Reasoning That Can Be Verified: Alpamayo-R1 and Chain-of-Causation — Part 4/6
  • Benchmarks as Reasoning: NAVSIM v2, Bench2Drive, WOD-E2E — Part 3/6
  • World Models Stop Making Pretty Videos, Start Measuring Policy — Part 5/6
NT

Nguyễn Anh Tuấn

Robotics & AI Engineer. Building VnRobo — sharing knowledge about robot learning, VLA models, and automation.

Khám phá VnRobo

Fleet MonitoringROS 2 IntegrationAMR Solutions
adas-e2e-2026 — Phần 6/6
← World Models Stopped Being About Pretty Video

Related Posts

NEWResearch
World model thôi làm video đẹp, chuyển sang đo lường policy
adasautonomous-drivingself-drivingPart 5
adas

World model thôi làm video đẹp, chuyển sang đo lường policy

World model 2026 không còn đo bằng FVD: tiêu chí mới là môi trường sinh có đo đúng chất lượng policy hay không. GAIA-4, Orbis 2, WorldLens.

8/19/202613 min read
NT
Research
Suy luận kiểm tra được: Alpamayo-R1 và Chain-of-Causation
adasautonomous-drivingself-drivingPart 4
adas

Suy luận kiểm tra được: Alpamayo-R1 và Chain-of-Causation

NVIDIA Alpamayo-R1 chứng minh xe tự lái giải thích được lý do phanh bằng chuỗi nhân-quả kiểm chứng được, không phải văn bản trang trí.

8/15/202610 min read
NT
Deep Dive
Benchmark chính là lập luận: NAVSIM v2, Bench2Drive, WOD-E2E
adasautonomous-drivingself-drivingPart 3
adas

Benchmark chính là lập luận: NAVSIM v2, Bench2Drive, WOD-E2E

Ba loại điểm benchmark KHÔNG thay thế được cho nhau. Hiểu điều này trước khi đọc bất kỳ paper autonomous driving nào năm 2026.

8/11/202612 min read
NT
VnRobo logo

AI infrastructure for next-generation industrial robots.

Product

  • Features
  • Pricing
  • Knowledge Base
  • Services

Company

  • About Us
  • Blog
  • Contact

Legal

  • Privacy Policy
  • Terms of Service

© 2026 VnRobo. All rights reserved.

Made with♥in Vietnam