Waymo Explains Their Multimodal Approach to Autonomy [VIDEO]

Nehal Malik

The autonomous driving industry remains split on how self-driving vehicles should perceive the world. While Tesla’s vision-only architecture gains ground with consumer Full Self-Driving and the commercial Robotaxi platform, Alphabet’s Waymo continues to advocate for a multimodal sensor suite combining cameras, radar, and LiDAR.

Speaking at Y Combinator's Startup School 2026 event, Waymo Co-CEO Dmitri Dolgov shared key technical lessons learned from scaling Waymo’s driverless robotaxis from a demo to a commercial product. During his talk, Dolgov detailed why Waymo relies on fusing data across multiple physical sensing modalities to solve physical AI challenges in real-world traffic.

Fusing Cameras, Radar, and LiDAR for Redundancy

Not a Tesla App

Waymo’s technical architecture centers around a multimodal sensor setup. Rather than relying solely on optical cameras, Waymo vehicles combine high-resolution cameras, active LiDAR, and radar sensors into a unified perception model.

Dolgov explained that while cameras provide rich color and high resolution, passive optical sensors suffer in heavy glare, direct sunlight, complete darkness, or harsh weather conditions. Active LiDAR provides direct 3D structural measurements regardless of ambient lighting, while radar penetrates dense fog, rain, and snow using Doppler measurements to measure object velocity.

According to Waymo, single-modality systems create a safety curve that flattens out too early when attempting to achieve superhuman driving performance. By equipping vehicles with multiple distinct sensor types, overlapping coverage protects against edge-case failures, such as road debris partially obscuring a camera lens.

Waymo vs. Tesla: Two Divergent Philosophies

Waymo’s multimodal methodology sits in stark contrast to Tesla’s vision-only approach. While Waymo claims a camera-only approach provides an easier early ramp that eventually hits a performance ceiling, Tesla remains committed to a single sensing modality, adamant that it can solve autonomy for edge cases by simply feeding its models enough real-world data.

Tesla notoriously walked away from radar and LiDAR to go all-in on vision, arguing that because human drivers navigate using two eyes and a brain, self-driving vehicles should operate identically using cameras and neural networks. Instead of active sensors, Tesla models the physical world with vision alone, using cameras to reconstruct 3D space and processing billions of video frames to solve long-tail edge cases.

Not a Tesla App

This architectural disagreement highlights two very different business models. Tesla’s vision-only system keeps hardware costs low, allowing driver-assist features to ship natively on millions of customer vehicles worldwide. On the flip side, Waymo’s multimodal suite adds expensive hardware, but it clears current regulatory frameworks more easily thanks to its fundamentals of redundancy and proven safety metrics.

While humans can drive safely using visual perception alone, adding active radar and 3D LiDAR creates extra safety margins that could allow autonomous systems to react faster than humanly possible, and potentially even drive faster than humans safely can. As both platforms scale their unsupervised driverless operations, real-world intervention metrics will determine which technical path reaches widespread fleet deployment first. You can watch Dolgov’s full talk in the video below: