Multimodal AI · Google Maps · 2023

Lens in Maps

Point the camera down a street and ask a question. Lens in Maps was the first Google Maps feature built on Gemini's models, answering what you see with knowledge of where you stand.

Role

UX Design Manager

Team

Cross-functional UX team of ~7 — designers, UX engineers, motion

Duration

2023 – 2025

Defining the problem

Search assumes you can name what you want. Standing on a street, you often can't.

Live View had proven the camera could locate you better than GPS, but its value was gated on navigating somewhere. The next step was letting people ask about the world directly: point the camera at an unfamiliar street and learn what is worth knowing about it. In 2023, large language and vision models made that plausible for the first time, and nobody had established methods for designing with them in a production product.

My team was the first in Google Maps to adopt LLMs and vision-language models. Lens in Maps combined Gemini's image understanding with place context from the Visual Positioning System, so a photo and a question produced an answer grounded in the actual place in front of you rather than a generic caption.

Process

Designing for a model means designing its behavior, not just its screens. We worked out context and prompt engineering as a production design practice, defined evals and golden datasets with UX writers, and prototyped the experience with people before building it: in the Amazing Race NYC exercise, a Wizard-of-Oz study I designed, researchers role-played a voice agent for participants walking around New York, surfacing what people actually ask outside the app and what context Maps already held in those moments.

Field research in New YorkMultimodal query explorations

Ground answers in place

Image understanding alone produces a caption. Fused with the Visual Positioning System, it produces an answer about this door, this restaurant, this street.

Evals are design material

Golden datasets and evals, written with UX writers, are how you design a model's behavior deliberately instead of discovering it in production.

Prototype the agent with people

Role-playing the voice agent on real streets told us more about real queries than a built demo could, before a line of engineering was spent.

The camera is an accessibility device

With screen reader support, Lens in Maps describes places aloud through VoiceOver — a real-world superpower for blind and low-vision users.

What we learned

Maps' first Gemini feature, and the working model the next ones ran on.

Lens in Maps shipped in 2023 as the first Google Maps feature to use Gemini's language and vision models, and the screen reader launch I advocated for made the camera useful to people who cannot see what it sees. The methods became the team's working model for AI products: context engineering, evals, and golden datasets as standard design practice.

That groundwork — semantic scene understanding plus natural-language interaction — fed directly into Gemini in Navigation, the first voice assistant in Google Maps, which I led the design of as a Staff designer through its 2025 launch.

Next case study

Live View

Case Study