Multimodal AI · Google Maps · 2023
Point the camera down a street and ask a question. Lens in Maps was the first Google Maps feature built on Gemini's models, answering what you see with knowledge of where you stand.
Defining the problem
Live View had proven the camera could locate you better than GPS, but its value was gated on navigating somewhere. The next step was letting people ask about the world directly: point the camera at an unfamiliar street and learn what is worth knowing about it. In 2023, large language and vision models made that plausible for the first time, and nobody had established methods for designing with them in a production product.
My team was the first in Google Maps to adopt LLMs and vision-language models. Lens in Maps combined Gemini's image understanding with place context from the Visual Positioning System, so a photo and a question produced an answer grounded in the actual place in front of you rather than a generic caption.
Process
Designing for a model means designing its behavior, not just its screens. We worked out context and prompt engineering as a production design practice, defined evals and golden datasets with UX writers, and prototyped the experience with people before building it: in the Amazing Race NYC exercise, a Wizard-of-Oz study I designed, researchers role-played a voice agent for participants walking around New York, surfacing what people actually ask outside the app and what context Maps already held in those moments.


Image understanding alone produces a caption. Fused with the Visual Positioning System, it produces an answer about this door, this restaurant, this street.
Golden datasets and evals, written with UX writers, are how you design a model's behavior deliberately instead of discovering it in production.
Role-playing the voice agent on real streets told us more about real queries than a built demo could, before a line of engineering was spent.
With screen reader support, Lens in Maps describes places aloud through VoiceOver — a real-world superpower for blind and low-vision users.
What we learned
Lens in Maps shipped in 2023 as the first Google Maps feature to use Gemini's language and vision models, and the screen reader launch I advocated for made the camera useful to people who cannot see what it sees. The methods became the team's working model for AI products: context engineering, evals, and golden datasets as standard design practice.
That groundwork — semantic scene understanding plus natural-language interaction — fed directly into Gemini in Navigation, the first voice assistant in Google Maps, which I led the design of as a Staff designer through its 2025 launch.