I am a Director and Research Scientist at Google DeepMind, leading the research teams and technical strategy behind Gemini’s video understanding and multimodal foundation models.
My team drives the full spectrum of spatiotemporal reasoning: from web-scale exocentric analysis deployed across YouTube and Vertex AI, to low-latency egocentric reasoning powering Gemini Live. Our native multimodal architectures set state-of-the-art benchmarks across core capabilities (see the Gemini 3.0 blog and companion vision blog).
Most recently, we introduced agentic video understanding, advancing a core thesis in multimodal intelligence: coupling native spatiotemporal scaling with active perception, enabling models to dynamically reason and control what they observe across space and time.
Prior to DeepMind, I worked on Generative AI and large-scale vision systems. I completed my PhD at ETH Zurich, focused on efficient data selection and discrete optimization algorithms.