ZurichCV #10
[about]
Felix Wimbauer (Technical University of Munich/Google): Learning to Understand the 3D World from Large-Scale Video Datasets
Felix will cover three of his works (MonoRec, BehindTheScenes, S4C) for understanding 3D scenes from video data, highlighting approaches that shift from traditional supervised learning to more data-efficient, self-supervised techniques. Each method addresses the common challenges of estimating depth, handling dynamic objects, and synthesizing new views, all while working with minimal labeled data. These techniques share a focus on leveraging large video datasets to improve depth prediction, semantic segmentation, and camera motion estimation without relying on specialized sensors. Building on that, he will present his recent CVPR 2025 paper AnyCam, which builds on these principles by estimating camera motion and scene geometry from casual videos, aiming for more reliable performance in diverse environments.
Joan Puigcerver (Google DeepMind): Scaling Computer Vision: Transformers and Sparse Mixture-of-Experts Models
The field of Computer Vision shifted from ad-hoc models trained on a few thousands of examples, to more general models trained on large and diverse datasets, able to solve many tasks. We will review how Transformers have been a fundamental part of this shift, both in Computer Vision and Natural Language Processing, and more recently Sparse Mixture-of-Experts (MoE) enabled to further increase the model capacity without increasing the training or evaluation time complexity. We will explore different types of MoE models, and discuss some particularities of vision models compared to language models.
[speakers]
Felix Wimbauer
Technical University of Munich and Google
Joan Puigcerver
Google DeepMind
[details]
- time
- 07 aug 2025 18:00
- location
- ETH AI Center
- address
- Andreasstrasse 5, OAT, 14th floor, 8050, Zürich
- format
- talk
- status
- finished
- tags
- access
- OAT building.