zurich.tech

[about]

Felix Wimbauer (Technical University of Munich/Google): Learning to Understand the 3D World from Large-Scale Video Datasets

Felix will cover three of his works (MonoRec, BehindTheScenes, S4C) for understanding 3D scenes from video data, highlighting approaches that shift from traditional supervised learning to more data-efficient, self-supervised techniques. Each method addresses the common challenges of estimating depth, handling dynamic objects, and synthesizing new views, all while working with minimal labeled data. These techniques share a focus on leveraging large video datasets to improve depth prediction, semantic segmentation, and camera motion estimation without relying on specialized sensors. Building on that, he will present his recent CVPR 2025 paper AnyCam, which builds on these principles by estimating camera motion and scene geometry from casual videos, aiming for more reliable performance in diverse environments.

Joan Puigcerver (Google DeepMind): Scaling Computer Vision: Transformers and Sparse Mixture-of-Experts Models

The field of Computer Vision shifted from ad-hoc models trained on a few thousands of examples, to more general models trained on large and diverse datasets, able to solve many tasks. We will review how Transformers have been a fundamental part of this shift, both in Computer Vision and Natural Language Processing, and more recently Sparse Mixture-of-Experts (MoE) enabled to further increase the model capacity without increasing the training or evaluation time complexity. We will explore different types of MoE models, and discuss some particularities of vision models compared to language models.

[speakers]

Felix Wimbauer

Technical University of Munich and Google

Joan Puigcerver

Google DeepMind

[details]

time
07 aug 2025 18:00
location
ETH AI Center
address
Andreasstrasse 5, OAT, 14th floor, 8050, Zürich
format
talk
status
finished
tags
#past#3d#computer-vision#zurichcv
access
OAT building.

[photos]

5