zurich.tech

[about]

Ivan Vulić (University of Cambridge/Google DeepMind): Guiding Vision-Language Models to Climb the Mountain of Spatial Reasoning

Large Vision-Language Models (VLMs) have demonstrated impressive performance in general vision-language tasks. However, even the most recent and most powerful VLMs still struggle even with simple spatial understanding and reasoning capabilities. In this talk, I will first provide a brief overview of our recent work on creating new benchmarks and improving evaluation of VLMs for a range of spatial reasoning tasks. I will then outline our novel methodology related to enhancing spatial reasoning capabilities of VLMs, with a focus on spatial navigation tasks, such as multi-modal visualization-of-thought and purely visual planning.

Tiago Pimentel (ETH Zurich): Duplicating Vocabularies to Analyse Generalisation in Language Models

In this talk, we will explore how duplicating a language model’s vocabulary can create controlled experiments, which we can leverage to address two research questions. First, we use vocabulary duplication to investigate lexical generalisation in LMs, looking at the effect of near duplicate subwords (e.g., vocabulary items such as Now and now) in an LM's performance. Second, we will analyse cross-linguistic generalisation, investigating how cross-linguistic data (im)balance affects performance LMs' performance.

[speakers]

Ivan Vulić

University of Cambridge and Google DeepMind

Tiago Pimentel

ETH Zurich

[details]

time
21 may 2025 18:00
location
ETH AI Center
address
Andreasstrasse 5, OAT, 14th floor, 8050, Zürich
format
talk
status
finished
tags
#past#nlp#vision-language-models#zurichnlp
access
OAT building.

[photos]

5