Description
Title: Learning Universal Multimodal Representations via a Global Workspace with Minimal Paired Data
Abstract: Learning multimodal representation using current deep learning methods requires a massive amount of multimodal paired data, relying on pure supervised learning. These paired data are difficult to obtain, limiting the possible amount available or requiring a lot of works to create them. On the opposite humans and animals in general are able to learn useful multimodal representation using only few paired samples. Recent works have begun tackling this challenge by reducing reliance on full supervision and dense pairing, yet limitations remain.
This work builds on these efforts, adapting the Global Workspace Theory of human cognition to multimodal representation learning under limited supervision. Central to this theory, the broadcasting mechanism, implemented with a set of modality-specific encoders and decoders, enables training through unsupervised and self-supervised losses using cycle-consistency.
We evaluated our approach on retrieval, generation, and latent space arithmetic tasks across varying amounts of paired samples. Our method outperformed current state-of-the-art approaches in very low paired data regimes and scaled more effectively as the amount of paired data increased.
Bio: Léopold Maytié received the Ph.D. degree in computer science from the University of Toulouse, France, in 2025, with a thesis on multimodal representation learning for robotics. He is currently a postdoctoral researcher at the University of Innsbruck, Austria, where he works on latent representations for reinforcement learning, with a focus on discovering diverse strategies to solve a same problem. His research interests include representation learning and world models for planning and problem solving.