Reconstructing dynamic preferred visual experiences during naturalistic navigation

Tianjiao Zhang
Equal contribution
, Cheol Jun Cho
Equal contribution
, Jack L. Gallant | 2026 | bioRxiv
Summary
The brain is selective for complex, high-dimensional combinations of features. These tuning properties are difficult to understand intuitively. We use a VAE to decode these tuning properties back into image space to create preferred visual experiences that tell us what a region is most active for, which serves as an intuitive proxy for its tuning.
Abstract
Complex, high-dimensional brain representations are often difficult to understand intuitively. Recent advances in generative neural networks have made it possible to predict the optimal stimulus that maximizes the activity in a specific brain region. Because these stimuli are in image space, they provide an intuitive understanding of the functional properties of that region. Previous work on optimized stimuli has focused on passive perceptual tasks. Yet real-world perception operates in a closed loop with cognition and action, and so passive stimuli are unlikely to be optimal for many brain regions. In a previous study, we used fMRI to record BOLD activity from participants actively navigating through a virtual world, and identified a network of 11 cortical regions that support active navigation and are tuned for complex feature combinations. Here, we used a variational autoencoder to generate video clips predicted to correspond to maximal activity in each region. These preferred visual experiences capture dynamic scenes that incorporate both world events and participant actions. These preferred visual experiences provide intuitive descriptions of the functional properties of each region, and suggest that these regions are more engaged during dynamic than static scenes. These patterns also generate novel, data-driven hypotheses for future study.