My research spans self-supervised learning for MRI, cross-subject transfer, and multimodal brain decoding. I'm interested in how learned representations capture spatiotemporal structure and biological variability, and how dataset priors affect out-of-distribution generalization.
Selected publications
ICML 2026
Scaling Vision Transformers for Functional MRI with Flat Maps
Connor Lane, Mihir Tripathy, Leema Krishna Murali, Ratna Sagari Grandhi, Shamus Sim Zi Yang, Sam Gijsen, Debojyoti Das, Manish Ram, Utkarsh Kumar Singh, Cesar Kadir Torrico Villanueva, Yuxiang Wei, Will Beddow, Gianfranco Cortés, Suin Cho, Daniel Z. Kaplan, Benjamin Warner, Tanishq M. Abraham, Paul S. Scotti
We adapted Vision Transformers to fMRI using cortical flat maps and trained CortexMAE on 2,100 hours of open brain activity data. We also released Brainmarks to evaluate trait prediction and cognitive state decoding.
Predicting Brain Responses To Natural Movies With Multimodal LLMs
Cesar Kadir Torrico Villanueva, Jiaxin Cindy Tu, Mihir Tripathy, Connor Lane, Rishab Iyer, Paul S. Scotti
Our multimodal pipeline combines video, audio, and language features to predict brain responses to movies. It placed fourth in Algonauts 2025, with a mean Pearson correlation of 0.2085 on withheld, out-of-distribution movies.
Transformer-Based Foundation Model for Functional Neuroimaging
Teddy Akiki, Mihir Tripathy, Xue Zhang, Adam Pines, Josue Ortega Caro, Syed Rizvi, Christopher Averill, David van Dijk, Chadi Abdallah, Leanne Williams
A transformer-based foundation model approach for functional neuroimaging, published in Biological Psychiatry, Volume 97, Issue 9, pages S160–S161.
MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of Data
Paul S. Scotti, Mihir Tripathy, Cesar Kadir Torrico Villanueva, Reese Kneeland, Tong Chen, Ashutosh Narang, Charan Santhirasegaran, Jonathan Xu, Thomas Naselaris, Kenneth A. Norman, Tanishq Mathew Abraham
We trained a shared model across seven participants, then adapted it to a new participant using as little as one hour of fMRI data for image retrieval and reconstruction.