Vighnesh Subramaniam

I am a third year MIT EECS PhD student at CSAIL advised by Boris Katz in the MIT Infolab. I also work closely with Brian Cheung and previously worked with Andrei Barbu.

I was recently a student researcher at Google working on the Cloud AI Research team with Yale Song and Chun-Liang Li.

Before my PhD, I obtained my bachelors (2023) and MEng (2024) in Computer Science at MIT. During this time, I was a research assistant in the Infolab specifically working on problems in deep learning, multimodal processing, and computational neuroscience. During my MEng, I also had the opportunity to work on research with Shuang Li, Yilun Du, Igor Mordatch, and Antonio Torralba on problems related to generative modeling. I am extremely fortunate to be supported by the NSF Graduate Research Fellowship as well as the Robert J Shillman (1974) Fund Fellowship.

Email  /  Scholar  /  Github  /  Twitter

profile photo

Research

My research focuses on studying the science of deep learning, and more specifically, the science of neural network design. This has mostly involved studying and understanding the inductive biases of neural networks and how these biases relate to architecture and training dynamics. But also more recently, this has branched into new training methods for autoregressive transformers, metrics for comparing neural representations, and improvements of multimodal understanding in deep learning models. I have also led projects related to generative modeling and test-time scaling.

In the past, I conducted research on many topics, including multimodal models, finetuning with large language models, and deep learning applications to neuroscience/cognitive science. I highlight some publications below -- see my scholar for a more complete list.

Optimization Prior Diagram **NEW** PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents
Vighnesh Subramaniam, Boris Katz, Brian Cheung, Chun-Liang Li, Tomas Pfister, Yale Song
Under Review
Project Page / arXiv

We introduce PreviewDiff, a new test-time scaling method for diffusion models. Our method operates on diffusion latents, extracting a denoised preview of the image/video that is evaluated by an multimodal LLM judge, applying edits in a tree structure. Our tree branching operates over edits to the input prompt to the diffusion model, introduced by the MLLM judge. We evaluate our approach across different image and video diffusion models, showing strong test-time scaling trends! Spending more compute on branching with MLLM edits leads to increasingly strong results. Our analysis also explores components that matter for scaling and improvements across different settings.

Optimization Prior Diagram **NEW** Transferring Architecture as an Optimization Prior
Vighnesh Subramaniam, Boris Katz, Brian Cheung
Neural Information Processing Systems, 2026
arXiv (Coming soon!)

We introduce representational similarity as an optimization prior. This is an initialization-only procedure that aligns a target network to the representational geometry of a randomly initialized guide network before any downstream training occurs. Crucially, the guide network transfers no learned knowledge: it is frozen, never sees labels. We demonstrate that this cross-architecture transfer of fate works in two key domains: from deeper to shallower networks and from fine-grained to coarse-grained networks. Ultimately, this establishes that a model's architectural destiny can be decoupled from its physical structure and explicitly programmed at initialization.

Network of Theseus Diagram Network of Theseus (like the ship)
Vighnesh Subramaniam, Colin Conwell, Boris Katz, Andrei Barbu, Brian Cheung
Neural Information Processing Systems, 2026
International Conference on Learning Representations: Workshop on Scientific Methods for Understanding Deep Learning, 2026
arXiv

We introduce the Network of Theseus (NoT), a method to part-by-part convert between architectures. We show massive conversions where untrained architectures can be used to guide a conversion to a completely new architecture. This leads to some pretty amazing findings on RNNs and MLPs. Furthermore, we find some even more striking results for making larger architectures smaller using untrained teachers!

Training the Untrainable Diagram Training the Untrainable: Introducing Inductive Bias via Representational Alignment
Vighnesh Subramaniam, David Mayo, Colin Conwell, Tomaso Poggio, Boris Katz, Brian Cheung, Andrei Barbu
Neural Information Processing Systems, 2025
Neural Information Processing Systems: Workshop on Unifying Representations in Neural Models, 2024
Project Page / arXiv / Code

We investigate the relationship between inductive biases in neural networks and their representation space by designing methods to transfer aspects that make certain networks trainable to networks that are difficult to train. This is done by a per-training-step alignment of the untrainable network activations and trainable network activations. Our findings are really surprising and we make some of these networks really competitive like RNNs!

Multiagent Finetuning diagram Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains
Vighnesh Subramaniam*, Yilun Du*, Joshua B. Tenenbaum, Antonio Torralba, Shuang Li, Igor Mordatch
International Conference on Learning Representations, 2025
Project Page / arXiv / Code

We designed a new self-improvement finetuning method for LLMs that preserves diversity. Our method builds on multiagent prompt frameworks like multiagent debate but finetuning sets of generation and critic models that interact by proposing more accurate solutions and critiquing solutions more accurately. We see pretty considerable improvements across several math-related tasks!


Website design credits to John Barron.