Hi, I am a third-year PhD Candidate at the University of Illinois Urbana‑Champaign, advised by Svetlana Lazebnik. I also closely collaborate with Unnat Jain and Heng Ji. I am working on post-training methods to build models for robotics that generalize. Some keywords to easily convey my research: MoT, VLAs, WAMs, VLMs, and multi‑modal representation learning.
Previously, I completed my B.Tech from Indian Institute of Technology, Gandhinagar, where I was awarded the Institute Gold medal for graduating at the top of my discipline.
I had great summers at Amazon (Ring AI, 2026) and Microsoft Research (2025), where I worked on long-video understanding with MLLMs and on vision language navigation with MLLM agents, respectively. Check out the projects: Repairable Uncertainty (Amazon) and AgentNav (Microsoft Research).
Amazon
Summer 2026 |
Microsoft Research
Summer & Fall 2025 |
University of British Columbia
Summer 2023 |
USC (AIISC)
2022‑2024 |
ISRO
Spring 2023 |
TCS Research
Winter & Spring 2023 |
EFICENS
Summer 2022 |
Selected PublicationsAll Publications » |
Generalizable VLA Finetuning via Representation Anchoring and Language-Action AlignmentDwip Dalal, Shivansh Patel, Chahit Jain, Jeonghwan Kim, Utkarsh Mishra, Alex Baratian, Hyeonjeong Ha, Heng Ji, Svetlana Lazebnik*, Unnat Jain* Under Review Paper | Code | Project Page |
|
|
Coming soon!
|
Repairable Uncertainty: Selective Temporal Reacquisition for Long-Video Understanding with MLLMsDwip Dalal, Keval Doshi, Amar Kumar, Wei Wang, Sowndarya Sundar, Noranart Vesdapunt, Jim Thomas, Kah Kuen Fu, Svetlana Lazebnik Under Review |
Constructive Distortion: Improving MLLMs with Attention-Guided Image WarpingDwip Dalal, Gautam Vashishtha, Utkarsh Mishra, Jeonghwan Kim, Madhav Kanda, Hyeonjeong Ha, Svetlana Lazebnik, Heng Ji, Unnat Jain ICLR 2026 Paper | Code | Project Page |
|
Can MLLMs Find Their Way in a City? Exploring Emergent Navigation from Web-Scale KnowledgeDwip Dalal*, Utkarsh Mishra*, Narendra Ahuja, Nebojsa Jojic EACL 2026 (Oral) Paper | arXiv | Code | Project Page |
|
Compositional Reasoning via Joint Image and Language DecompositionDwip Dalal*, Madhav Kanda*, Zhenhailong Wang, Heng Ji, Unnat Jain EACL 2026 Paper | Project Page |
|
|
Learning Robust Deep Visual Representations from EEG Brain RecordingsPrajwal Singh, Dwip Dalal, Gautam Vashishtha, Shanmuganathan Raman, Krishna Prasad Miyapuram WACV 2024 Paper | Code | Project Page | WACV Daily | Best of WACV 2024 |
|
|
Flow Symmetrization for Parameterized Constrained DiffeomorphismsDwip Dalal*, Aalok Gangopadhyay*, Progyan Das*, Shanmuganathan Raman Arxiv, 2024 Paper | Code | Project Page |
|
Single Image LDR to HDR Conversion Using Conditional DiffusionDwip Dalal, Gautam Vashishtha, Prajwal Singh, Shanmuganathan Raman International Conference on Image Processing (ICIP'23) (Oral) Paper | Project Page |
|
ODESolvers are also Wayfinders: Neural ODEs for Multi-Agent PathplanningDwip Dalal*, Progyan Das*, Anirban Dasgupta NeurIPS 2023 Workshop - Deep Learning and Differential Equations III Paper | Code |
|
FACTIFY-5WQA: 5W Aspect-based Fact Verification through Question AnsweringAnku Rani, SM Tonmoy, Dwip Dalal, Shreya Gautam, Megha C., Aman Chadha, Amit Sheth, Amitava Das ACL 2023 Paper | Code & Dataset |