|
Shiyi Lan 兰石懿
I am a Staff Research Scientist at Waymo. My goal is to build general-purpose multimodal agents that perceive, reason, plan, and act across digital and physical environments. I focus on language-grounded perception, long-horizon decision-making, and robust interaction under uncertainty.
My research spans VLM post-training, multimodal state and world modeling, Reinforcement Learning, and closed-loop perception-to-action learning. I have applied these capabilities to embodied decision-making and end-to-end autonomous driving. I received my Ph.D. in Computer Science from the University of Maryland, College Park in 2022, advised by Prof. Larry S. Davis, and my B.S. from Fudan University in 2018.
Email: voidrank [at] gmail [dot] com
CV
/
Google Scholar
/
GitHub
|
|
|
Experience
|
|
Waymo, Staff Research Scientist 2026–Present
Building Gemini-based multimodal agents that translate complex visual observations into language-grounded state representations, enabling long-horizon reasoning, decision-making, and robust action in long-tail scenarios. Developing post-training methods for diffusion-based Gemini VLMs and improving RL infrastructure for diffusion models, with autonomous driving as the primary embodied application.
NVIDIA Research, Senior Research Scientist 2023–2026
Core member of Nemotron-Diffusion-VL, responsible for data recipes, evaluation, VLM implementation, and training infrastructure. Led research on multimodal perception-to-action agents, including language-grounded scene understanding, trajectory generation, teacher-student distillation, closed-loop evaluation, and reinforcement learning, applied to end-to-end autonomous driving through the HydraMDP family and OmniDrive.
NVIDIA Research, Research Scientist 2022–2023
Core team member of NV-CLIP; led pre-training and resolved loss-divergence issues. Developed MMAL, the CVPR 2023 mask auto-labeling framework.
NVIDIA Research, Research Intern 2020
Researched weakly supervised instance segmentation on MS COCO.
ByteDance AI Lab, Research Intern 2018
Contributed to a recommendation-system warmup project that reduced daily warmup cost by five million CNY.
|
|
Signature Projects
|
- Nemotron-Diffusion-VL — Core member of NVIDIA's first diffusion VLM.
- End-to-End Driving at Scale — 1st place at CVPR 2025 and CVPR 2024 (Team Lead).
- 3D Occupancy Challenge — 1st place at CVPR 2023.
- Robust Vision Challenge — 1st place in the 2022 semantic segmentation track.
- MS COCO 2017 — 1st place in object detection and 2nd place in instance segmentation.
|
|
|
HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving
Wenhao Yao, Xinglong Sun, Zhenxin Li, Shiyi Lan, Zi Wang, Jose M. Alvarez, Zuxuan Wu
Under submission, 2026
paper
|
|
|
DriveSuprim: Towards Precise Trajectory Selection for End-to-End Planning
W Yao, Z Li, Shiyi Lan, Z Wang, X Sun, JM Alvarez, Z Wu
AAAI, 2026
paper /
arXiv
|
|
|
Hydra-MDP: End-to-End Multimodal Planning with Multi-Target Hydra-Distillation
Zhenxin Li, Kailin Li, Shihao Wang, Shiyi Lan, Zhiding Yu, Yishen Ji, Zhiqi Li, Ziyue Zhu, Jan Kautz, Zuxuan Wu, Yu-Gang Jiang, Jose M. Alvarez
arXiv, 2024
paper /
code
|
|
|
Hydra-MDP++: Advancing End-to-End Driving via Expert-Guided Hydra-Distillation
K Li, Z Li, Shiyi Lan, Y Xie, Z Zhang, J Liu, Z Wu, Z Yu, JM Alvarez
arXiv, 2025
paper /
code
|
|
|
Hydra-NeXt: Robust Closed-Loop Driving with Open-Loop Training
Z Li, S Wang, Shiyi Lan, Z Yu, Z Wu, JM Alvarez
ICCV, 2025
paper /
code
|
|
|
Centaur: Robust End-to-End Autonomous Driving with Test-Time Training
C Sima, K Chitta, Z Yu, Shiyi Lan, P Luo, A Geiger, H Li, JM Alvarez
arXiv, 2025
paper
|
|
|
Enhancing Autonomous Driving Safety with Collision Scenario Integration
Z Wang, Shiyi Lan, X Sun, N Chang, Z Li, Z Yu, JM Alvarez
arXiv, 2025
paper
|
|
|
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Z Li, G Chen, S Liu, S Wang, V VS, Y Ji, Shiyi Lan, H Zhang, Y Zhao
arXiv, 2025
paper
|
|
|
Cosmos World Foundation Model Platform for Physical AI
N Agarwal, A Ali, M Bala, Y Balaji, E Barker, T Cai, P Chattopadhyay, Shiyi Lan, et al.
arXiv, 2025
paper
|
|
|
StreamChat: Chatting with Streaming Video
J Liu, Z Yu, Shiyi Lan, S Wang, R Fang, J Kautz, H Li, JM Alvarez
arXiv, 2024
paper
|
|
|
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
S Wang, Z Yu, X Jiang, Shiyi Lan, M Shi, N Chang, J Kautz, Y Li, JM Alvarez
CVPR, 2025
paper
|
|
|
MDP: Multidimensional Vision Model Pruning with Latency Constraint
X Sun, B Lakshmanan, M Shen, Shiyi Lan, J Chen, JM Alvarez
CVPR, 2025
paper
|
|
|
A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties
J Xiao, Z Zhou, W Li, Shiyi Lan, J Mei, Z Yu, B Zhao, A Yuille, Y Zhou, C Xie
ECCV, 2024
paper /
code
|
|
|
FocalFormer3D: Focusing on Hard Instance for 3D Object Detection
Y Chen, Z Yu, Y Chen, Shiyi Lan, A Anandkumar, J Jia, JM Alvarez
ICCV, 2023
paper
|
|
|
FB-OCC: 3D Occupancy Prediction based on Forward-Backward View Transformation
Z Li, Z Yu, D Austin, M Fang, Shiyi Lan, J Kautz, JM Alvarez
CVPR Workshop on End-to-End Autonomous Driving, 2023
paper /
code
|
|
|
Vision Transformers Are Good Mask Auto-Labelers
Shiyi Lan, X Yang, Z Yu, Z Wu, JM Alvarez, A Anandkumar
CVPR, 2023
paper
|
|
|
1st Place Solution of The Robust Vision Challenge (RVC) 2022 Semantic Segmentation Track
J Xiao, Z Xu, Shiyi Lan, Z Yu, A Yuille, A Anandkumar
arXiv, 2022
paper
|
|
|
DiscoBox: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision
Shiyi Lan,
Zhiding Yu,
Christopher Choy,
Subhashree Radhakrishnan,
Guilin Liu,
Yuke Zhu,
Larry S. Davis,
Anima Anandkumar
ICCV, 2021
arXiv
/
code
The first-place weakly supervised instance segmentation model using box labels.
|
|
|
M3DETR: Multi-representation, Multi-scale, Mutual-relation 3D Object Detection with Transformers
Tianrui Guan*,
Jun Wang*,
Shiyi Lan†,
Rohan Chandra,
Zuxuan Wu,
Larry S. Davis,
Dinesh Manocha
† Correspondence Author. * means equal contribution.
WACV, 2022
arXiv
/
code
The multi-representation, multi-scale, mutual-relation 3D object detector with transformers.
|
|
|
InfoFocus: 3D Object Detection for Autonomous Driving with Dynamic Information Modeling
Jun Wang*,
Shiyi Lan*,
Mingfei Gao,
Yi Wu,
Larry S. Davis,
* means equal contribution.
ECCV, 2020
arXiv
3D Object Detection with the effective dynamic attention module.
|
|
|
AdaViT: Adaptive Vision Transformers for Efficient Image Recognition
Lingchen Meng, Hengduo Li, Bor-Chun Chen, Shiyi Lan, Zuxuan Wu, Yu-Gang Jiang, Ser-Nam Lim
CVPR, 2022
paper
|
|
|
SaccadeNet: A Fast and Accurate Object Detector
Shiyi Lan,
Zhou Ren,
Yi Wu,
Larry S. Davis,
Gang Hua
CVPR, 2020
arXiv
/
code
A real-time and high-performance object detector.
|
|
|
Modeling Local Geometric Structure of 3D Point Clouds Using Geo-CNN
Shiyi Lan,
Ruichi Yu,
Gang Yu,
Larry S. Davis
CVPR, 2019
arXiv
/
code
Modeling point cloud by leveraging geometric information of point clouds.
|
|
|
FastMask: Segment Multi-scale Object Candidates in One Shot
Hexiang Hu*,
Shiyi Lan*,
Yuning Jiang,
Zhimin Cao,
Fei Sha
* means equal contribution.
CVPR, 2017 (Spotlight)
arXiv
/
code
Generating instance segmentation candidates by leveraging multi-scale feature pyramids.
|
|
Additional Honors
|
- ICPC 2015 Shenyang Regional, Silver Medal (Rank 18/≈300)
- National Olympiad in Informatics of China 2013, Bronze (Rank 122/≈400)
|
|
Patents
|
- Temporal-based perception for autonomous systems and applications, US Patent 12,625,926 (2026) (Choi, Lopez, Lan, Asgarieh, Yu)
- Occupancy prediction using forward-backward view transformation, US Patent 12,515,706 (2026) (Li, Yu, Austin, Lan, Kautz, Lopez)
- Point-level supervision for video instance segmentation, US Patent App. 18/395,198 (Yu, Huang, Huang, Lan, Radhakrishnan, Alvarez)
- Class agnostic object mask generation, US Patent 12,614,284 (2026) (Lan, Yu, Radhakrishnan, Lopez, Anandkumar)
- Object detection, instance segmentation, and semantic correspondence from bounding box supervision, US Patent App. 17/177,068 (Yu, Lan, Choy, Radhakrishnan, Liu, Zhu, Anandkumar)
|
|
Academic Services
|
|
Conference Reviewer: AAAI'20, CVPR'21, ICCV'21, CVPR'22, ECCV'22; Journal Reviewer: IJCV'20
|
|