Yuxin DU

Embodied AI, Vision-Language-Action (VLA) models, and egocentric data.

Yuxin DU

I am a PhD student at the School of Artificial Intelligence, Shanghai Jiao Tong University.

My research focuses on Embodied AI, Vision-Language-Action (VLA) models, and egocentric data for embodied systems that perceive, understand, reason, and execute in physical environments.

Portrait of Yuxin DU

About

I am currently a PhD student at the School of Artificial Intelligence (SAI), Shanghai Jiao Tong University. My research interests include Embodied AI, Vision-Language-Action (VLA) models, and egocentric data. Before joining SJTU, I received my MPhil degree from the Artificial Intelligence Thrust at The Hong Kong University of Science and Technology (Guangzhou). I received my bachelor's degree in Computer Science and Technology from Beihang University. My work explores how perception, understanding, reasoning, and execution can be integrated to support embodied systems that understand and interact with physical environments.

Research Interests

Publications

Current papers are ordered by Y Du's author position, then by citation count.

Citation counts updated on 2026/07/08.

Representative figure for Segvol publication

Segvol: Universal and interactive volumetric medical image segmentation

Y Du, F Bai, T Huang, B Zhao

Advances in Neural Information Processing Systems 37, 2025

Citations: 178

Representative figure for Focusable Monocular Depth Estimation publication

Focusable Monocular Depth Estimation

Y Du, T Lin, Z Zhong, R Li, X Chen, J Liu, C Liu, YC Chen, Y Fu, B Zhao

arXiv preprint, 2026

Citations: 1

Representative figure for M3D publication

M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

F Bai, Y Du, T Huang, MQH Meng, B Zhao

arXiv preprint, 2024

Citations: 243

Representative figure for Evo-Depth publication

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model

T Lin, Y Du, J Liu, N Zhu, Y Li, Y Fu, Y Chen, H Cai, Z Ye, B Cheng, K Ye, ...

arXiv preprint, 2026

Citations: 3

Representative figure for L4VLA publication

L4VLA: Learning to Act without Seeing via Language-Action Pretraining

T Lin, Y Du, Y Mao, Z Ye, Y Zhong, B Cheng, Y Wang, J Liu, Y Tian, J Yan, ...

arXiv preprint, 2026

Representative figure for Evo-1 publication

Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment

T Lin, Y Zhong, Y Du, J Zhang, J Liu, Y Chen, E Gu, Z Liu, H Cai, Y Zou, ...

arXiv preprint, 2025

Citations: 21

Placeholder figure for Semantic-aligned reinforced attention model publication

Semantic-aligned reinforced attention model for zero-shot learning

Z Yang, Y Zhang, Y Du, C Tong

Image and Vision Computing 128, 2022

Citations: 7

Representative figure for Evo-0 publication

Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding

T Lin, G Li, Y Zhong, Y Zou, Y Du, J Liu, E Gu, B Zhao

arXiv preprint, 2025

Citations: 52

Representative figure for Touchstone benchmark publication

Touchstone benchmark: Are we on the right way for evaluating AI algorithms for medical segmentation?

PRAS Bassi, W Li, Y Tang, F Isensee, Z Wang, J Chen, YC Chou, ...

Advances in Neural Information Processing Systems 37, 2025

Citations: 57

Awards & Highlights

Outstanding RBM Project 2026 (2nd rank out of 82 groups).

Award image for Outstanding RBM Project 2026

Contact

Email: yuxindu444@gmail.com

Google Scholar: Scholar profile

I welcome academic discussions and opportunities for research collaboration.