Nan DUAN (段楠)
Google Scholar | LinkedIn | 京东探索研究院
We are hiring researchers and interns: duannan@jd.com (official) or nanduan.nlp@outlook.com (personal).
Dr. Nan Duan is Vice President of JD.COM and Deputy Director of Joy Future Academy, where he leads foundation model research spanning language, audio, vision, and embodied AI. Previously, he served as a Technical Fellow at StepFun and as a Senior Principal Researcher at Microsoft Research Asia. Dr. Duan is an adjunct professor and Ph.D. supervisor at the University of Science and Technology of China, Xi’an Jiaotong University, Xiamen University, and Tianjin University. His research interests include natural language processing, code agents, multimodal foundation models, and embodied intelligence. He has published more than 200 research papers in top-tier conferences and journals, with over 40,000 citations and an h-index of 85+, and holds more than 20 patents. In 2019, he was named a CCF-NLPCC Distinguished Young Scientist for his contributions to natural language processing. In 2023, he was recognized as one of DeepTech China’s Intelligent Computing Innovators for his contributions to AI foundation models.
段楠博士,现任京东集团副总裁、京东探索研究院副院长,领导语言、语音、视觉和具身智能领域的基础模型研究。此前,他曾任阶跃星辰 Technical Fellow、微软亚洲研究院资深首席研究员。段博士现为中国科学技术大学、西安交通大学、厦门大学和天津大学兼职教授、博士生导师。他的研究方向包括自然语言处理、代码智能体、多模态基础模型和具身智能。在顶级会议和期刊上发表论文200余篇,论文累计引用超过40,000次,h-index达85+,拥有20余项专利。2019年,他因在自然语言处理领域的贡献获评CCF-NLPCC杰出青年科学家。2023年,他因在人工智能基础模型领域的贡献入选DeepTech中国智能计算创新人物。
Highlight
- JoyAI-Echo & JoyAI-Echo-1.5 (Technical Report, 2026), SoTA audio-video generation & interactive world models.
- JoyAI-VL-Interaction (Technical Report, 2026), a SoTA real-time vision-language interaction model.
- JoyAI-Video-Edit (Technical Report, 2026), a SoTA real-time streaming video editing model.
- JoyAI-Image-Edit (Technical Report, 2026), a unified image understanding and generation model with spatial intelligence.
- Step-Video-T2V (Technical Report, 2025), a SoTA open-source text-to-video model.
- scGPT (Nature Methods, 2024), a Generative Pre-trained Transformer for single-cell biology.
- Not All Tokens Are What You Need (NeurIPS, 2024), NeurIPS 2024 Best Paper Runner-Up.
- Visual ChatGPT (Preprint, 2023), pioneering work on multimodal AI agents, with 34K+ GitHub stars.
- VL-InterpreT (CVPR, 2022), recipient of the CVPR 2022 Best Demo Award.
- NUWA(女娲) (ECCV, 2022) & NUWA-Infinity (NeurIPS, 2022), pioneering work on video generation models, reviewed by Bill Gates and cited by OpenAI Sora.
- CodeBERT (EMNLP, 2020) & CodeXGLUE (NeurIPS, 2021), pioneering work on code foundation models, cited by OpenAI Codex.
- Unicoder (EMNLP, 2019) & Unicoder-VL (AAAI, 2020), the 1st multilingual & multimodal pre-trained models deployed in Microsoft Bing for 100+ languages.
Academic Service & Award
- Adjunct Ph.D. Supervisor at University of Science and Technology of China (中国科学技术大学), Xi’an Jiaotong University (西安交通大学), Xiamen University (厦门大学), and Tianjin University (天津大学).
- Program Committee Chair of NLPCC, 2023.
- Senior Action Editor & Senior Area Chair for ACL Rolling Review (ARR), NeurIPS, ACL, EMNLP, NAACL, and SIGKDD.
- Standing Reviewer for TACL, 2020-present.
- Executive Member of the China Society of Image and Graphics (CSIG), 2025-present.
- Executive Member of the CCF Technical Committee of NLP, 2018-present.
- Senior Member of IEEE.
- Distinguished Member of CCF.
- AI 2000 Most Influential Scholar Award Honorable Mention in NLP, 2025.
- Stanford/Elsevier World’s Top 2% Scientists, 2022-present.
- Intelligent Computing Innovators China (中国智能计算科技创新人物), 2023.
- CCF-NLPCC Distinguished Young Scientist Award (CCF-NLPCC青年科学家奖), 2019.
- NeurIPS Best Paper Runner-Up Award, 2024.
- CVPR Best Demo Award, 2022.