Dong Won Lee

Hi, my name is Dong Won. My advisors and friends usually call me “Don" :)

I’m a final-year PhD Student at MIT, where I develop multimodal foundation models for human-robot social interaction to build real-world/real-time robotic systems for speech, perception, expression powering robots like Jibo, Sprout, Reachy Mini, and Astro. I am grateful to be advised by Prof. Cynthia Breazeal, Dr. Hae Won Park, and Prof. Louis-Philippe Morency. Prior to MIT, I graduated from CMU with M.S./B.S. in Machine Learning. My research is generously supported by the Amazon AI Fellowship . I also serve as an E14 VC Fund Fellow at MIT.

Google Scholar  /  LinkedIn

Email: dongwonl_at_mit_dot_edu

profile photo
News

06/2026: Successfully proposed my PhD thesis, Multimodal Foundation Models as Communicative Robot Action Models.
04/2026: Our paper, Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models’ Social Reasoning , was accepted to RSS 2026.
09/2025: Honored to be named an Amazon AI Research Innovation Fellow , supporting my doctoral research on multimodal social intelligence and foundation models for embodied AI.
09/2025: Selected as a 2025–2027 E14 VC Fund Fellow at MIT , working with the E14 Fund on early-stage venture activities.
09/2025: Our paper, Aligning Dialogue Agents with Global Feedback via Large Language Model Multimodal Reward Decomposition , was accepted to EMNLP 2025 Findings.
06/2025: Completed the MIT Research Mentoring Certificate Program, focused on effective and inclusive research mentoring.
04/2025: Gave a Rising Stars Student Spotlight talk at the MIT Media Lab Members’ Event on Embodied AI Agents in Real-World Social Interactions.
09/2024: Our paper, Global Reward to Local Rewards: Multimodal-Guided Decomposition for Improving Dialogue Agents , was accepted to EMNLP 2024 as an oral presentation.
06/2024: Joined Microsoft Research NYC as a research intern working on Human-Oriented AI.
08/2023: Joined MIT MAS.630: Advanced Seminar: Affective Computing and Ethics as a graduate teaching assistant.
07/2023: Our paper, Lecture Presentations Multimodal Dataset: Towards Understanding Multimodality in Educational Videos , was accepted to ICCV 2023.
07/2023: Our paper, HIINT: Historical, Intra-and Inter-personal Dynamics Modeling with Cross-person Memory Transformer , was accepted to ICMI 2023.
04/2023: Our paper, Multipar-T: Multiparty-Transformer for Capturing Contingent Behaviors in Group Conversations , was accepted to IJCAI 2023 as an oral presentation.
03/2023: Our proposal for the 1st Workshop on Social and Affective Intelligence (SAI) was accepted at ACII 2023.
03/2022: Our paper, Low-resource Adaptation for Personalized Co-Speech Gesture Generation , was accepted to CVPR 2022.
05/2021: We organized the First Workshop on Crossmodal Social Animation at ICCV 2021.
09/2020: Our paper, No Gestures Left Behind: Learning Relationships between Spoken Language and Freeform Gestures , was accepted to EMNLP 2020 Findings.
07/2020: Our paper, Style Transfer for Co-Speech Gesture Animation: A Multi-Speaker Conditional Mixture Approach , was accepted to ECCV 2020.
Selected Publications
paper thumbnail A Modern System Recipe for Situated Embodied Human–Robot Conversation with Real-Time Multimodal LLMs and Tool-Calling
Dong Won Lee, Sarah Gillet, Louis-Philippe Morency, Cynthia Breazeal, Hae Won Park
RSS Beyond the Lab Workshop, 2026
paper / website

We present a minimal system recipe for situated embodied conversation that pairs a real-time multimodal language model with a small set of tools for attention and active perception, enabling robots to interleave dialogue with “what to look at, when to look, and what to say” under tight latency constraints.

paper thumbnail Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models’ Social Reasoning
Dong Won Lee, Yubin Kim, Denison Guvenoz, Sooyeon Jeong, Parker Malachowsky, Louis-Philippe Morency, Cynthia Breazeal, Hae Won Park
RSS, 2026
website / paper

We introduce a large-scale collection of datasets of real-world human-robot interaction videos with 10K+ annotations to benchmark AI models' ability to identify and reason about social interactions, providing a foundation for advancing socially intelligent AI.

paper thumbnail Aligning Dialogue Agents with Global Feedback via Large Language Model Reward Decomposition
Dong Won Lee, Hae Won Park, Cynthia Breazeal, Louis-Philippe Morency
EMNLP (Findings), 2025
paper

We introduce a framework that uses a frozen large language model (LLM) to decompose global session-level feedback into fine-grained turn-level rewards for dialogue agents. Our method works in both text-only and multimodal settings using cues such as pitch and gaze, enabling reinforcement learning without dense supervision. The resulting reward models improve dialogue quality in human evaluations.

paper thumbnail Global Reward to Local Rewards: Multimodal-Guided Decomposition for Improving Dialogue Agents
Dong Won Lee, Hae Won Park, Yoon Kim, Cynthia Breazeal, Louis-Philippe Morency
EMNLP, 2024 (Oral)
paper / code / huggingface

We introduce an approach named GELI, which automatically decomposes a single Global Explicit post-interaction score while incorporating Local Implicit feedback from multimodal signals to adapt a language model to become more conversational.

paper thumbnail HIINT: Historical, Intra-and Inter-personal Dynamics Modeling with Cross-person Memory Transformer
Yubin Kim, Dong Won Lee, Paul Pu Liang, Sharifa Algohwinem, Cynthia Breazeal, Hae Won Park
ICMI, 2023
paper

We model Historical, Intra-and Inter-personal (HIINT) Dynamics in conversation by incorporating memory modules in the Cross-person Memory Transformer to address temporal coherence and better represent the context of conversational behaviors.

paper thumbnail Multipar-T: Multiparty-Transformer for Capturing Contingent Behaviors in Group Conversations
Dong Won Lee, Yubin Kim, Rosalind Picard, Cynthia Breazeal, Hae Won Park
IJCAI, 2023 (Oral)
paper

We introduce a new transformer architecture to model contingent behaviors in multiparty group conversations.

paper thumbnail Low-resource Adaptation for Personalized Co-Speech Gesture Generation
Chaitanya Ahuja, Dong Won Lee, Louis-Philippe Morency
CVPR, 2022
paper

We propose a new approach in crossmodal generative modeling in low-resource settings to create personalized gesture generation models with limited data from a new speaker.

paper thumbnail No Gestures Left Behind: Learning Relationships between Spoken Language and Freeform Gestures
Chaitanya Ahuja, Dong Won Lee, Ryo Ishii, Louis-Philippe Morency
EMNLP, Findings, 2020
paper / code

We study relationships between spoken language and co-speech gestures to account for the long tail of the text-gesture distribution.

paper thumbnail Style Transfer for Co-Speech Gesture Animation: A Multi-Speaker Conditional Mixture Approach
Chaitanya Ahuja, Dong Won Lee, Yukiko I. Nakano, Louis-Philippe Morency
ECCV, 2020
project page / paper / code

We propose a new style transfer model to learn individual styles of speakers' gestures.

Teaching
MIT Affective Computing MIT MAS.630: Advanced Seminar: Affective Computing and Ethics
Graduate TA, Fall 2023
CMU LTI CMU LTI 11-777: Multimodal Machine Learning
Graduate TA, Spring 2022
CMU Machine Learning CMU MLD 10-725: Convex Optimization
Graduate TA, Spring 2021
CMU Statistics and Data Science CMU Stat & DS 36-202: Statistics & Data Science Methods
Undergraduate TA, Fall 2019, Spring 2020, Fall 2020 (3 Semesters)
CMU Statistics and Data Science CMU Stat & DS 36-200: Reasoning with Data
Undergraduate TA, Fall 2020, Spring 2021 (2 Semesters)
Service
EMNLP Empirical Methods in Natural Language Processing (EMNLP 2021, 2024)
Reviewer
ACII 2023 Social and Affective Intelligence (SAI) @ ACII 2023
Co-Organizing Chair
workshop page
CVF First Workshop on Crossmodal Social Animation @ ICCV 2021
Publication Chair
workshop page / video
ACM International Conference on Multimodal Interaction (ICMI 2021)
Reviewer

Website Credits Here: source code