Program
Sunday, September 27, 2026
Room 328 · Pittsburgh, PA · EDT (UTC−4)
Invited talks: 30 min + 5 min Q&A; Stefan Leutenegger: 25 min + 5 min Q&A.
Expand “Abstract & speaker bio” for submitted details.
Morning
Opening Remarks

Hanbyul Joo
Seoul National University
From Capturing People to Teaching Robots
Abstract & speaker bio
Abstract
Equipping AI and robotic systems with the ability to understand human behavior is essential for enabling them to assist people across a wide range of everyday applications. This need is more pressing than ever: the heaviest consumers of such knowledge are no longer perception systems alone, but robot policies that must learn to act in the physical world. Yet the high-quality 3D human motion data required to learn this knowledge remains extremely scarce.
In this talk, I will present our lab's efforts to scale and enrich 3D human motion data by capturing everyday movements and natural human-object interactions, with the ultimate goal of teaching robots to move like humans.
I will first introduce our multi-year effort in building multi-camera capture systems, from Panoptic Studio to ParaHome, and most recently OmniRoboHome, a new system designed to capture human-object interactions in natural home environments. Next, I will present a complementary direction: learning everyday interactions and affordances from generative image and video models, which offer a scalable source of human behavior priors that capture systems alone cannot reach. Finally, I will discuss the missing pieces that vision and image models cannot provide, including physics, contact, and the gap across diverse embodiments, together with our recent efforts to fill them.
Bio
Hanbyul Joo is an associate professor at Seoul National University (SNU) in the Department of Computer Science and Engineering. Before joining SNU, Hanbyul was a Research Scientist at Facebook AI Research (FAIR), Menlo Park. Hanbyul received his PhD from the Robotics Institute at Carnegie Mellon University. Hanbyul is a recipient of the Best Student Paper Award at CVPR 2018.
Speaker profile
Guanya Shi
Carnegie Mellon University
Humanoid Learning from Human Data with Physics Grounding
Abstract
Abstract
Humanoid robots offer a versatile platform for performing a wide range of tasks in complex, human-centered environments. Their human-like embodiment makes human data a promising foundation for scalable skill learning. However, translating human skills into robot control requires bridging cross-embodiment gaps, particularly in dynamics and actuation.
In this talk, we present a general recipe for learning humanoid skills from human data with physics grounding. Human data provides high-level intent and skill structure, while a physics-grounded layer translates them into motor-level actions. This approach enables versatile, adaptive, and robust skill learning across a range of loco-manipulation and dexterous manipulation tasks.

Stefan Leutenegger
ETH Zurich · Remote
Robust, Real-Time, Multi-Sensor Perception for Open-World Robot Autonomy
Abstract & speaker bio
Abstract
For mobile robots to work full shifts, they need to cope with the messiness and unpredictability of the real world for hours at a time without accumulating failures—potentially even in open-world settings. Robust perception and 3D/4D scene understanding therefore remain key challenges. In this talk, I will present several recent works on real-time perception that exploit complementary sensing modalities, including event cameras, to improve robustness. I will also discuss approaches that lift open-vocabulary vision-language features into 3D representations, enabling semantic navigation and higher-level task execution on drones and other mobile robots.
Bio
Stefan Leutenegger is Associate Professor of Mobile Robotics at ETH Zurich, where he leads the Mobile Robotics Lab. Before joining ETH in 2025, he held positions at the Technical University of Munich and Imperial College London. He received his BSc, MSc, and PhD degrees from ETH Zurich. His research focuses on robust perception and autonomy for mobile robots, including visual-inertial navigation, SLAM, 3D scene understanding, and robot learning.
Speaker profile
Lukas Rosenberger Schmid
University of Technology Nuremberg
Embodied Spatio-Temporal AI: Long-term dynamic scene understanding in real-time
Abstract & speaker bio
Abstract
The ability to build an actionable understanding of the environment of a robot is crucial for autonomy and prerequisite to a large variety of applications, ranging from home, service, care and consumer robots to autonomous vehicles, augmented reality, and disaster response. Notably, much of this depends on long-term operation in human-centric domains that are complex, widely diverse, and highly dynamic. This talk presents an autonomy pipeline to address these challenges. At the core, I argue that time, and thus memory, dynamics, and adaptation, should be an integral component of AI systems. I will introduce methods to capture the present and past of complex and dynamic scenes through symbolic abstractions, which further facilitate predicting future scene outcomes. I will show how robots as embodied agents can leverage our actionable scene representations and predictions to complete tasks such as actively gathering data that helps them improve their scene models and perception capabilities, and how all these tools can combine for robots to fully autonomously learn over time.
The presented methods are demonstrated on-board fully autonomous aerial and ground robots, run in real-time on the limited hardware available, and are released as open-source software.
Bio
Lukas Rosenberger Schmid (formerly Schmid) is a Tenure-Track Professor of Machine Intelligence at UTN. Before that, he was a Research Scientist and Postdoctoral Fellow at the SPARK Lab led by Prof. Luca Carlone at MIT, and a Postdoctoral Researcher at the Autonomous Systems Lab (ASL) led by Prof. Roland Siegwart at ETH Zürich. He earned his PhD in 2022 from ASL at ETHZ, where he also was a visiting researcher at the Microsoft Spatial AI Lab led by Prof. Marc Pollefeys. His work has been recognized by several honors, including RSS Pioneers 2025, NOKOV New Generation Star at IROS 2025, the RSS Outstanding Systems Paper Award 2024, two ETH Medals for outstanding PhD and M.Sc. Theses, the Willi Studer Prize for the best graduate of the year at ETHZ, the first place in the 2024 Hilti SLAM challenge, and a Swiss National Science Foundation (SNSF) Postdoc Fellowship.
His research focuses on embodied spatio-temporal AI for human-centric robot autonomy. This includes research on scene representations and abstraction, on detection, prediction, and understanding of moving and changing entities, active perception and information gathering, as well as lifelong learning for continuous adaptation to the robot environment, embodiment, task, and human preference.
Speaker profile
Ram Vasudevan
University of Michigan
Any Curve Will Do: Finite-Time Heat Flow for Millisecond Trajectory Optimization
Abstract & speaker bio
Abstract
Trajectory optimization can make robots reach, dodge, and walk, but on high-dimensional systems such as humanoids it typically takes seconds to minutes and requires a good initial guess. This talk presents FLINT (Finite-time heat-fLow INTegration), a heat-flow method that needs no initial guess: it can start from any curve that connects the start to the goal. Drawing on the theory of fixed-time stability, FLINT discretizes the heat flow with a step size that grows as the flow approaches a solution, and we prove that it converges in a number of steps that is bounded independently of the requested accuracy. Across arms and humanoids, FLINT is 30 to 530 times faster than state-of-the-art direct, DDP-based, and earlier heat-flow solvers, with equal or better success when every solver is judged by the same criterion. FLINT plans collision-free arm motions among obstacles given as sensed point clouds in milliseconds, solves a 22-joint humanoid in 14 ms, and produces seven walking gaits for the Talos humanoid, including stair climbing, each from a standing start in under 0.3 s.
Bio
Ram Vasudevan is a professor in Mechanical Engineering and Robotics at the University of Michigan. He received a BS in Electrical Engineering and Computer Sciences, an MS degree in Electrical Engineering, and a PhD in Electrical Engineering all from the University of California, Berkeley. He is a recipient of the NSF CAREER Award, the ONR Young Investigator Award, and the 1938E Award from the University of Michigan. His work has received best paper awards at the IEEE Conference on Robotics and Automation, the ASME Dynamics Systems and Controls Conference, and IEEE International Conference on Biomedical Robotics and Biomechatronics, and has been finalist for best paper at Robotics: Science and Systems.
Speaker profilePoster / Demo Session I
Lunch & Afternoon
Lunch Break
Poster / Demo Session II

Jiajun Wu
Stanford University
Building Physical Agents via Structured Representations
Abstract & speaker bio
Abstract
Agentic systems have been widely deployed in the digital space and are now entering the physical world. In this talk, I will discuss our recent efforts in bringing such systems into robotic manipulation, with their strengths and limitations. I will in particular focus on how coding agents may connect to robotic task planning, and present our thinking on the connections between learning systems and structured representations across multiple levels of abstraction.
Bio
Jiajun Wu is an Assistant Professor of Computer Science and, by courtesy, of Psychology at Stanford University, working on computer vision, machine learning, robotics, and computational cognitive science. Before joining Stanford, he was a Visiting Faculty Researcher at Google Research. He received his PhD in Electrical Engineering and Computer Science from the Massachusetts Institute of Technology. Wu's research has been recognized through the IJCAI Computers and Thought Award, the Young Investigator Programs (YIP) by ONR and by AFOSR, the Early Career Program (ECP) by ARO, the NSF CAREER award, the Okawa research grant, and paper awards and finalists at ICCV, CVPR, SIGGRAPH Asia, ICRA, CoRL, and IROS.
Speaker profile
Steve Waslander
University of Toronto
Long-term Memory for Agentic Warehouse Robots
Abstract & speaker bio
Abstract
Agentic reasoning for robots is rapidly becoming a reality, allowing flexible natural language interaction with human operators and enabling a wide range of navigation, object handling and recall tasks in a variety of settings. In this talk, Prof. Waslander will discuss the ongoing efforts in his lab to make useful agentic robots for the warehouse automation, by integrating open world perception with agentic reasoning for reliable open world navigation, and by adding multi-faceted memory - spatial, descriptive and visual - to enable long-term experience recall for visual question answering. Together, these advances enable wide variety of spatial, semantic, functional and temporal tasks.
Bio
Prof. Steven Waslander is a leading authority on autonomous robotics, including self-driving cars and multirotor drones. He received his B.Sc.E.in 1998 from Queen’s University, his M.S. in 2002 and his Ph.D. in 2007, both from Stanford University in Aeronautics and Astronautics. He was recruited to the University of Waterloo from Stanford in 2008, where he led the Autonomoose project, the first self-driving car to be tested on public roads by a Canadian university. In 2018, he joined the University of Toronto Institute for Aerospace Studies (UTIAS), and founded the Toronto Robotics and Artificial Intelligence Laboratory (TRAILab).
Speaker profile
Qingyuan Jiang
Apple
Robot Planning for Human-Centered Perception
Abstract & speaker bio
Abstract
Robotic systems are increasingly expected to work closely and smoothly around humans, but the dynamic nature of human behavior makes this hard: a robot must continuously predict and respond to evolving human behavior rather than merely react to the current observation. This talk studies this challenge through the Robot Follow-Ahead (RFA) problem, where a robot must move in front of a walking human to maintain visibility of high-priority frontal areas such as facial expressions. I'll present a progression of methods that couple human motion prediction with robot path planning in a closed loop, with the help of diffusion models and Bayesian methods.
Bio
Qingyuan Jiang is a Machine Learning Engineer at Apple Camera Incubation team. He received his PhD degree from the University of Minnesota, advised by Prof. Volkan Isler. His research lies at the intersection of human motion prediction and robot motion planning, with a focus on perception-oriented human-robot interaction. His PhD work was supported in part by the 3M Fellowship, and an NSF NRI Award.
Speaker profileLloyd Lei
ArenaLabs · CMU
Lightning Talk
Speaker bio
Bio
Lloyd Lei is a CMU MSCS student who is currently on leave to build his startup, ArenaLabs, as its founder and CEO. The company just finished a $750K angel round and is building the largest across-embodiment dataset in the world.
He was one of the youngest YC China partners. He has held research internships at ByteDance (TikTok), Broadcom, and the Robotics and AI Institute (RAI), and participated in StarBot’s seed round. Before his academic career in AI, he was a visiting researcher at Lawrence Berkeley National Laboratory, studying particle physics.
His current research interests and investments mainly focus on Whole-Body Control, real-time and low-latency inference, and across-embodiment representation learning.
Industry Panel
Joshua Joseph
AMR Deployment Engineer · Tesla
Joshua Joseph is an Autonomous Mobile Robot (AMR) deployment engineer at Tesla, where he scales robotic systems inside Gigafactories. An active contributor to the A3 R15.08 and ASTM F45, F50 committees, his work bridges real-world factory operations, physical AI, and standards governance. Recognized as an ASME Watchlist honoree, Joshua serves as Director of the Logistics and Supply Chain Division at IISE and on the Industry Advisory Board at Northeastern University's MIE department.
ProfileQingyuan Jiang
Machine Learning Engineer · Apple
Qingyuan Jiang is a Machine Learning Engineer at Apple Camera Incubation team. He received his PhD degree from the University of Minnesota, advised by Prof. Volkan Isler. His research lies at the intersection of human motion prediction and robot motion planning, with a focus on perception-oriented human-robot interaction. His PhD work was supported in part by the 3M Fellowship, and an NSF NRI Award.
ProfileHao Li (Leo)
Founding Scientist & Head of Academic Partnerships · Ropedia · MMLab, NTU
Hao Li (Leo) is a Visiting Student at MMLab, Nanyang Technological University, Singapore, advised by Prof. Ziwei Liu, and a Founding Scientist & Head of Academic Partnerships at Ropedia. Previously, he held research internships at Meituan LongCat, ByteDance Seed, StepFun, and Baidu. He serves as Principal Investigator on a project funded by the National Natural Science Foundation of China (NSFC) under its Young Student Basic Research Program (Ph.D. Track). His research focuses on multimodal large language models, spatial intelligence, and embodied AI, with an emphasis on enabling intelligent agents to perceive, reason about, and interact with the physical world.
ProfileFan Wang
Senior Research Scientist · Amazon Robotics
Fan Wang received her B.Sc. (Hons.) in Electrical and Mechanical Engineering from the University of Edinburgh, U.K., followed by an M.Sc. in Electrical and Computer Engineering from the University of Delaware, USA. She earned her Ph.D. in Electrical and Computer Engineering from Duke University. Currently, she is a Senior Research Scientist at Amazon Robotics, where she focuses on robot manipulation for warehouse automation. Her research interests include robotics, planning, manipulation, machine learning, computer vision, and perception-based control of autonomous systems.
ProfileShumo Chu
Cofounder & CEO · General Intelligence Labs
Shumo Chu is the Cofounder and CEO of General Intelligence Labs. He received his Ph.D. in Computer Science from the University of Washington and was previously an Assistant Professor at the University of California, Santa Barbara. He founded two startups before General Intelligence Labs.
ProfileHanbyul Joo
Associate Professor · Seoul National University
Hanbyul Joo is an associate professor at Seoul National University (SNU) in the Department of Computer Science and Engineering. Before joining SNU, Hanbyul was a Research Scientist at Facebook AI Research (FAIR), Menlo Park. Hanbyul received his PhD from the Robotics Institute at Carnegie Mellon University. Hanbyul is a recipient of the Best Student Paper Award at CVPR 2018.
ProfileBowen Weng
Assistant Professor of Computer Science · Iowa State University
Dr. Bowen Weng is an Assistant Professor of Computer Science at Iowa State University. His research addresses how to build robots capable of dynamic physical interaction with and near humans, and how to establish with statistical rigor that they perform reliably and safely. His work is supported by the National Science Foundation and the National Institute of Standards and Technology. From 2016 until joining Iowa State, Dr. Weng worked as a Research Engineer, Technical Specialist, and Research Scientist with Transportation Research Center Inc., on assignment to the National Highway Traffic Safety Administration (NHTSA), U.S. Department of Transportation. During this period, he completed his Ph.D. in Electrical and Computer Engineering at The Ohio State University as a part-time student, graduating in 2023. He also currently chairs ASTM Subcommittee F45.06 on Legged Robot Systems.
ProfileWrap-up, Awards & Closing Remarks
Sponsors





