Please select the desired project time frame:
- July 2026
- January 2026
- July 2025
- January 2025
- July 2024
- January 2024
- July 2023
- January 2023
- July 2022
- January 2022
- July 2021
- January 2021
- July 2020
- January 2020
- July 2019
- January 2019
- July 2018
- January 2018
- July 2017
- January 2017
- July 2016
- January 2016
- July 2015
- January 2015
- July 2014
- January 2014
- July 2013
- January 2013
- July 2012
- January 2012
- July 2011
- January 2011
- July 2010
- January 2010
- July 2009
- January 2009
- July 2008
- January 2008
- July 2007
- January 2007
- July 2006
- January 2006
- July 2005
- January 2005
- July 2004
- January 2004
- July 2003
- January 2003
- July 2002
- January 2002
- July 2001
- January 2001
Start of funding 01.01.2025
Unpacking Failure Modes of Generative Robot Policies
Prof. Dr. Matthias Althoff
Technische Universität München
Informatik 6 - Lehrstuhl für Robotik, Künstliche Intelligenz und Echtzeitsysteme
Christopher Agia
Stanford University
Interactive Perception and Robot Learning Lb, Department of Computer Science
Robot behavior policies trained via imitation learning are prone to failure under conditions that deviate from their training data. Thus, algorithms that monitor learned policies at test time and provide early warnings of failure are necessary to facilitate scalable deployment. We propose a runtime monitoring framework that splits the detection of failures into two complementary categories: 1) Erratic failures, which we propose to detect using a statistical measure of temporal action consistency, and 2) task progression failures, where we use VLMs to detect when the policy confidently and consistently takes actions that do not solve the task. Our approach has two key strengths. First, because learned policies exhibit diverse failure modes, combining complementary detectors leads to significantly higher accuracy at failure detection. Second, using a statistical temporal action consistency measure ensures that we quickly detect when multi-modal, generative policies exhibit erratic behavior at negligible computation cost. In contrast, we only use VLMs to detect failure modes that are less time sensitive. We demonstrate our approach in the context of diffusion policies trained on a multi-modal domain and robotic mobile manipulation domains both in simulation and in the real world.