M30.7 CONNECT THE MECHANISM
Connect visual instructions to a robot’s action interface
"Put the chipped mug in the sink." To follow that, a robot needs web-scale knowledge and precise motor control. See how RT-2, OpenVLA, and π0 turn a language model into a robot policy.
LESSON OVERVIEW14 min lesson
Lesson overview
"Put the chipped mug in the sink." To follow that, a robot needs web-scale knowledge and precise motor control. See how RT-2, OpenVLA, and π0 turn a language model into a robot policy.
What you’ll explore
- Explain how a vision-language-action model turns images and an instruction into robot actions, convert continuous actions to and from 256-bin tokens, and compare token outputs (RT-2, OpenVLA) with a continuous flow-matching action head (π0) in precision, speed, and embodiment.
GO TO THE SOURCE
Original explanations, connected to the research.
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control (Brohan et al., 2023)OpenVLA: An Open-Source Vision-Language-Action Model (Kim et al., 2024)π0: A Vision-Language-Action Flow Model for General Robot Control (Black et al., 2024)Open X-Embodiment: Robotic Learning Datasets and RT-X Models (2023)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.