Robots that learn a new task from a dozen tries.
We build control policies that pick up new work from a handful of demonstrations instead of a warehouse of logged data. The next step is doing it from none.
- Sample efficiency
- A new task costs demonstrations you can count, not months of teleoperation. Twelve is an ordinary number here.
- Transfer
- One policy holds up across grippers, objects and lighting it was never shown while training.
- Generalization
- The long bet: a robot that is told what to do rather than taught, and gets it right the first time.