After injecting user data as a prior, we can train a new expert in ~20 minutes.
No complicated reward engineering. No expensive reward-model fine-tuning. Just sparse, simple rewards are enough to adapt the policy across tasks.
A scalable way to produce large numbers of high-quality robot experts with very low compute and human cost.
显示更多