注册并分享邀请链接,可获得视频播放与邀请奖励。

huaijiang 的个人资料封面
huaijiang 的头像

huaijiang (@huaijiangzhu)

@huaijiangzhu
0 正在关注    0 粉丝
ok, so if i understood this correctly, the “in-context learning” ability was essentially baked into the model at training time. roughly speaking, the policy is trained to do something like: (demo video, current obs) → action this is quite different from in-context learning in llms, where the model is simply trained on next-token prediction over arbitrary context. fwiw, this is pretty clever because it could also solve another problem: they could pair an egocentric video-only demonstration with a UMI/teleop trajectory for the same task. the target trajectory provides the action labels, while the video-only demonstration provides the task context. that way, the policy can learn from action-labeled robot data while also learning how to translate an unlabeled video demonstration into robot actions.
显示更多
0
6
110
7
转发到社区
my spicy theory is that chinese ai labs keep winning because of culture, not talent or resources. chinese labs still have a deeply hands-on engineering culture. deepseek and kimi are flat organizations where science and engineering are fused: the same people move between algorithms, data, and infra, doing whatever it takes to make the model work. but sf tech bros have decided that “researcher” is the high-status title while infra is merely support work. every new sf ai startup calls itself a “lab,” every ambitious engineer quietly upgrades their title to “research engineer,” and the infra work is left to whoever failed to escape it. but at frontier scale, infra determines experiment velocity, and experiment velocity determines research output. at frontier scale, infra **is** research.
显示更多
0
253
5.9K
363
转发到社区