扶摇AI知识笔记AI 前沿知识库
首页 / 物理 AI / 正文
物理 AI

PointZero: 3D Point Track Completion for Learning Transferable 3D Dynamics

来源:arXiv cs.CV 论文速递 约 2159 字 robot
arXiv cs.CV
转载

本文转载自 arXiv cs.CV,版权归原作者及原发布平台所有。本站仅作知识整理与转载分享,如涉版权问题请联系客服删除。

01核心要点

  • World models endow perceptual systems with the ability to predict how scenes evolve under interaction.
  • They are most beneficial when trained on diverse volumes of data, to instill a rich prior into downstream applications.
  • Existing methods typically require robot action labels to learn action-conditioned 3D dynamics, which excludes web video data from the training pool.

02正文全文

Abstract:World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into downstream applications. Existing methods typically require robot action labels to learn action-conditioned 3D dynamics, which excludes web video data from the training pool. We study 3D point track completion as a pre-training objective for learning transferable 3D dynamics without robot data. Given a single RGB-D observation and sparse partial 3D trajectories (tracks), we predict future 3D tracks of all observed points. We show this objective produces a rich 3D dynamics prior, without requiring robot action labels. We contribute a diverse dataset of 2.9 million synthetic frames spanning deformable, articulated, and rigid objects, and use it to train PointZero. We show that a flexible and expressive transformer, PointZero, outperforms prior methods on the same data. We demonstrate the utility of our pre-training objective by post-training PointZero for two downstream applications: (1) action-conditioned 3D dynamics prediction and (2) imitation learning. When fine-tuned to condition on end-effector pose, PointZero outperforms the baselines on the recent PGND 3D dynamics benchmark. When fine-tuned to predict robot actions and 3D tracks, PointZero outperforms or matches the baselines on 6/7 simulated and real-world robot manipulation tasks. We furthermore evaluate training PointZero from scratch to isolate the benefits of our proposed architecture from those of our proposed pre-training objective and dataset. We release the dataset, checkpoints, and full training recipe.

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

03原文直达

本文内容转载自 arXiv cs.CV,如需查看原排版、配图与最新修订,请访问原始出处。

阅读原文(arXiv cs.CV)

下载论文 PDF

正在校验阅读权限…
RELATED

相关阅读

更多 物理 AI