Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation

Gervet, Theophile; Xian, Zhou; Gkanatsios, Nikolaos; Fragkiadaki, Katerina

Computer Science > Robotics

arXiv:2306.17817 (cs)

[Submitted on 30 Jun 2023 (v1), last revised 19 Oct 2023 (this version, v2)]

Title:Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation

Authors:Theophile Gervet, Zhou Xian, Nikolaos Gkanatsios, Katerina Fragkiadaki

View PDF

Abstract:3D perceptual representations are well suited for robot manipulation as they easily encode occlusions and simplify spatial reasoning. Many manipulation tasks require high spatial precision in end-effector pose prediction, which typically demands high-resolution 3D feature grids that are computationally expensive to process. As a result, most manipulation policies operate directly in 2D, foregoing 3D inductive biases. In this paper, we introduce Act3D, a manipulation policy transformer that represents the robot's workspace using a 3D feature field with adaptive resolutions dependent on the task at hand. The model lifts 2D pre-trained features to 3D using sensed depth, and attends to them to compute features for sampled 3D points. It samples 3D point grids in a coarse to fine manner, featurizes them using relative-position attention, and selects where to focus the next round of point sampling. In this way, it efficiently computes 3D action maps of high spatial resolution. Act3D sets a new state-of-the-art in RL-Bench, an established manipulation benchmark, where it achieves 10% absolute improvement over the previous SOTA 2D multi-view policy on 74 RLBench tasks and 22% absolute improvement with 3x less compute over the previous SOTA 3D policy. We quantify the importance of relative spatial attention, large-scale vision-language pre-trained 2D backbones, and weight tying across coarse-to-fine attentions in ablative experiments. Code and videos are available on our project website: this https URL.

Subjects:	Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2306.17817 [cs.RO]
	(or arXiv:2306.17817v2 [cs.RO] for this version)
	https://doi.org/10.48550/arXiv.2306.17817

Submission history

From: Theophile Gervet [view email]
[v1] Fri, 30 Jun 2023 17:34:06 UTC (11,851 KB)
[v2] Thu, 19 Oct 2023 19:36:31 UTC (19,528 KB)

Computer Science > Robotics

Title:Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Robotics

Title:Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators