GaussianGrasper: 3D Language Gaussian Splatting for Open-vocabulary Robotic Grasping

Zheng, Yuhang; Chen, Xiangyu; Zheng, Yupeng; Gu, Songen; Yang, Runyi; Jin, Bu; Li, Pengfei; Zhong, Chengliang; Wang, Zengmao; Liu, Lina; Yang, Chao; Wang, Dawei; Chen, Zhen; Long, Xiaoxiao; Wang, Meiqing

Computer Science > Robotics

arXiv:2403.09637 (cs)

[Submitted on 14 Mar 2024]

Title:GaussianGrasper: 3D Language Gaussian Splatting for Open-vocabulary Robotic Grasping

Authors:Yuhang Zheng, Xiangyu Chen, Yupeng Zheng, Songen Gu, Runyi Yang, Bu Jin, Pengfei Li, Chengliang Zhong, Zengmao Wang, Lina Liu, Chao Yang, Dawei Wang, Zhen Chen, Xiaoxiao Long, Meiqing Wang

View PDF HTML (experimental)

Abstract:Constructing a 3D scene capable of accommodating open-ended language queries, is a pivotal pursuit, particularly within the domain of robotics. Such technology facilitates robots in executing object manipulations based on human language directives. To tackle this challenge, some research efforts have been dedicated to the development of language-embedded implicit fields. However, implicit fields (e.g. NeRF) encounter limitations due to the necessity of processing a large number of input views for reconstruction, coupled with their inherent inefficiencies in inference. Thus, we present the GaussianGrasper, which utilizes 3D Gaussian Splatting to explicitly represent the scene as a collection of Gaussian primitives. Our approach takes a limited set of RGB-D views and employs a tile-based splatting technique to create a feature field. In particular, we propose an Efficient Feature Distillation (EFD) module that employs contrastive learning to efficiently and accurately distill language embeddings derived from foundational models. With the reconstructed geometry of the Gaussian field, our method enables the pre-trained grasping model to generate collision-free grasp pose candidates. Furthermore, we propose a normal-guided grasp module to select the best grasp pose. Through comprehensive real-world experiments, we demonstrate that GaussianGrasper enables robots to accurately query and grasp objects with language instructions, providing a new solution for language-guided manipulation tasks. Data and codes can be available at this https URL.

Subjects:	Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2403.09637 [cs.RO]
	(or arXiv:2403.09637v1 [cs.RO] for this version)
	https://doi.org/10.48550/arXiv.2403.09637

Submission history

From: Yupeng Zheng [view email]
[v1] Thu, 14 Mar 2024 17:59:46 UTC (5,358 KB)

Computer Science > Robotics

Title:GaussianGrasper: 3D Language Gaussian Splatting for Open-vocabulary Robotic Grasping

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Robotics

Title:GaussianGrasper: 3D Language Gaussian Splatting for Open-vocabulary Robotic Grasping

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators