Ponymation: Learning 3D Animal Motions from Unlabeled Online Videos

Sun, Keqiang; Litvak, Dor; Zhang, Yunzhi; Li, Hongsheng; Wu, Jiajun; Wu, Shangzhe

Computer Science > Computer Vision and Pattern Recognition

arXiv:2312.13604v1 (cs)

[Submitted on 21 Dec 2023 (this version), latest version 31 Jul 2024 (v3)]

Title:Ponymation: Learning 3D Animal Motions from Unlabeled Online Videos

Authors:Keqiang Sun, Dor Litvak, Yunzhi Zhang, Hongsheng Li, Jiajun Wu, Shangzhe Wu

View PDF HTML (experimental)

Abstract:We introduce Ponymation, a new method for learning a generative model of articulated 3D animal motions from raw, unlabeled online videos. Unlike existing approaches for motion synthesis, our model does not require any pose annotations or parametric shape models for training, and is learned purely from a collection of raw video clips obtained from the Internet. We build upon a recent work, MagicPony, which learns articulated 3D animal shapes purely from single image collections, and extend it on two fronts. First, instead of training on static images, we augment the framework with a video training pipeline that incorporates temporal regularizations, achieving more accurate and temporally consistent reconstructions. Second, we learn a generative model of the underlying articulated 3D motion sequences via a spatio-temporal transformer VAE, simply using 2D reconstruction losses without relying on any explicit pose annotations. At inference time, given a single 2D image of a new animal instance, our model reconstructs an articulated, textured 3D mesh, and generates plausible 3D animations by sampling from the learned motion latent space.

Comments:	Project page: this https URL. The first two authors contributed equally to this work. The last two authors contributed equally
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2312.13604 [cs.CV]
	(or arXiv:2312.13604v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2312.13604

Submission history

From: Keqiang Sun [view email]
[v1] Thu, 21 Dec 2023 06:44:18 UTC (11,751 KB)
[v2] Tue, 30 Jul 2024 15:49:16 UTC (10,313 KB)
[v3] Wed, 31 Jul 2024 18:40:17 UTC (10,313 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Ponymation: Learning 3D Animal Motions from Unlabeled Online Videos

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Ponymation: Learning 3D Animal Motions from Unlabeled Online Videos

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators