Skepxels: Spatio-temporal Image Representation of Human Skeleton Joints for Action Recognition

Liu, Jian; Akhtar, Naveed; Mian, Ajmal

Computer Science > Computer Vision and Pattern Recognition

arXiv:1711.05941 (cs)

[Submitted on 16 Nov 2017 (v1), last revised 3 Aug 2018 (this version, v4)]

Title:Skepxels: Spatio-temporal Image Representation of Human Skeleton Joints for Action Recognition

Authors:Jian Liu, Naveed Akhtar, Ajmal Mian

View PDF

Abstract:Human skeleton joints are popular for action analysis since they can be easily extracted from videos to discard background noises. However, current skeleton representations do not fully benefit from machine learning with CNNs. We propose "Skepxels" a spatio-temporal representation for skeleton sequences to fully exploit the "local" correlations between joints using the 2D convolution kernels of CNN. We transform skeleton videos into images of flexible dimensions using Skepxels and develop a CNN-based framework for effective human action recognition using the resulting images. Skepxels encode rich spatio-temporal information about the skeleton joints in the frames by maximizing a unique distance metric, defined collaboratively over the distinct joint arrangements used in the skeletal image. Moreover, they are flexible in encoding compound semantic notions such as location and speed of the joints. The proposed action recognition exploits the representation in a hierarchical manner by first capturing the micro-temporal relations between the skeleton joints with the Skepxels and then exploiting their macro-temporal relations by computing the Fourier Temporal Pyramids over the CNN features of the skeletal images. We extend the Inception-ResNet CNN architecture with the proposed method and improve the state-of-the-art accuracy by 4.4% on the large scale NTU human activity dataset. On the medium-sized N-UCLA and UTH-MHAD datasets, our method outperforms the existing results by 5.7% and 9.3% respectively.

Comments:	Submitted to a journal
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1711.05941 [cs.CV]
	(or arXiv:1711.05941v4 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1711.05941

Submission history

From: Naveed Akhtar Dr. [view email]
[v1] Thu, 16 Nov 2017 06:03:52 UTC (3,508 KB)
[v2] Tue, 9 Jan 2018 11:44:46 UTC (3,509 KB)
[v3] Mon, 16 Apr 2018 06:37:21 UTC (3,581 KB)
[v4] Fri, 3 Aug 2018 06:57:39 UTC (5,701 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Skepxels: Spatio-temporal Image Representation of Human Skeleton Joints for Action Recognition

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Skepxels: Spatio-temporal Image Representation of Human Skeleton Joints for Action Recognition

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators