Journals / Turkish Journal of Electrical Engineering and Computer Sciences / 2020 / Cilt: 28 - Sayı: 3

Deep temporal motion descriptor (DTMD) for human action recognition

Pages
1371–1385
DOI
—

Abstract

Spatiotemporal features have significant importance in human action recognition, as they provide the actor’sshape and motion characteristics specific to each action class. This paper presents a new deep spatiotemporal humanaction representation, the deep temporal motion descriptor (DTMD), which shares the attributes of holistic and deeplearned features. To generate the DTMD descriptor, the actor’s silhouettes are gathered into single motion templatesby applying motion history images. These motion templates capture the spatiotemporal movements of the actor andcompactly represent the human actions using a single 2D template. Then deep convolutional neural networks are usedto compute discriminative deep features from motion history templates to produce the DTMD. Later, DTMD is usedfor learning a model to recognize human actions using a softmax classifier. The advantage of DTMD are that DTMDis automatically learned from videos and contains higher-dimensional discriminative spatiotemporal representations ascompared to handcrafted features; DTMD reduces the computational complexity of human activity recognition as allthe video frames are compactly represented as a single motion template; and DTMD works effectively for single andmultiview action recognition. We conducted experiments on three challenging datasets: MuHAVI-Uncut, iXMAS, andIAVID-1. The experimental findings reveal that DTMD outperforms previous methods and achieves the highest actionprediction rate on the MuHAVI-Uncut dataset.