Motion PointNet: Solving Dynamic Capture in Point Cloud Video Human Action
Zhuoxu Huang, Zhenkun Fan, Tao Xu, Jungong Han
OpenReview ground truth
Abstract
Motion representation plays a pivotal role in understanding video data, thereby elevating the dynamic capture to the forefront of action recognition tasks based on point cloud video. Previous works mainly compute the motion information in an unguided way, e.g. aggregate the spatial variations on adjacent point cloud frames using 4D convolutions or capture a point trajectory with kinematic computation like scene flow. However, the former fails to explicitly consider motion representation in corresponding frames, and the latter's reliance on tracking point trajectories becomes impractical in real-life applications due to the potential inter-frame migration of points. In this paper, we tackle the dynamic capture in point cloud video action by formulating it as solvable partial differential equations (PDEs) in feature space. Based on this intuitive design, we propose Motion PointNet, a novel method that improves the dynamic capture in point cloud video human action by constructing clear guidance for network learning. Motion PointNet is composed of a lightweight yet effective PointNet-like encoder and a PDEs-solving module for dynamic capture. Remarkably, our Motion PointNet, with merely 0.72 M parameters and 0.82 G FLOPs, achieves an impressive accuracy of 97.52 % on the MSRAction-3D dataset, surpassing the current state-of-the-art in all aspects. The code and the trained models will be released for reproduction.
Author context
Most prolific author: 2 submissions (credibility 1.00).
No mass-submission penalty for this paper (authors within normal submission volume).
Aggregate statistics only — no individual author rankings.
Ranking trajectory
Percentile by tournament round — convergence indicates rating stability.
Battle history — 40 comparisons
Ranked above opponent in 53% of matchups.
- ▲ beat PASTA: Pretrained Action-State Transformer… ×6
- ▲ beat Heterogeneity of Regularization between ad… ×4
- ▼ lost to Demystifying CLIP Data ×4
- ▲ beat PATHS: Parameter-wise Adaptive Two-Stage T… ×4
- ▲ beat Extending to New Domains without Visual an… ×4
Judge assessments
Mean overall score 0.0 ± 0.0 (n = 40)