Асосий контентга ўтиш
Мақола

Motion-Guided Dynamic-Graph Construction with Kinematic-Aware Transformer for Skeleton Action Recognition

Kabul KhudaybergenovDepartment of Applied Informatics, Kimyo International University in Tashkent, Tashkent 100121, UzbekistanMarakhimov Avazjon RakhimovichDepartment of Information Processing and Management Systems, Tashkent State Technical University, Tashkent 100095, Uzbekistan
2026en
ABI

Аннотация

Skeleton-based action recognition has attracted considerable research interest because skeleton data are inherently robust to illumination changes, viewpoint variation, background clutter, and camera motion. Nevertheless, extracting informative representations from skeleton sequences remains a challenging problem, as it requires capturing both the spatial co-occurrence patterns among body joints and the fine-grained kinematic cues that distinguish different actions. In this paper, we propose a single-stream architecture that constructs an action-specific skeleton graph directly from motion and processes it with a kinematic-aware Transformer. Rather than relying on a fixed skeleton topology, a motion-guided dynamic-graph construction module infers a per-frame adjacency matrix from short-term motion cues through a differentiable edge predictor and Gumbel-Softmax sparsification, allowing the model to discover action-driven connections between distant joints that lack direct bone connectivity (e.g., coordinated hand motion during clapping). Each joint is described by kinematic node features that combine its 3D position, instantaneous velocity, and limb-angle encodings within a single descriptor, so that both motion dynamics and higher-order limb configurations are available to the spatial encoder from the outset. A graph-attention network (GAT) encodes the spatial configuration of every frame over the learned graph, and the resulting sequence of frame descriptors is processed by a Transformer encoder that models long-range temporal dependencies; a learnable classification token aggregates the sequence, and a multi-layer perceptron (MLP) produces the final action classification. The entire model is trained end-to-end from action labels alone. We conduct a comprehensive ablation study and evaluate the proposed method on the large-scale NTU RGB+D 60 and NTU RGB+D 120 benchmarks, where the results demonstrate that our approach achieves competitive performance compared to state-of-the-art architectures.

Ҳали таржима қилинмаган

Идентификаторлар

Иқтибослар ва манбалар