Skip to main content
Article

Hybrid Invariant Latent Feature Graph Transformer for Skeleton-Based Human Action Recognition

Kabul KhudaybergenovDepartment of Applied Informatics, Kimyo International University in Tashkent, Tashkent 100121, UzbekistanMarakhimov Avazjon RakhimovichDepartment of Information Processing and Management Systems, Tashkent State Technical University, Tashkent 100174, Uzbekistan
2026en
ABI

Abstract

Skeleton-based human action recognition is an important problem in applied vision systems, yet many existing approaches depend on a single skeleton descriptor or a single feature-learning mechanism. This restriction can weaken the representation of local body kinematics, long-range joint relations, and temporal dependencies within an action sequence. To address these limitations, this paper proposes HILF-GT (Hybrid Invariant Latent Feature Graph Transformer), a hybrid Graph Convolutional Network (GCN)-Transformer framework based on multiple spatio-temporal invariant latent features. The representation module constructs complementary structured tensors from skeleton graphs, inter-joint distances, adjacent-frame joint displacements, and inter-limb angles. Instead of transforming these descriptors into image-like maps for separate Convolutional Neural Network (CNN)-based classification, HILF-GT keeps their graph and temporal organization during learning. A local GCN branch models skeleton-aware kinematic patterns, whereas a graph-aware Transformer branch uses biased self-attention and cross-attention to capture dependencies among distant joints, frames, and latent-feature streams. A Perceiver-style latent bottleneck is further introduced to reduce the memory cost of global attention over frame-joint tokens. Experiments were conducted on four standard benchmark datasets, including NTU-RGB+D 60, NTU-RGB+D 120, NW-UCLA, and UTD-MHAD. The proposed method achieved 93.1% and 97.20% accuracy on the NTU-RGB+D 60 Cross-Subject and Cross-View protocols, 88.15% and 90.20% on the NTU-RGB+D 120 Cross-Subject and Cross-Setup protocols, 98.50% on NW-UCLA, and 97.50% on UTD-MHAD.

Identifiers

Citations and references

Cited by 062 references