Asosiy kontentga oʻtish
Bob

Transformer-Based Models for Video Recognition and Scene Understanding

Ergashev NuriddinDepartment of Information Systems and Technologies, Karshi State Technical University, UzbekistanXurramov RuslanDepartment of Information Technology and Exact Sciences, Termez University of Economics and Service, Termez, UzbekistanMavlanov NormuminDepartment of Scientific Research and Innovation, Tashkent State University of Economics, Tashkent, UzbekistanAnorgul AshirovaDepartment of General Professional Sciences, Mamun University, Khiva, UzbekistanKhayrulla UrozboevDepartment of International Scientific Journals and Ratings, Alfraganus University, Tashkent, Uzbekistan
2026
ABI

Annotatsiya

Transformer-based models have significantly advanced video recognition and scene understanding by effectively capturing long-range spatial and temporal dependencies in video data. Unlike traditional convolutional approaches, Transformers leverage self-attention mechanisms to model complex relationships across frames, improving performance in tasks such as action recognition, event detection, and video classification. This work reviews the principles, architectures, applications, and challenges of Transformer-based video models, highlighting recent developments and future research directions.

Hali tarjima qilinmagan

Identifikatorlar

Iqtiboslar va manbalar

0 ta iqtibos0 ta foydalanilgan manba