Vision Language Models for Video Captioning, Video Retrieval, and Video Question Answering
Jamalova GulchexraDepartment of Information Systems and Technologies, Karshi State Technical University, Karshi, UzbekistanBegimov OktamIsmatulla XayrullayevDepartment of Information Technology and Exact Sciences, Termez University of Economics and Service, Termez, UzbekistanAbdullayev Abdulla Fayzulla UgliDepartment of Specialization, Social Humanities, and Exact Sciences, Tashkent State University of Economics, Tashkent, UzbekistanAnorgul AshirovaDepartment of General Professional Sciences, Mamun University, Khiva, Uzbekistan
2026
ABI
Annotatsiya
Vision-language models have significantly improved intelligent video understanding by integrating visual and textual information for tasks such as video captioning, video retrieval, and question answering. These models learn rich multimodal representations that enhance semantic understanding, enabling accurate content description, efficient search, and natural language interaction with video data. This study reviews recent advancements, applications, and challenges of vision-language models in developing intelligent and scalable video analysis systems.
Hali tarjima qilinmagan
Identifikatorlar
Iqtiboslar va manbalar
0 ta iqtibos0 ta foydalanilgan manba