Мақола

Applying Web Crawler Technologies for Compiling Parallel Corpora as one Stage of Natural Language Processing

Nilufar AbdurakhmonovaNational University of Uzbekistan,Uzbek linguistics department,Tashkent,UzbekistanIsmailov AlisherAAndijan machine building institute,Innovative Educational department,AAndijan,UzbekistanGuli ToirovaBukhara State University,Uzbek linguistics department,Bukhara,Uzbekistan

2022 7th International Conference on Computer Science and Engineering (UBMK)conference2022en

ABI

Аннотация

over the past decade, the amount of information on the internet has increased. A large amount of unstructured data, referred to as big data on the web, has been created. Finding and extracting data on the internet is called information retrieval. In the search for information, there are web crawler tools, which are a program that scans information on the internet and downloads web documents automatically. Search robot applications can be used in various fields, such as news, finance, medicine, etc. In this article, we will discuss the basic principle and characteristics of search engines as an example to build parallel corpora, as well as the classification of modern popular crawlers, strategies and current applications of crawlers. Finally, we will end this article with a discussion of future directions for research on crawlers.

Ҳали таржима қилинмаган

Мавзулар

Web Data Mining and Analysis

Идентификаторлар

DOI: 10.1109/ubmk55850.2022.9919521

Иқтибослар ва манбалар

6 та иқтибос0 та фойдаланилган манба

Кўрсаткичлар — AkademScholar