Dominos: A New Web Crawler's Design
Chabane Djeraba
- 发表年份
- 2004
- 引用次数
- 16
摘要
Today’s search engines are equipped with specialized agents known as Web crawlers (download robots) dedicated to crawling large Web contents on line. These contents are then analyzed, indexed and made available to users. Crawlers interact with thousands of Web servers over periods extending from a few weeks to several years. This type of crawling process therefore means that certain judicious criteria need to be taken into account, such as the robustness, exibilit y and maintainability of these crawlers. In the present paper, we will describe the design and implementation of a realtime distributed system of Web crawling running on a cluster of machines. The system crawls several thousands of pages every second, includes a high-performance fault manager, is platform independent and is able to adapt transparently to a wide range of congurations without incurring additional hardware expenditure. We will then provide details of the system architecture and describe the technical choices for very high performance crawling. Finally, we will discuss the experimental results obtained, comparing them with other documented systems.
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991