Home /Research /Dominos: A New Web Crawler's Design
OTHER

Dominos: A New Web Crawler's Design

Chabane Djeraba

Year
2004
Citations
16

Abstract

Today’s search engines are equipped with specialized agents known as Web crawlers (download robots) dedicated to crawling large Web contents on line. These contents are then analyzed, indexed and made available to users. Crawlers interact with thousands of Web servers over periods extending from a few weeks to several years. This type of crawling process therefore means that certain judicious criteria need to be taken into account, such as the robustness, exibilit y and maintainability of these crawlers. In the present paper, we will describe the design and implementation of a realtime distributed system of Web crawling running on a cluster of machines. The system crawls several thousands of pages every second, includes a high-performance fault manager, is platform independent and is able to adapt transparently to a wide range of congurations without incurring additional hardware expenditure. We will then provide details of the system architecture and describe the technical choices for very high performance crawling. Finally, we will discuss the experimental results obtained, comparing them with other documented systems.

Keywords

Web crawlerCrawlingComputer scienceWorld Wide WebWeb serverWeb pageMaintainabilityFocused crawlerServerTroubleshooting

Related papers

Browse all OTHER papers