首页 /研究 /A large-scale study of robots.txt
OTHER

A large-scale study of robots.txt

Yang Sun, Ziming Zhuang, C. Lee Giles

发表年份
2007
引用次数
42

摘要

Search engines largely rely on Web robots to collect information from the Web. Due to the unregulated open-access nature of the Web, robot activities are extremely diverse. Such crawling activities can be regulated from the server side by deploying the Robots Exclusion Protocol in a file called robots.txt. Although it is not an enforcement standard, ethical robots (and many commercial) will follow the rules specified in robots.txt. With our focused crawler, we investigate 7,593 websites from education, government, news, and business domains. Five crawls have been conducted in succession to study the temporal changes. Through statistical analysis of the data, we present a survey of the usage of Web robots rules at the Web scale. The results also show that the usage of robots.txt has increased over time.

关键词

Scale (ratio)RobotComputer scienceArtificial intelligenceGeographyCartography

相关论文

查看 OTHER 分类全部论文