首页 /研究 /SearchGen
OTHER

SearchGen

Huajing Li, Wang-Chien Lee, Anand Sivasubramaniam, C. Lee Giles

发表年份
2007
引用次数
6

摘要

Due to the popularity of web applications and their heavy usage, it is important to obtain a good understanding of their workloads in order to improve performance of search services. Existing works have typically focused on generic web workloads without putting emphasis on specific domains. In this paper, we analyze the usage logs of CiteSeer, a scientific literature digital library and search engine, to characterize workloads for both robots and users. Essential ingredients that contribute to workloads are proposed. Among them we find the access intervals show high variance, and thus cannot be predicted well with time-series models. On the other hand, client visiting path and semantics can be well captured with probabilistic models and Zipf-law. Based on the findings, we propose SearchGen, a synthetic workload generator to output traces for scientific literature digital libraries and search engines. A comparison between synthetic workloads and actual logged traces suggests that the synthetic workload fits well.

关键词

Computer scienceWorkloadPopularityDigital libraryWeb crawlerProbabilistic logicZipf's lawSearch engineGenerator (circuit theory)Variance (accounting)

相关论文

查看 OTHER 分类全部论文