A gleaning subsystem for CINDI
Tong Zhang
- 发表年份
- 2004
- 引用次数
- 2
- 访问权限
- 开放获取
摘要
Internet search engines typically use Internet crawlers, or robots, for the purpose of constructing and maintaining a searchable index of resources on the Web. Topic-specific robots will become popular in the next generation. They gather information on the Internet in specific domains by means of information filtering technology. The CINDI Robot System is such an application in academic domain. This research is concerned with a structure-based gleaning subsystem for CINDI. The system separates theses, technical reports, academic papers, and FAQs as resources while e-mails, letters, resumes, graphics, and discussion groups are considered as chaff. This system makes decisions based on weight, which is carefully assigned to each resource by matching its structure with predefined Document Type Definitions (DTDs). The DTDs for the typical structure for the specific document types are built based on some predefined profiles. The system also features conversion subsystem in Windows environment to unify document formats for CINDI. (Abstract shortened by UMI.)
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991