AN APPROACH TO BUILD A WEB CRAWLER USING CLUSTERING BASEDK-MEANS ALGORITHM

[摘要] Central to any data-mining project is having sufficient amounts of data that can be processed to provide meaningful and statistically relevant information. But getting the unstructured data is only the initial stage and that data must be transformed into a structured format which is suitable for further processing. In this paper we have proposed architecture for the web-crawling and arrange their unstructured data using cluster based algorithm. . The clustering process is based on the k-means algorithm. This paper is completely based on the focused crawler mechanism that only scans the pages by using general crawling policies.

[发布日期] [发布机构]

[效力级别] [学科分类]

[关键词] web-mining;clustering;data-mining-means;focused crawler [时效性]

浏览次数：1

统一登录查看全文激活码登录查看全文