By Topic

Detecting near-replicas on the Web by content and hyperlink analysis

Sign In

Cookies must be enabled to login.After enabling cookies , please use refresh or reload or ctrl+f5 on the browser for the login options.

Formats Non-Member Member
$31 $13
Learn how you can qualify for the best price for this item!
Become an IEEE Member or Subscribe to
IEEE Xplore for exclusive pricing!
close button

puzzle piece

IEEE membership options for an individual and IEEE Xplore subscriptions for an organization offer the most affordable access to essential journal articles, conference papers, standards, eBooks, and eLearning courses.

Learn more about:

IEEE membership

IEEE Xplore subscriptions

5 Author(s)
Di Iorio, E. ; Dipt. di Ingegneria dell''Informazione, Siena Univ., Italy ; Diligenti, M. ; Gori, M. ; Maggini, M.
more authors

The presence of near-replicas of documents is very common on the Web. Documents may be replicated completely or partially for different reasons (versions, mirrors, etc.), or the same resource can be associated to different URLs (dynamically generated pages, etc.). Whilst replication can improve information accessibility by the users, the presence of near-replicated documents can hinder the effectiveness of search engines (for example, decreasing the coverage). We propose a method to detect similar pages, in particular replicas and near-replicas, which is based on a pair of signatures. The first signature is obtained by a random projection of the bag-of-words vector representing the page contents. The second signature is computed by a recursive equation which exploits the connectivity among the Web pages to code the context of each page. The accuracy of the proposed approach is analyzed and validated by experimental results which show that on the given dataset near-replicas can be detected with a precision-recall of 93%.

Published in:

Web Intelligence, 2003. WI 2003. Proceedings. IEEE/WIC International Conference on

Date of Conference:

13-17 Oct. 2003