By Topic

Assessing the effectiveness of feature groups in author recognition tasks with the SOM model

Sign In

Cookies must be enabled to login.After enabling cookies , please use refresh or reload or ctrl+f5 on the browser for the login options.

Formats Non-Member Member
$33 $13
Learn how you can qualify for the best price for this item!
Become an IEEE Member or Subscribe to
IEEE Xplore for exclusive pricing!
close button

puzzle piece

IEEE membership options for an individual and IEEE Xplore subscriptions for an organization offer the most affordable access to essential journal articles, conference papers, standards, eBooks, and eLearning courses.

Learn more about:

IEEE membership

IEEE Xplore subscriptions

1 Author(s)
G. Tambouratzis ; Inst. for Language & Speech Process., Athens, Greece

The present paper focuses on studying the effectiveness of the self-organizing map (SOM) when applied to the task of categorizing a corpus of texts according to the style of their authors. This task is of particular importance for information retrieval applications using very large databases of documents. The emphasis of this article is to determine the extent to which the SOM possesses the ability to analyze such data, successfully uncovering the stylistic differences among authors in an unsupervised manner. To that end, a variety of feature vectors are studied, each of which either 1) comprises a single category of linguistic features or 2) spans several different categories of linguistic features, in order to determine the effectiveness of each feature category. It is shown that the highest accuracy is achieved when using a vector covering multiple linguistic categories. A comparison of the results obtained to the results of statistical methods indicates the ability of the SOM network to reveal the clustering potential of isolated parameter groups and its effectiveness in handling efficiently high-dimensional data vectors. Potential extensions to related text-organization techniques, such as the WEBSOM, thus become evident.

Published in:

IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews)  (Volume:36 ,  Issue: 2 )