Skip to Main Content
The research presented in this paper focuses on the pre-processing stage of the clustering process, proposing a novel indexing technique which goes beyond the syntax of terms; trying to capture their unambiguous meaning from their context and to derive a set of concepts to be used to represent the documents. This approach overcomes some of the major drawbacks deriving from the use of bag of words and term frequency based indexing techniques. The proposed approach is evaluated by using unsupervised performance measures and by comparing the clustering results achieved against the ones obtained when using a traditional indexing method. The experimental results show that better clustering results are achieved through the use of the proposed indexing approach, which also led to a substantial reduction of the index term dimension.