Abstract:
This paper describes an experiment in applying a standard supervised machine learning algorithm (C4.5) to the problem of developing subject classification rules for docum...Show MoreMetadata
Abstract:
This paper describes an experiment in applying a standard supervised machine learning algorithm (C4.5) to the problem of developing subject classification rules for documents. This algorithm is found to produce surprisingly concise models of document classifications. While the models are highly accurate on the training sets, evaluation over test sets or through cross-validation shows a significant decrease in classification accuracy. Given the difficult nature of the experimental task, however, the results of this investigation are promising and merit further study. An additional algorithm, 1R, is shown to be highly effective in generating lists of candidate terms for subject descriptions.
Published in: Proceedings 1995 Second New Zealand International Two-Stream Conference on Artificial Neural Networks and Expert Systems
Date of Conference: 20-23 November 1995
Date Added to IEEE Xplore: 06 August 2002
Print ISBN:0-8186-7174-2