Conferences >2016 12th IAPR Workshop on Do...

Interactive Definition and Tuning of One-Class Classifiers for Document Image Classification

Download PDF
Download References
Request Permissions
Save to
Alerts

Abstract:

With mass of data, document image classification systems have to face new trends like being able to process heterogeneous data streams efficiently. Generally, when proces...Show More

Metadata

Abstract:

With mass of data, document image classification systems have to face new trends like being able to process heterogeneous data streams efficiently. Generally, when processing data streams, few knowledge is available about the content of the possible streams. Furthermore, as getting labelled data is costly, the classification model has to be learned from few available labelled examples. To handle such specific context, we think that combining one-class classifiers could be a very interesting alternative to quickly define and tune classification systems dedicated to different document streams. The main interest of one-class classifiers is that no interdependence occurs between each classifier model allowing easy removal, addition or modification of classes of documents. Such reconfiguration will not have any impact on the other classifiers. It is also noticeable that each classifier can use a different set of features compared to the other to handle the same class or even different classes. In return, as only one class is well-specified during the learning step, one-class classifiers have to be defined carefully to obtain good performances. It is more difficult to select the representative training examples and the discriminative features with only positive examples. To overcome these difficulties, we have defined a complete framework offering different methods that can help a system designer to define and tune one-class classifier models. The aims are to make easier the selection of good training examples and of suitable features depending on the class to recognize into the document stream. For that purpose, the proposed methods compute different measures to evaluate the relevance of the available features and training examples. Moreover, a visualization of the decision space according to selected examples and features is proposed to help such a choice and, an automatic tuning is proposed for the parameters of the models according to the class to recognize when a val...

Published in: 2016 12th IAPR Workshop on Document Analysis Systems (DAS)

Date of Conference: 11-14 April 2016

Date Added to IEEE Xplore: 13 June 2016

Electronic ISBN:978-1-5090-1792-8

DOI: 10.1109/DAS.2016.46

Conference Location: Santorini, Greece

References is not available for this document.

Contents

I. Introduction

Since several years companies are interested in document dematerialization process, for different reasons like sharing information, ecological purposes as well as space storage reduction. A typical example of dematerialization procedure is the digitization of administrative documents to facilitate the processing and indexing of received mail streams for example. In this context, this article focuses on facilitating the creation, adaptation and tuning of document image classification (DIC) methods.

Select All

S. S. Khan and M. G. Madden, “One-class classification: taxonomy of study and review of techniques,” The Knowledge Engineering Review, vol. 29, pp. 345–374, 6 2014.

CrossRef Google Scholar

B. Schölkopf, J. C. Platt, J. C. Shawe-Taylor, A. J. Smola, and R. C. Williamson, “Estimating the support of a high-dimensional distribution,” Neural Computation, vol. 13, no. 7, pp. 1443–1471, Jul. 2001.

Interactive Definition and Tuning of One-Class Classifiers for Document Image Classification

Alerts

Abstract:

Metadata

Abstract:

I. Introduction

Authors

Figures

References

Citations

Keywords

Metrics

Footnotes

References

IEEE Account

Purchase Details

Profile Information

Need Help?