INFORMATICA

Informatica

0868-4952 0868-4952

inf18302

10.15388/Informatica.2007.181

Research article

Assessment of Classification Models with Small Amounts of Data

Brumen

Boštjan

bostjan.brumen@uni-mb.si Jurič

Matjaž B.

Welzer

Tatjana

Rozman

Ivan

Jaakkola

Hannu

Papadopoulos

Apostolos

Faculty of Electrical Engineering and Computer Science, University of Maribor, Smetanova 17, Si-2000 Maribor, Slovenia Tampere University of Technology, Pori, PO BOX 300, Fi-28101 Pori, Finland Department of Informatics, Aristotle University, PO BOX 451, Thessaloniki, GR-54124, Greece

01 01 2007

18 3 343 362 01 10 2006

One of the tasks of data mining is classification, which provides a mapping from attributes (observations) to pre-specified classes. Classification models are built by using underlying data. In principle, the models built with more data yield better results. However, the relationship between the available data and the performance is not well understood, except that the accuracy of a classification model has diminishing improvements as a function of data size. In this paper, we present an approach for an early assessment of the extracted knowledge (classification models) in the terms of performance (accuracy), based on the amount of data used. The assessment is based on the observation of the performance on smaller sample sizes. The solution is formally defined and used in an experiment. In experiments we show the correctness and utility of the approach.

Keywords assessment classification accuracy learning curve sampling