Abstract

ID3 is most successful when used with sets of training and testing data that contain no missing attribute values. Many times, however, real-world domains have attributes with missing values. Sometimes these attribute values may not be needed to classify an instance. Such attribute values are called don't-care attribute values. In other cases, the values are needed but are unavailable. These values are called unknown attribute values. This paper describes the difference between unknown and don't-care attribute values and discusses several ways of identifying don't-care attribute values in ID3. Numerical results are described which validate the practicality of these approaches.

Department(s)

Computer Science

Second Department

Mathematics and Statistics

Comments

The first Author is a Graduate Student

This report is substantially the M.S. thesis of the first author, completed May, 1992.

This thesis has been prepared in the style utilized by the Association for Computing Machinery (ACM). Pages 1-52 will be presented for publication in the journal Communications of the ACM. Appendices A, B, and C have been added for purposes normal to thesis writing.

Keywords and Phrases

Automated Induction, Machine Learning, Knowledge Representation

Report Number

CSc-92-09

Document Type

Technical Report

Document Version

Final Version

File Type

text

Language(s)

English

Rights

© 1992 University of Missouri - Rolla, All rights reserved

Publication Date

1 May, 1992

Share

 
COinS