Skip to main navigation Skip to search Skip to main content

Preliminary approach on synthetic data sets generation based on class separability measure

  • Núria Macià*
  • , Ester Bernadó-Mansilla
  • , Albert Orriols-Puig
  • *Corresponding author for this work

    Research output: Book chapterConference contributionpeer-review

    21 Citations (Scopus)

    Abstract

    Usually, performance of classifiers is evaluated on real-world problems that mainly belong to public repositories. However, we ignore the inherent properties of these data and how they affect classifier behavior. Also, the high cost or the difficulty of experiments hinder the data collection, leading to complex data sets characterized by few instances, missing values, and imprecise data. The generation of synthetic data sets solves both issues and allows us to build problems with a minor cost and whose characteristics are predefined. This is useful to test system limitations in a controlled frame-work. This paper proposes to generate synthetic data sets based on data complexity. We rely on the length of the class boundary to build the data sets, obtaining a preliminary set of benchmarks to assess classifier accuracy. The study can be further matured to identify regions of competence for classifiers.

    Original languageEnglish
    Title of host publication2008 19th International Conference on Pattern Recognition, ICPR 2008
    PublisherInstitute of Electrical and Electronics Engineers Inc.
    ISBN (Print)9781424421756
    DOIs
    Publication statusPublished - 2008

    Publication series

    NameProceedings - International Conference on Pattern Recognition
    ISSN (Print)1051-4651

    Fingerprint

    Dive into the research topics of 'Preliminary approach on synthetic data sets generation based on class separability measure'. Together they form a unique fingerprint.

    Cite this