Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

jm1

jm1’, the cleaned version by Shepperd et al., described as jm1’ here. jm1’’, the cleaned version by Shepperd et al., described as jm1’’ here. Title/Topic: JM1/software defect prediction Sources: Creators: NASA, then the NASA Metrics Data

👤
CreatorTim Menzies
📅
Published2004-12-02
🔗
DOI10.5281/zenodo.268514
📊
Downloads897
⚖️
Licensecc-by-4.0
File Size1.9 MB
Data TypeDataset
Published2004
Licensecc-by-4.0
Total Views3,010
Total Downloads897
  • jm1’, the cleaned version by Shepperd et al., described as jm1’ here.
  • jm1’’, the cleaned version by Shepperd et al., described as jm1’’ here.

Title/Topic: JM1/software defect prediction

  • Sources:
    • Creators: NASA, then the NASA Metrics Data Program
    • Contacts:
      • Mike Chapman, Galaxy Global Corporation (Robert.Chapman@ivv.nasa.gov) +1-304-367-8341;
      • Pat Callis, NASA, NASA project manager for MDP (Patrick.E.Callis@ivv.nasa.gov) +1-304-367-8309
  • Past usage:
    • How Good is Your Blind Spot Sampling Policy?; 2003; Tim Menzies and Justin S. Di Stefano; 2004 IEEE Conference on High Assurance Software Engineering (http://menzies.us/pdf/03blind.pdf).
      • Results:
        • Very simple learners (ROCKY) perform as well in this domain as more sophisticated methods (e.g. J48, model trees, model trees) for predicting detects
        • Many learners have very low false alarm rates.
        • Probability of detection (PD) rises with effort and rarely rises above it.
        • High PDs are associated with high PFs (probability of failure)
        • PD, PF, effort can change significantly while accuracy remains essentially stable
        • With two notable exceptions, detectors learned from one data set (e.g. KC2) have nearly they same properties when applied to another (e.g. PC2, KC2). Exceptions:
          • LinesOfCode measures generate wider inter-data-set variances;
          • Precision’s inter-data-set variances vary wildly
    • “Assessing Predictors of Software Defects”, T. Menzies and J. DiStefano and A. Orrego and R. Chapman, 2004, Proceedings, workshop on Predictive Software Models, Chicago, Available from http://menzies.us/pdf/04psm.pdf
      • Results:
        • From JM1, Naive Bayes generated PDs of 25% with PF of 20%
        • Naive Bayes out-performs J48 for defect detection
        • When learning on more and more data, little improvement is seen after processing 300 examples.
        • PDs are much higher from data collected below the sub-sub-system level.
        • Accuracy is a surprisingly uninformative measure of success for a defect detector. Two detectors with the same accuracy can have widely varying PDs and PFs.
  • Relevant information:
    • JM1 is written in “C” and is a real-time predictive ground system: Uses simulations to generate predictions
    • Data comes from McCabe and Halstead features extractors of source code. These features were defined in the 70s in an attempt to objectively characterize code features that are associated with software quality. The nature of association is under dispute. Notes on McCabe can be found here and notes on Halstead can be found here. These metrics are widely used for defect

      📤 Share this page

      Found this useful? Share it with your network.

      ✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
jm1 (Full Dataset)1.9 MB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Tim Menzies (2004). jm1. https://doi.org/10.5281/zenodo.268514