Image segmentations produced by BAMF under the AIMI Annotations initiative
The Imaging Data Commons (IDC)(https://imaging.datacommons.cancer.gov/) [1] connects researchers with publicly available cancer imaging data, often linked with other types of cancer data. Many of the collections have limited annotations due to
The Imaging Data Commons (IDC)(https://imaging.datacommons.cancer.gov/) [1] connects researchers with publicly available cancer imaging data, often linked with other types of cancer data. Many of the collections have limited annotations due to the expense and effort required to create these manually. The increased capabilities of AI analysis of radiology images provide an opportunity to augment existing IDC collections with new annotation data. To further this goal, we trained several nnUNet [2] based models for a variety of radiology segmentation tasks from public datasets and used them to generate segmentations for IDC collections.
To validate the model’s performance, roughly 10% of the AI predictions were assigned to a validation set. For this set, a board-certified radiologist graded the quality of AI predictions on a Likert scale. If they did not ‘strongly agree’ with the AI output, the reviewer corrected the segmentation.
This record provides the AI segmentations, Manually corrected segmentations, and Manual scores for the inspected IDC Collection images.
Only 10% of the AI-derived annotations provided in this dataset are verified by expert radiologists . More details, on model training and annotations are provided within the associated manuscript to ensure transparency and reproducibility.
This work was done in two stages. Versions 1.x of this record were from the first stage. Versions 2.x added additional records. In the Version 1.x collections, a medical student (non-expert) reviewed all the AI predictions and rated them on a 5-point Likert Scale, for any AI predictions in the validation set that they did not ‘strongly agree’ with, the non-expert provided corrected segmentations. This non-expert was not utilized for the Version 2.x additional records.
Likert Score Definition:
Guidelines for reviewers to grade the quality of AI segmentations.
- 5 Strongly Agree – Use-as-is (i.e., clinically acceptable, and could be used for treatment without change)
- 4 Agree – Minor edits that are not necessary. Stylistic differences, but not clinically important. The current segmentation is acceptable
- 3 Neither agree nor disagree – Minor edits that are necessary. Minor edits are those that the review judges can be made in less time than starting from scratch or are expected to have minimal effect on treatment outcome
- 2 Disagree – Major edits. This category indicates that the necessary edit is required to ensure correctness, and sufficiently significant that user would prefer to start from the scratch
- 1 Strongly disagree – Unusable. This category indicates that the quality of the automatic annotations is so bad that they are unusable.
Zip File Folder Structure
Each zip file in the collection correlates to a specific segmentation tas
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.