Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

RiceMind Dataset

Abstract This repository contains the core datasets supporting the RiceMind platform, a comprehensive, structurally standardized, and fine-grained knowledge dataset for rice (Oryza sativa). To address the explosive growth of rice-related scientific liter

👤
CreatorLiu, Yawen
📅
Published2026-04-15
🔗
DOI10.5281/zenodo.19702501
📊
Downloads100
⚖️
Licensecc-by-4.0
File Size615.2 MB
Data TypeDataset
Published2026
Licensecc-by-4.0
Total Views159
Total Downloads100

Abstract This repository contains the core datasets supporting the RiceMind platform, a comprehensive, structurally standardized, and fine-grained knowledge dataset for rice (Oryza sativa). To address the explosive growth of rice-related scientific literature, we employed automated, large-scale Natural Language Processing (NLP) strategies to mine experimentally validated associations from abstracts and full-text articles (PubMed/PMC).

This dataset not only provides extensive text-mined evidence for Gene-Trait Associations (GTAs), Gene-Variety Associations (GVAs), and Trait-Variety Associations (TVAs), but also seamlessly integrates these findings with multi-dimensional omics data and standardized ontologies. It serves as a robust resource for researchers in plant genomics, molecular breeding, and bioinformatics.

Data Standardization & Identifier Mapping
To ensure the highest level of structural consistency and cross-database interoperability, rigorous standardization pipelines were applied:

  • Phenotypic Trait Standardization: All extracted trait descriptions have been systematically mapped and unified to recognized semantic ontologies, including the Gene Ontology (GO), Plant Trait Ontology (TO), Plant Ontology (PO), and the Rice Trait Ontology (RTO). The dataset centers around Oryza sativa.
  • Gene Nomenclature Standardization: Gene entities sourced from external databases including Oryzabase, RAP-DB, Ensembl Plants, and Planteome were standardized to the unified RAP ID system. For instance, genes from Oryzabase were anchored via their annotated RAP IDs, and data from Planteome were mapped using a “Protein-Gene-RAP ID” trajectory to ensure nomenclature consistency.

Data Records Overview The repository includes the following files:

1. Text Corpus & Evidence Data

  • keyword_filtered_rice_sentences.jsonl: Raw, keyword-filtered sentence segments extracted from full-text scientific literature. This serves as the foundational text corpus for downstream NLP processing.

  • rice_context_sentences_compressed.tsv: Compressed contextual sentences providing the exact narrative evidence and semantic context from which the associations were mined.

2. NLP-Extracted Association Databases

  • NLP_Rice_GTA_Database.tsv: The core text-mined Gene-Trait Association (GTA) dataset. It contains fine-grained, explicitly

    📤 Share this page

    Found this useful? Share it with your network.

    ✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
RiceMind Dataset (Full Dataset)615.2 MB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Liu, Yawen (2026). RiceMind Dataset. https://doi.org/10.5281/zenodo.19702501