MeSpEn_Parallel-Corpora
MeSpEn consists of a resource of heterogeneous health related documents in Spanish and English useful to build parallel corpora for training and evaluating Spanish <-> English medical machine translation systems, to generate multilingual automatic term extraction tools, and develop other Sp
MeSpEn consists of a resource of heterogeneous health related documents in Spanish and English useful to build parallel corpora for training and evaluating Spanish <-> English medical machine translation systems, to generate multilingual automatic term extraction tools, and develop other Spanish medical NLP components. MeSpEn provides the combination and harmonization of various bibliographic datasets of biomedical and clinical literature from Spain and Latin America or web-content with trusted information sources about diseases, conditions, and wellness issues for patients.
MeSpEn was used to generate automatically bilingual health related-glossaries through automatic term detection and named entity recognition in English and target candidate term extraction in Spanish through sentence alignment approaches, implying potentially the generation of Silver Standard annotated health texts in Spanish.
MeSpEn was used to generate automatically bilingual health related-glossaries through automatic term detection and named entity recognition in English and target candidate term extraction in Spanish through sentence alignment approaches, implying potentially the generation of Silver Standard annotated health texts in Spanish (see Villegas, et al. "The MeSpEN resource for English-Spanish medical machine translation and terminologies: census of parallel corpora, glossaries and term translations." Proc. LREC 2018 Workshop MultilingualBIO: Multilingual Biomedical Text Processing).
The MeSpEn resource aggregates several datasets, mainly from 4 principal sources: IBECS, SciELO, Pubmed and MedlinePlus:
IBECS (Spanish Bibliographical Index in Health Sciences) is a bibliographical database that collects scientific journals covering multiple fields in health sciences. It is maintained by the Spanish National Health Sciences Library (BNCS), at the Carlos III Health Institute.
This corpus contains titles and abstracts from 168,198 records in English and Spanish. Users can find the metadata of each record written in Dublin Core format. The original XML file of the record provided by IBECS is provided as well.
For more information about IBECS parallel corpora, see IBECS_README file.
SciELO (Scientific Electronic Library Online) gathers electronic publications of complete full text articles from scientific journals of Latin America, South Africa and Spain. Currently is present in 15 countries and supported by the Sao Paulo Research Foundation (FAPESP) and the Brazilian National Council for Scientific and Technological Development (BI
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.