Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

Annotated dataset of Russian -nie nominalisations

Annotated dataset of Russian -nie nominalisations Daria Seres (University of Graz) Marko Simonović (University of Graz) Predrag Kovačević (University of  Novi Sad) Goal and rationale This dataset was created as a modern Russian comparison dataset for th

👤
CreatorSeres, Daria
📅
Published2026-05-22
🔗
DOI10.5281/zenodo.20344473
📊
Downloads54
⚖️
Licensecc-by-4.0
File Size572.8 KB
Data TypeDataset
Published2026
Licensecc-by-4.0
Total Views38
Total Downloads54

Annotated dataset of Russian -nie nominalisations

Daria Seres (University of Graz)

Marko Simonović (University of Graz)

Predrag Kovačević (University of  Novi Sad)

Goal and rationale

This dataset was created as a modern Russian comparison dataset for the analysis of -nie nominalisations in Simonović, Kovačević and Milićev (accepted). Its purpose is to provide comparable modern Russian evidence against which the Slavonic-Serbian -nie nominalisation data can be contrasted.

Dataset description

Each row represents one attested token of a modern Russian -nie nominalisation drawn from the Russian National Corpus (RNC; ruscorpora.ru). The data were extracted from RNC texts classified as учебно-научная ‘academic/educational-scientific’ and публицистика ‘journalistic writing’ and restricted to texts produced after 2000. The sample was randomly selected and then manually checked and annotated.

The sample contains:

  • 1,705 tokens

  • 435 lemmas

The data were extracted from the Russian National Corpus using the following lemma query:

[lemma=”.*(т|н)ие”]

This query targets lemmas ending in -ние or -тие. Since the query targets formal lemma endings, it can retrieve false positives. In this dataset, such items were not removed, but their status is explicitly marked in Departicipial and by blank or uncertain values in the subsequent annotation columns.

In what follows, we describe each column. The columns Departicipial, Compound stem, and Perfective base contain manual linguistic annotation; the remaining columns contain identifiers, concordance context, or metadata from the Russian National Corpus.

In what follows, we describe each column. The columns Departicipial, Compound stem, and Perfective base contain manual linguistic annotation; the remaining columns contain identifiers, concordance context, or metadata from the Russian National Corpus.

Column A — ID

ID assigns an arbitrary unique number to each example in the dataset.

Column B — Citation form

Contains the citation form of the nominalisation, for example разрешение ‘permission’, клонирование ‘cloning’, исследование ‘research’, образование ‘education’, исполнение ‘execution; fulfilment’, and создание ‘creation’.

Due to Russian inflectional morphology, the citation form may differ from the attested surface form in the Center column. For example, the citation form клонирование ‘cloning’ corresponds to attested forms such as клонирования, and разрешение ‘permission’ may correspond to разрешение or разрешения depending on case and number.

Column C — Departicipial

Marks whether the item is analysed as belonging to the target deverbal/departicipial -nie nominalisation class.</

📤 Share this page

Found this useful? Share it with your network.

✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
Annotated dataset of Russian -nie nominalisations (Full Dataset)572.8 KB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Seres, Daria (2026). Annotated dataset of Russian -nie nominalisations. https://doi.org/10.5281/zenodo.20344473