Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

Identifying Machine-Paraphrased Plagiarism

README.txt Title: Identifying Machine-Paraphrased Plagiarism Authors: Jan Philip Wahle, Terry Ruas, Tomas Foltynek, Norman Meuschke, and  Bela Gipp contact email: wahle@gipplab.org; ruas@gipplab.org; Venue: iConference Year: 2022 =========================

👤
CreatorWahle, Jan Philip
📅
Published2021-01-14
🔗
DOI10.5281/zenodo.3608000
📊
Downloads1,290
⚖️
Licensecc-by-4.0
File Size5.6 GB
Data TypeDataset
Published2021
Licensecc-by-4.0
Total Views6,275
Total Downloads1,290

README.txt

Title: Identifying Machine-Paraphrased Plagiarism
Authors: Jan Philip Wahle, Terry Ruas, Tomas Foltynek, Norman Meuschke, and  Bela Gipp
contact email: wahle@gipplab.org; ruas@gipplab.org;
Venue: iConference
Year: 2022
================================================================
Dataset Description:

Training:
200,767 paragraphs (98,282 original, 102,485paraphrased) extracted from 8,024 Wikipedia (English) articles (4,012 original, 4,012 paraphrased using the SpinBot API).

Testing:
SpinBot: 
    arXiv         – Original – 20,966;    Spun – 20,867
    Theses        – Original – 5,226;        Spun – 3,463
    Wikipedia    – Original – 39,241;    Spun – 40,729
    
SpinnerChief-4W: 
    arXiv         – Original – 20,966;    Spun – 21,671
    Theses        – Original – 2,379;        Spun – 2,941
    Wikipedia    – Original – 39,241;    Spun – 39,618
    
SpinnerChief-2W: 
    arXiv         – Original – 20,966;    Spun – 21,719
    Theses        – Original – 2,379;        Spun – 2,941
    Wikipedia    – Original – 39,241;    Spun – 39,697

================================================================
Dataset Structure:

[human_evaluation] folder: human evaluation to identify human-generated text and machine-paraphrased text. It contains the files (original and spun) as for the answer-key for the survey performed with human subjects (all data is anonymous for privacy reasons).

NNNNN.txt – whole document from which an extract was taken for human evaluation
    key.txt.zip – information about each case (ORIG/SPUN)
    results.xlsx – raw results downloaded from the survey tool (the extracts which humans judged are in the first line)
    results-corrected.xlsx – at the very beginning, there was a mistake in one question (wrong extract). These results were excluded.

[automated_evaluation]: contains all files used for the automated evaluation considering [spinbot] (https://spinbot.com/API) and [spinnerchief] (http://developer.spinnerchief.com/API_Document.aspx).

  • Each paraphrase tool folder contains:
  • [corpus]

    📤 Share this page

    Found this useful? Share it with your network.

    ✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
Identifying Machine-Paraphrased Plagiarism (Full Dataset)5.6 GB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Wahle, Jan Philip (2021). Identifying Machine-Paraphrased Plagiarism. https://doi.org/10.5281/zenodo.3608000