Datasets:

Modalities:
Text
Formats:
parquet
Languages:
German
License:

Request access to GerAV-Reddit

GerAV-Reddit is available exclusively for academic research. Please review the access conditions below before submitting an access request.

By submitting this request, you confirm that you have read and agree to the access conditions described below.

Log in or Sign Up to review the conditions and access this dataset content.

GerAV-Reddit

Dataset Description

GerAV-Reddit is a German authorship verification benchmark belonging to the publication Towards New Heights in German Authorship Verification using Fine-Tuned LLMs on a New Benchmark (Kiefer et al., 2026). The dataset is designed for research on authorship verification and related tasks in computational authorship analysis.

Configurations

The dataset contains four configurations representing different evaluation settings:

  • cross_domain: Authorship verification across different topical domains
  • in_domain: Authorship verification within the same topical domain.
  • profile_based: Profile-based authorship verification concatenating user texts.
  • mixed_source: A mixed-source setting combining different data sources.

All configurations contain train, validation, and test splits.

Please refer to the paper and the corresponding GerAV GitHub repository for further information.

Data Format

Field Type Description
label boolean Whether the two texts were written by the same author
subreddit_a string Subreddit associated with the first text
domain_a string Domain of the first text
subreddit_b string Subreddit associated with the second text
domain_b string Domain of the second text
permalink_a string Permalink of the first text
permalink_b string Permalink of the second text

Intended Use

GerAV-Reddit is intended for academic and non-commercial research on authorship verification, stylometry, NLP, and related areas.

Access Conditions

You need to satisfy the following conditions for your access request to be accepted:

  • The dataset may be used for academic research only. Such use must comply with the applicable Reddit Terms of Use and respect the privacy of Reddit users.
  • Access is restricted to researchers affiliated with a research or educational institution. Please provide an institutional email address with your request.
  • The dataset may not be used for commercial purposes or for mass surveillance.
  • Access requests are reviewed individually. Additional restrictions or exclusions may be applied based on the information provided in the access request.
  • The dataset may not be redistributed or made available to third parties.

By requesting access, you confirm that you have read and agree to these conditions.

Citation

If you use GerAV-Reddit in your research, please cite:

@inproceedings{kiefer-etal-2026-gerav,
    title = "{G}er{AV}: Towards New Heights in {G}erman Authorship Verification using Fine-Tuned {LLM}s on a New Benchmark",
    author = "Kiefer, Lotta  and
      Leiter, Christoph  and
      Takeshita, Sotaro  and
      Schmidt, Elena  and
      Eger, Steffen",
    editor = "Liakata, Maria  and
      Moreira, Viviane P.  and
      Zhang, Jiajun  and
      Jurgens, David",
    booktitle = "Findings of the {A}ssociation for {C}omputational {L}inguistics: {ACL} 2026",
    month = jul,
    year = "2026",
    address = "San Diego, California, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2026.findings-acl.1991/",
    doi = "10.18653/v1/2026.findings-acl.1991",
    pages = "40050--40069",
    ISBN = "979-8-89176-395-1"
}
Downloads last month
-