Datasets:
Request access to GerAV-Reddit
GerAV-Reddit is available exclusively for academic research. Please review the access conditions below before submitting an access request.
By submitting this request, you confirm that you have read and agree to the access conditions described below.
Log in or Sign Up to review the conditions and access this dataset content.
GerAV-Reddit
Dataset Description
GerAV-Reddit is a German authorship verification benchmark belonging to the publication Towards New Heights in German Authorship Verification using Fine-Tuned LLMs on a New Benchmark (Kiefer et al., 2026). The dataset is designed for research on authorship verification and related tasks in computational authorship analysis.
Configurations
The dataset contains four configurations representing different evaluation settings:
- cross_domain: Authorship verification across different topical domains
- in_domain: Authorship verification within the same topical domain.
- profile_based: Profile-based authorship verification concatenating user texts.
- mixed_source: A mixed-source setting combining different data sources.
All configurations contain train, validation, and test splits.
Please refer to the paper and the corresponding GerAV GitHub repository for further information.
Data Format
| Field | Type | Description |
|---|---|---|
label |
boolean | Whether the two texts were written by the same author |
subreddit_a |
string | Subreddit associated with the first text |
domain_a |
string | Domain of the first text |
subreddit_b |
string | Subreddit associated with the second text |
domain_b |
string | Domain of the second text |
permalink_a |
string | Permalink of the first text |
permalink_b |
string | Permalink of the second text |
Intended Use
GerAV-Reddit is intended for academic and non-commercial research on authorship verification, stylometry, NLP, and related areas.
Access Conditions
You need to satisfy the following conditions for your access request to be accepted:
- The dataset may be used for academic research only. Such use must comply with the applicable Reddit Terms of Use and respect the privacy of Reddit users.
- Access is restricted to researchers affiliated with a research or educational institution. Please provide an institutional email address with your request.
- The dataset may not be used for commercial purposes or for mass surveillance.
- Access requests are reviewed individually. Additional restrictions or exclusions may be applied based on the information provided in the access request.
- The dataset may not be redistributed or made available to third parties.
By requesting access, you confirm that you have read and agree to these conditions.
Citation
If you use GerAV-Reddit in your research, please cite:
@inproceedings{kiefer-etal-2026-gerav,
title = "{G}er{AV}: Towards New Heights in {G}erman Authorship Verification using Fine-Tuned {LLM}s on a New Benchmark",
author = "Kiefer, Lotta and
Leiter, Christoph and
Takeshita, Sotaro and
Schmidt, Elena and
Eger, Steffen",
editor = "Liakata, Maria and
Moreira, Viviane P. and
Zhang, Jiajun and
Jurgens, David",
booktitle = "Findings of the {A}ssociation for {C}omputational {L}inguistics: {ACL} 2026",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.findings-acl.1991/",
doi = "10.18653/v1/2026.findings-acl.1991",
pages = "40050--40069",
ISBN = "979-8-89176-395-1"
}
- Downloads last month
- -