Skip to content

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

Fediverse Deepfake-Text Corpus

This repository contains a large-scale deepfake-text corpus from the Fediverse. The corpus contains human-written text (HWT, label 0) collected from pre-LLM Fediverse snapshots and AI-generated text (AIGT, label 1) generated based on that HWT corpus.

Data Collection

The retrieval run was conducted from approximately October 29, 2025, 14:00 AoE to October 30, 2025, 09:00 AoE using five servers located in Wisconsin, United States. Each server was equipped with two Intel® Xeon® Silver 4114 CPUs (20 cores / 40 threads in total), 188 GB of RAM, and a 1 TB HDD, and ran Ubuntu 20.04.6 LTS. These dates describe when the snapshots were retrieved, not when the source posts were created.

To collect historical HWT posts, we manually identified publicly available Mastodon statuses from around November 1 of 2017, 2018, and 2019 and used their status IDs as cursor anchors (max_id) for FediLive's public-timeline collection. The anchors were used to initialize backward pagination near the end of each target period; the resulting snapshots cover the corresponding October 18–November 1 windows.

Corpus files

The corpus is provided as nine standalone JSON Lines shards:

Files Records
corpus/corpus_01.jsonl through corpus/corpus_09.jsonl 1,003,993

Each line represents one post object. Each shard is below 45 MB and ends at a JSONL record boundary. The corpus contains 523,048 HWT records and 480,945 AIGT records across 263 Mastodon communities.

Record schema

Each line is one JSON object. The public corpus uses these fields:

Field Type Description
text string Cleaned post text.
label integer 0 for HWT and 1 for AIGT.
community_id string Mastodon instance/community domain.
is_reply boolean Whether the post is a reply.
language string Language metadata released by the collection source, or unknown.
post_id string or null Post identifier for HWT records (hash value).
original_id string or null HWT source identifier associated with AIGT records (hash value).
generation_method string or null AIGT generation strategy.
llm_model string or null Target LLM selected for generation.

Note

  • The public post_id and original_id values are lowercase SHA-256 hex digests of the original identifier for user privacy and anonymity. The original_id to post_id linkage between AIGT and HWT records is preserved.

  • The language field is collection-source metadata, not a newly inferred language label. The release has 97 raw values, including unknown and source-specific variants. unknown occurs in 591,963 records (58.96%). After excluding unknown and normalizing documented variants such as ja-IM to ja, there are 89 language categories. For downstream language-aware analysis, we suggest using a language identification model such as FastText instead of directly utilizing the source metadata.

  • The shards preserve records without text deduplication. Across exact text values, 9,140 distinct strings occur more than once; together, these repeats account for 19,187 rows beyond the first occurrence of each string.

AIGT generation

The generation pipeline uses seven representative LLM families: GPT-4-Turbo, GPT-4o-mini, Claude-3.5-Sonnet, Claude-Sonnet-4, Gemini-2.0-Flash, Qwen-2.5-32B-Instruct, and LLaMA3-8B-Instruct. The recorded strategies are polish, complete, and 3-iteration paraphrase.

Research Note

Exact-copy generated text

There are 295 AIGT records whose text exactly matches their linked HWT source: 213 from 3-iteration paraphrase, 79 from polish, and 3 from complete. These 295 records represent 0.061% of the 480,945 AIGT records. Labels reflect record provenance: 0 identifies collected HWT, while 1 identifies generation-pipeline outputs, including outputs that preserve the source text unchanged. These outputs are retained in the archival corpus with their generation-provenance labels.

More broadly, 288 distinct text strings occur under both labels, covering 1,003 records. These conflicting labels cannot be distinguished from text alone. For downstream usage, we recommend excluding all records with these cross-label text values; the partitioning script does so by default.

Pair-aware splits

The corpus release does not include a train/test split. To reduce leakage in a custom split, keep records in the same group when they share an HWT source (post_id/original_id) or an exact text value, then assign whole groups to partitions. This keeps linked generations and identical strings together. Record the split seed and grouping policy when publishing results.

Below, we provide an example partitioning.

Suggested Partitioning

Run scripts/create_splits.py from the corpus repository root to create splits from the nine shards. Linked HWT/AIGT records and rows with identical text stay in the same partition. Exact text found under both labels is excluded by default; pass --keep-cross-label-exact-text to retain those rows.

Create train/validation files with a 20% validation fraction:

python scripts/create_splits.py \
  --output-dir splits \
  --validation-fraction 0.2 \
  --seed 42

Create train/validation/test files with 10% validation and 10% test fractions:

python scripts/create_splits.py \
  --output-dir splits \
  --validation-fraction 0.1 \
  --test-fraction 0.1 \
  --seed 42

The output contains train.jsonl, validation.jsonl, optional test.jsonl, and split_manifest.json. Source shards are unchanged. The script uses only the Python standard library; no additional packages are required.

Requested split fractions are approximate because linked records and identical text remain grouped. The manifest records the actual record and label counts for each partition.

Use a separate output directory for each split configuration. Existing split outputs are rejected by default; pass --overwrite to replace them. When no test partition is requested, --overwrite also removes any previous test.jsonl. Other files in the output directory are left unchanged.

Instance metadata

The instance_metadata/instance_metadata_*.jsonl files contain one JSON object per line with crawl results for Mastodon instances. These files provide community context and are separate from the post JSONL shards.

instance_metadata/metadata_errors.csv records crawl failures. Because ActivityPub instances can freely connect or disconnect, an unavailable instance may become reachable on a later crawl; retries must follow applicable platform policies.

Instance-level statistics and distribution

Top-10 Instances Statistics

Instance Posts HWT AIGT HWT % AIGT %
mastodon.social 198,827 109,075 89,752 54.86 45.14
pawoo.net 153,665 77,780 75,885 50.62 49.38
mstdn.maud.io 55,918 28,306 27,612 50.62 49.38
chaosphere.hostdon.jp 48,292 24,172 24,120 50.05 49.95
imastodon.net 44,370 22,305 22,065 50.27 49.73
mamot.fr 38,991 20,513 18,478 52.61 47.39
rewa.mobi 29,088 15,774 13,314 54.23 45.77
eletusk.club 24,218 12,223 11,995 50.47 49.53
mstdn.guru 23,338 11,725 11,613 50.24 49.76
pokemon.mastportal.info 16,085 8,160 7,925 50.73 49.27

Post-volume distribution across instances

Post-volume distribution across instances

Figure: Distribution of post volumes across Mastodon instances. The x-axis is shown on a logarithmic scale.

Example

The examples below illustrate how community context and paired HWT/AIGT text appear across representative instances. Ellipses indicate shortened examples.

Instance Description and community context HWT examples AIGT examples
mastodon.social
109,075 posts
The original server operated by the Mastodon gGmbH non-profit.

Rules: No misinformation, no harassment, no violence incitement...
Annabel found some Halloween-ready lighting.

I've also applied to put up a shop on Designed by Humans, since I think...

okay it's food truck o'clock I guess...
Yo, peep the Halloween swag Annabel just found!

I've also applied to open a shop on Designed by Humans. Their shipping rates...

so I better eat something quick before the show starts...
pawoo.net
77,780 posts
Pawoo, a Mastodon instance operated by The Social Coop Limited...

Rules: No rules.
i could listen to the 100 poets i replay for days

[Update] IconTweak 1.0.1 - Show App Versions in Menu

RT: Oh, I just love #Japan! #Tokyo #subway
I could listen to those 100 poets over and over again, replaying for days!

[Update] IconTweak 1.0.1 is here! Now you can see app versions directly in the menu. Super handy, right?

RT: Oh, I just love #Japan! The culture, the food, and the scenery...
mstdn.maud.io
28,306 posts
The place to express you more freely.

Rules: Comply with law, no disruption, check updates regularly...
The Pirate Bay was recently down for over a week due to a DDoS attack

What is CMAF? Threat or Opportunity?

GitHub - neuecc/Utf8Json: Definitely Fastest and Zero Allocation JSON Serializer...
The Pirate Bay was offline for more than a week recently because of a DDoS attack.

Is CMAF a chance or a risk?

Definitely a game-changer if you're looking for performance. The speed is unreal, and zero allocations mean...
mastodon.art
5,397 posts
Your friendly home on the Fediverse for all things creative.

Rules: No AI or NFTs, credit + commentary required, respect user boundaries...
Once upon a time I started this. I think I'll continue it, now. #MastoArt

Compliment my costume and I'll give you more candy

second october patreon reward teaser #digitalart...
Been sitting on this project for a while now...

Compliment my costume, and I'll give you some extra candy!

second october patreon reward teaser #digitalsketch...

Verification

The release includes a standalone, standard-library-only verification script that recomputes the corpus claims from the public files. From the repository root, run:

python scripts/verify_statistics.py

The same command also works from inside scripts/ as python verify_statistics.py. Expected values are recorded in expected/corpus_statistics.txt, which contains the manuscript-oriented claims checked by the script. The environment is specified by requirements.txt and environment.yml; no third-party Python packages are required.

Ethics Considerations

Because this study involves the large-scale collection, analysis, and release of publicly available social-media data, we considered potential privacy and research-ethics risks throughout the study. The human-written Fediverse data used in this study were collected using the open-source FediLive framework. The collection accessed only publicly shared content, respected platform rate limits and the Robots Exclusion Protocol (robots.txt), and did not access private, restricted, or otherwise non-public content.

Privacy risks were further minimized in the subsequent use and release of the data. Our analyses operate at the corpus and community levels rather than at the level of individual users: FediAID uses aggregate community-level representations instead of individualized retrieval, user histories, or user-level profiling. Consistent with this design, the released corpus excludes user-level information and account identifiers, and focuses on textual content together with community-level labels and metadata. Because the corpus retains publicly posted text, some posts may nevertheless be traceable to their original public sources through text search; we therefore do not claim that the released text is fully anonymous, but avoid providing user-level identifiers or mappings that facilitate such linkage.

The protocol governing the collection, research use, and release of the publicly available Fediverse data underwent institutional ethics review by our institution's Institutional Review Board (IRB). In consultation with institutional ethics staff, the study was considered as involving “no more than minimal risk” and received approval prior to data collection.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages