N8Programs's picture
Add dataset card
0dc4ea9 verified
|
Raw
History Blame Contribute Delete
1.73 kB
metadata
license: mit
task_categories:
  - text-generation
tags:
  - arc-agi
  - re-arc
  - test-time-training
  - synthetic-data
pretty_name: ARC Tiny Transformer corpus and evaluation candidates

ARC Tiny Transformer data

Data and vote-level artifacts for N8python/arc-tiny-transformer.

Corpus

corpus/documents_3k_per_task.jsonl.zst is the exact pretraining corpus:

  • 1,200,000 JSONL documents;
  • 3,000 documents for each of 400 RE-ARC generator families;
  • 3,401,127,124 ARC tokens;
  • 10,413,648,837 uncompressed bytes;
  • uncompressed SHA-256 5d1336f8d6f45358a377af2e3fd05d43c4bfe91621b0c6a44c53dad0e34ca430.

Each record contains token IDs and metadata for task ID, sampled difficulty band, demonstration count, rejection statistics, and generation seed. The fixed vocabulary and sampler are in the GitHub repository. manifests/ contains the complete task-level corpus provenance.

Evaluation artifacts

artifacts/ contains the full 128-candidate records for:

  • the frozen 50M model;
  • TTT replicas seeded 57, 58, and 59;
  • the 7M TTT run.

These files are sufficient to recompute rank-1, raw-frequency top-2, historical hierarchical top-2, oracle accuracy, candidate-scaling curves, and multi-replica ensembles without rerunning model inference.

viewer/data/ contains the compact per-query data backing the browser-based 384-candidate explorer.

Licensing note

Release code and metadata are MIT licensed. The corpus is procedurally generated with the pinned RE-ARC generators/verifiers; ARC-AGI task files are not duplicated here. Users should also review the upstream RE-ARC and ARC-AGI licenses for their use case.