DataBench

Rather than training one model, we trained multiple models, each on a different dataset. Everything else was kept the same, so that the only changing variable would be the dataset itself.

Model Architecture

  • Base Architecture: LlamaForCasualLM
  • Tokenizer: Harley-ml/Dillionv2-1.3M
  • Transformers Version: 5.13.1
  • Hidden Size: 128
  • Vocab Size: 2564
  • Number of Layers: 8
  • Number of Heads: 4
  • Number of KV Heads: 2
  • Intermediate Size: 344
  • Head Dim: 32
  • Max Position Embeddings: 256
  • RoPe Theta: 2500.0
  • Tie Word Embeddings: true
  • Hidden Activation: silu
  • MLP Bias: false
  • Initializer Range: 0.2
  • RMS Norm Eps: 1e-06
  • Pretraining Tp: 1
  • Use Cache: false
  • Total Parameters: 1,780,352

Training Setup

  • Epochs: 1
  • Max Steps: -1.0
  • Batch Size: 400
  • Sequence Length: 256
  • Gradient Accumulation: 2
  • Gradient Clipping: 1.0
  • Gradient Checkpointing: true
  • Learning Rate: 2.5e-3
  • Eval Split: 0.00165
  • Weight Decay: 0.01
  • Optimizer: AdamW
  • AdamW Betas: (0.9, 0.95)
  • AdamW Eps: 1e-8
  • Scheduler: WSD
  • WSD Warmup Ratio: 0.015
  • WSD Stable Ratio: 0.78
  • WSD Decay Ratio: 0.20
  • WSD Minium LR Ratio: 0.0
  • WSD Number of Cycles: 0.5
  • DType: float16
  • Torch.Compile: true
  • DataLoader Workers: 2
  • Seed: 311

Results

Accuracy is normalized by length and shown as a percentage.

Dataset ARC-Easy HellaSwag PIQA Avg ↑
FineWeb 27.82% 26.89% 52.83% 35.85%
DCLM-1.0-Baseline 29.08% 27.06% 52.23% 36.12%
DOAB 31.14% 28.00% 52.18% 37.11%
Project Gutenberg 25.84% 24.63% 50.21% 33.56%
EOT-2004-Raw 29.08% 27.54% 51.31% 35.98%
Wikipedia 28.28% 27.61% 51.14% 35.68%

(Will add ArithMark-3.0 later)

Notice

This is a work in progress and is currently not completed. By the end of this project, we aim to have tested over 50 datasets.

License

Apache 2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support