hypernix.3-mini

A 48.7M-parameter decoder-only language model, pretrained from scratch on 0.87B tokens with a fully custom 32,000-token vocabulary, on one gtx1080 — a consumer card — in 12d 3h.

No distillation, no fine-tune of someone else's checkpoint, no borrowed tokenizer. Random init to final weights, on hardware you can buy for the price of a nice dinner.

This is a base model. It completes text. It has never been instruction-tuned, chat-tuned, RLHF'd, or aligned. It will not follow instructions, and at this size it will not reliably answer questions. See What it is not.

Architecture hyperNix0x-v2 (8L, d_model 512, GQA 8/2, SwiGLU, RMSNorm, RoPE)
Parameters 48,706,048 (48.7M)
Training tokens 873,660,416 (~17.9 tokens/parameter)
Context length 512
Vocabulary 32,000 byte-level BPE, trained from scratch on this corpus
Precision fp32 (Pascal has no bf16 and runs fp16 GEMMs at 1/64 rate)
Training hardware 1x gtx1080
Training time 12d 3h
Training compute ~2.55e17 FLOPs

Benchmarks

Every number below was measured on the hardware named above, by bench_hypernix3_mini.py in this repo, through EleutherAI's lm-evaluation-harness v0.4.13. The baselines were run locally, by the same script, in the same session — none of these figures were copied from another model card.

Baselines are the base (not instruct) checkpoints, because comparing a base model to an instruction-tuned one measures the tuning rather than the pretraining.

Task (0-shot) hypernix.3-mini random
hellaswag — 25
arc_easy — 25
arc_challenge — 25
piqa — 50
winogrande — 50
openbookqa — 25
lambada_openai — —
wikitext ppl ↓ 13503.4 —
Footprint hypernix.3-mini
Parameters 48.7M
Weights (fp32) 0.19 GB
Peak VRAM, 1-stream decode —
Decode tok/s —

Partial evaluation. These numbers come from a --limit 500 pass — 500 documents per task, not the full sets. Treat them as indicative.

Zero-shot. Multiple-choice tasks report length-normalised accuracy (acc_norm); winogrande and lambada report acc; wikitext reports word-level perplexity. ± is the harness's standard error.

Reproduce it
pip install lm-eval transformers accelerate hypernix
python bench_hypernix3_mini.py \
  --model-dir ./hypernix.3-mini \
  --tokenizer-dir ./hypernix.3-mini \
  --device cuda

Full config, per-task metrics and standard errors are in bench_results.json. Run date: 2026-09-25T06:05:34Z; device: NVIDIA GeForce GTX 1080.

How to read this

hypernix.3-mini is several timesx smaller than the smallest baseline here and was trained on a tiny fraction of their token budgets. On knowledge and commonsense tasks it is not competitive with them, and this card is not going to pretend otherwise — at 48.7M parameters and 0.87B tokens, scores at or near the random baseline on the harder multiple-choice sets are the expected result, not a defect.

What the model is actually for is the other table: it is small enough to train end to end on one consumer GPU in days, small enough to run anywhere, and it is a clean, honest, fully-reproducible artifact of that process.


Usage

With hypernix

from hypernix.models.neo_oven import preheat_brewed

oven = preheat_brewed("ray0rf1re/hypernix.3-mini", tokenizer_source="ray0rf1re/hypernix.3-mini")
print(oven.complete("The reason the sky is blue is", max_new_tokens=64, temperature=0.8))

With plain PyTorch

import torch, json
from huggingface_hub import snapshot_download
from hypernix.training.brewer import BrewerConfig, BrewerModel
from transformers import AutoTokenizer

path = snapshot_download("ray0rf1re/hypernix.3-mini")
cfg = BrewerConfig.load(f"{path}/config.json")
model = BrewerModel(cfg).eval()
model.load_state_dict(torch.load(f"{path}/model.pt", map_location="cpu"))
tok = AutoTokenizer.from_pretrained(path)

ids = torch.tensor([tok.encode("Once upon a time")])
for _ in range(32):
    nxt = model(ids)[:, -1].argmax(-1, keepdim=True)
    ids = torch.cat([ids, nxt], dim=1)
print(tok.decode(ids[0]))

Note on sliding-window attention. This model was trained with use_sliding_window=false, and you should load it that way. The hypernix.training.brewer release this was trained against combines the causal and sliding-window masks with torch.maximum, which for additive masks keeps the less masked of the two — so enabling SWA lets every query on odd-numbered layers see up to sliding_window_size - 1 tokens of its own future. The training script in this repo ships a corrected attention forward and turns SWA off by default.


Architecture

hyperNix0x-v2 is a decoder-only transformer: RMSNorm (pre-norm), rotary position embeddings, grouped-query attention, and a SwiGLU feed-forward block, with the input embedding tied to the output head.

Layers 8
Model dim 512
FFN dim 2203 (SwiGLU)
Attention heads 8 query / 2 key-value (4:1 GQA)
Head dim 64
RoPE theta 100,000
Norm RMSNorm, eps 1e-05
Tied embeddings yes (34% of all parameters are the embedding matrix)
Max context 512

Training

Data HuggingFaceFW/fineweb-edu
Tokens seen 873,660,416
Optimizer pcv6 (betas 0.9/0.95, weight decay 0.1, grad clip 1.0)
LR schedule linear warmup 369 steps → cosine to 10% of peak
Peak LR 6.0e-4
Batch 32,768 tokens/step (8 x 512 x 8 accumulation)
Steps 18,464
Precision fp32 throughout
Final val loss 7.8398 (perplexity 2539.79)

The full training script is in this repo: train_hypernix3_mini.py. It streams the corpus, trains the vocabulary, packs tokens into a uint16 memmap, and trains — end to end, one command, resumable.

Tokenizer

A byte-level BPE trained from scratch on a 1.5 GB sample of the training corpus. 32,000 tokens, no unknown token, with <|endoftext|>, <|pad|> and the three fill-in-the-middle markers (<fim_prefix>, <fim_middle>, <fim_suffix>) reserved. Round-trips byte-exactly.

Data

Streamed from the sample-10BT shard of FineWeb-Edu — Common Crawl web text filtered by an educational-quality classifier. The stream was tokenized directly into a flat uint16 array and the raw text was never stored, so the whole corpus footprint on disk was the token file. Documents are separated by <|endoftext|> and packed contiguously; no additional filtering, deduplication or reweighting was applied beyond what the source dataset already does.


What it is not

  • Not an instruction-following model. No SFT, no chat template, no alignment. Prompt it like a text completer, not an assistant.
  • Not a knowledge model. 48.7M parameters and 0.87B tokens is not enough capacity or exposure to store reliable facts. It will state false things fluently.
  • Not safety-tuned. It reproduces the distribution of its web-derived training data, including its biases. Filter its output before showing it to anyone.
  • Not multilingual. The corpus and the vocabulary are English.
  • Not long-context. Trained at 512 tokens; quality past that is undefined.

Limitations and bias

The training corpus is derived from Common Crawl web text filtered by an educational quality classifier. It carries the demographic, geographic and viewpoint skew of the English-language web, and no additional debiasing, deduplication beyond the source dataset's own, or toxicity filtering was applied. At this scale the model has no meaningful capacity for factual recall or reasoning, so treat any confident-sounding output as a fluent guess.

License

Apache-2.0 for the weights and code in this repo. The training data (HuggingFaceFW/fineweb-edu) carries its own license — see the source dataset.

Citation

@misc{hypernix3mini,
  title  = {hypernix.3-mini: a 48.7M-parameter language model pretrained from scratch on a single GTX 1080},
  author = {Rayla (ray0rf1re)},
  year   = {2026},
  url    = {https://huggingface.co/ray0rf1re/hypernix.3-mini}
}

Built with HyperNix.

Downloads last month
311
Safetensors
Model size
48.7M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train ray0rf1re/HyperNix.3-mini

Collection including ray0rf1re/HyperNix.3-mini