hypernix.3-mini
A 48.7M-parameter decoder-only language model, pretrained from scratch on 0.87B tokens with a fully custom 32,000-token vocabulary, on one gtx1080 — a consumer card — in 12d 3h.
No distillation, no fine-tune of someone else's checkpoint, no borrowed tokenizer. Random init to final weights, on hardware you can buy for the price of a nice dinner.
This is a base model. It completes text. It has never been instruction-tuned, chat-tuned, RLHF'd, or aligned. It will not follow instructions, and at this size it will not reliably answer questions. See What it is not.
| Architecture | hyperNix0x-v2 (8L, d_model 512, GQA 8/2, SwiGLU, RMSNorm, RoPE) |
| Parameters | 48,706,048 (48.7M) |
| Training tokens | 873,660,416 (~17.9 tokens/parameter) |
| Context length | 512 |
| Vocabulary | 32,000 byte-level BPE, trained from scratch on this corpus |
| Precision | fp32 (Pascal has no bf16 and runs fp16 GEMMs at 1/64 rate) |
| Training hardware | 1x gtx1080 |
| Training time | 12d 3h |
| Training compute | ~2.55e17 FLOPs |
Benchmarks
Every number below was measured on the hardware named above, by
bench_hypernix3_mini.py in this repo, through
EleutherAI's lm-evaluation-harness
v0.4.13. The baselines were run locally, by the same script, in the same
session — none of these figures were copied from another model card.
Baselines are the base (not instruct) checkpoints, because comparing a base model to an instruction-tuned one measures the tuning rather than the pretraining.
| Task (0-shot) | hypernix.3-mini | random |
|---|---|---|
| hellaswag | — | 25 |
| arc_easy | — | 25 |
| arc_challenge | — | 25 |
| piqa | — | 50 |
| winogrande | — | 50 |
| openbookqa | — | 25 |
| lambada_openai | — | — |
| wikitext ppl ↓ | 13503.4 | — |
| Footprint | hypernix.3-mini |
|---|---|
| Parameters | 48.7M |
| Weights (fp32) | 0.19 GB |
| Peak VRAM, 1-stream decode | — |
| Decode tok/s | — |
Partial evaluation. These numbers come from a
--limit 500pass — 500 documents per task, not the full sets. Treat them as indicative.
Zero-shot. Multiple-choice tasks report length-normalised accuracy (acc_norm);
winogrande and lambada report acc; wikitext reports word-level perplexity.
± is the harness's standard error.
Reproduce it
pip install lm-eval transformers accelerate hypernix
python bench_hypernix3_mini.py \
--model-dir ./hypernix.3-mini \
--tokenizer-dir ./hypernix.3-mini \
--device cuda
Full config, per-task metrics and standard errors are in
bench_results.json. Run date: 2026-09-25T06:05:34Z; device: NVIDIA GeForce GTX 1080.
How to read this
hypernix.3-mini is several timesx smaller than the smallest baseline here and was trained on a tiny fraction of their token budgets. On knowledge and commonsense tasks it is not competitive with them, and this card is not going to pretend otherwise — at 48.7M parameters and 0.87B tokens, scores at or near the random baseline on the harder multiple-choice sets are the expected result, not a defect.
What the model is actually for is the other table: it is small enough to train end to end on one consumer GPU in days, small enough to run anywhere, and it is a clean, honest, fully-reproducible artifact of that process.
Usage
With hypernix
from hypernix.models.neo_oven import preheat_brewed
oven = preheat_brewed("ray0rf1re/hypernix.3-mini", tokenizer_source="ray0rf1re/hypernix.3-mini")
print(oven.complete("The reason the sky is blue is", max_new_tokens=64, temperature=0.8))
With plain PyTorch
import torch, json
from huggingface_hub import snapshot_download
from hypernix.training.brewer import BrewerConfig, BrewerModel
from transformers import AutoTokenizer
path = snapshot_download("ray0rf1re/hypernix.3-mini")
cfg = BrewerConfig.load(f"{path}/config.json")
model = BrewerModel(cfg).eval()
model.load_state_dict(torch.load(f"{path}/model.pt", map_location="cpu"))
tok = AutoTokenizer.from_pretrained(path)
ids = torch.tensor([tok.encode("Once upon a time")])
for _ in range(32):
nxt = model(ids)[:, -1].argmax(-1, keepdim=True)
ids = torch.cat([ids, nxt], dim=1)
print(tok.decode(ids[0]))
Note on sliding-window attention. This model was trained with
use_sliding_window=false, and you should load it that way. Thehypernix.training.brewerrelease this was trained against combines the causal and sliding-window masks withtorch.maximum, which for additive masks keeps the less masked of the two — so enabling SWA lets every query on odd-numbered layers see up tosliding_window_size - 1tokens of its own future. The training script in this repo ships a corrected attention forward and turns SWA off by default.
Architecture
hyperNix0x-v2 is a decoder-only transformer: RMSNorm (pre-norm), rotary position
embeddings, grouped-query attention, and a SwiGLU feed-forward block, with the input
embedding tied to the output head.
| Layers | 8 |
| Model dim | 512 |
| FFN dim | 2203 (SwiGLU) |
| Attention heads | 8 query / 2 key-value (4:1 GQA) |
| Head dim | 64 |
| RoPE theta | 100,000 |
| Norm | RMSNorm, eps 1e-05 |
| Tied embeddings | yes (34% of all parameters are the embedding matrix) |
| Max context | 512 |
Training
| Data | HuggingFaceFW/fineweb-edu |
| Tokens seen | 873,660,416 |
| Optimizer | pcv6 (betas 0.9/0.95, weight decay 0.1, grad clip 1.0) |
| LR schedule | linear warmup 369 steps → cosine to 10% of peak |
| Peak LR | 6.0e-4 |
| Batch | 32,768 tokens/step (8 x 512 x 8 accumulation) |
| Steps | 18,464 |
| Precision | fp32 throughout |
| Final val loss | 7.8398 (perplexity 2539.79) |
The full training script is in this repo:
train_hypernix3_mini.py. It streams the corpus,
trains the vocabulary, packs tokens into a uint16 memmap, and trains — end to end, one
command, resumable.
Tokenizer
A byte-level BPE trained from scratch on a 1.5 GB sample of the training corpus.
32,000 tokens, no unknown token, with <|endoftext|>, <|pad|> and the three
fill-in-the-middle markers (<fim_prefix>, <fim_middle>, <fim_suffix>) reserved.
Round-trips byte-exactly.
Data
Streamed from the sample-10BT shard of FineWeb-Edu — Common Crawl web text filtered by an educational-quality classifier. The stream was tokenized directly into a flat uint16 array and the raw text was never stored, so the whole corpus footprint on disk was the token file. Documents are separated by <|endoftext|> and packed contiguously; no additional filtering, deduplication or reweighting was applied beyond what the source dataset already does.
What it is not
- Not an instruction-following model. No SFT, no chat template, no alignment. Prompt it like a text completer, not an assistant.
- Not a knowledge model. 48.7M parameters and 0.87B tokens is not enough capacity or exposure to store reliable facts. It will state false things fluently.
- Not safety-tuned. It reproduces the distribution of its web-derived training data, including its biases. Filter its output before showing it to anyone.
- Not multilingual. The corpus and the vocabulary are English.
- Not long-context. Trained at 512 tokens; quality past that is undefined.
Limitations and bias
The training corpus is derived from Common Crawl web text filtered by an educational quality classifier. It carries the demographic, geographic and viewpoint skew of the English-language web, and no additional debiasing, deduplication beyond the source dataset's own, or toxicity filtering was applied. At this scale the model has no meaningful capacity for factual recall or reasoning, so treat any confident-sounding output as a fluent guess.
License
Apache-2.0 for the weights and code in this repo. The training data (HuggingFaceFW/fineweb-edu) carries its own license — see the source dataset.
Citation
@misc{hypernix3mini,
title = {hypernix.3-mini: a 48.7M-parameter language model pretrained from scratch on a single GTX 1080},
author = {Rayla (ray0rf1re)},
year = {2026},
url = {https://huggingface.co/ray0rf1re/hypernix.3-mini}
}
Built with HyperNix.
- Downloads last month
- 311