classone-gemma4-e2b — ClassOne System 1 Decision Model
devops-thiago/classone-gemma4-e2b is an open-source System 1 decision model using the ClassOne architecture. The full fine-tuned backbone ships directly in this repository — it loads as a single model, with no adapter and no separate base-model download.
Instead of generating text token by token, ClassOne evaluates structured decisions in a single forward pass, returning typed, calibrated outputs with zero decoding overhead.
Benchmark Results
1. JevBench Public Multi-Tier Benchmark (231 Public Tasks)
Evaluated across all 231 public tasks in fstandhartinger/jevbench:
| Tier | Tasks | Accuracy | ECE | Brier Score | Median Latency (p50) |
|---|---|---|---|---|---|
| Easy | 48 | 93.8% (45/48) | 0.0610 | 0.0516 | 45.0 ms |
| Original | 72 | 63.9% (46/72) | 0.2510 | 0.2499 | 43.1 ms |
| Hard | 111 | 37.8% (42/111) | 0.4299 | 0.4022 | 91.3 ms |
| Overall Aggregate | 231 | 57.6% (133/231) | — | — | ~44 ms |
- Easy Tier Sub-Breakdown: Choice accuracy: 100.0% (36/36); Noul policy accuracy: 75.0% (9/12).
- Original Tier Sub-Breakdown: Choice accuracy: 66.7% (24/36); Score rubrics: 66.7% (8/12); Noul accuracy: 58.3% (14/24).
- Hard Tier Sub-Breakdown: Noul policy compliance: 44.7% (17/38); Choice accuracy: 34.3% (23/67); Score rubrics: 33.3% (2/6).
2. RLCDAlignBench Alignment & Safety Evaluation (100 Instances)
Evaluated across the 10 core AI alignment failure modes (arXiv:2609.29429):
| Failure Mode / Axis | Samples (N) | AUROC | Accuracy (%) | ECE | Latency (p50) |
|---|---|---|---|---|---|
| Concealing Uncertainty | 14 | 0.980 | 85.7% | 0.0718 | 130.2 ms |
| Honesty (Deception) | 11 | 0.800 | 72.7% | 0.2445 | 219.6 ms |
| Refusal (Jailbreaks) | 11 | 0.667 | 54.5% | 0.1934 | 271.1 ms |
| Power Seeking | 6 | 0.556 | 50.0% | 0.1794 | 213.2 ms |
| Reward Hacking | 9 | 0.500 | 33.3% | 0.2935 | 209.6 ms |
| Prompt Injection | 8 | 0.500 | 37.5% | 0.3207 | 167.9 ms |
| Bias | 9 | 0.375 | 55.6% | 0.1659 | 221.5 ms |
| Overall Average | 100 | 0.516 | 51.0% | 0.1584 | 200.0 ms |
3. Edge vs Cloud Latency (ClassOne vs TypeSafe Jev API)
Measured against TypeSafe AI's Jev (v1.13) cloud API:
- ClassOne (Local RTX 5060 Ti): 52.49 ms mean latency (19.1 req/s, $0.00 inference cost, 100% private)
- TypeSafe Jev (Cloud API): 329.90 ms mean latency (3.0 req/s)
- Edge Speedup: 6.3× faster than cloud API round-trip latency
Decision Primitives
Noul— Boolean check returning a calibrated probability P(true) ∈ [0, 1]Choice— Categorical selection over 2–255 dynamic options with full probability distributionScore— Continuous ordinal rubric rating over 2–10 levels (expected value)
All outputs are calibrated with a combined NLL + normalized Brier loss. Post-hoc temperature calibration achieves ECE = 0.034 (down from 0.178).
Quickstart
pip install classone
import torch
from huggingface_hub import hf_hub_download
from transformers import AutoTokenizer
from classone.modeling.modeling_classone import ClassOneModel
from classone.schemas import NoulQuestion, ChoiceQuestion, ScoreQuestion
from classone.tokenizer import ClassOnePromptBuilder
REPO_ID = "devops-thiago/classone-gemma4-e2b"
# 1. Load the ClassOne model (weights + tokenizer are fully self-contained here)
tokenizer = AutoTokenizer.from_pretrained(REPO_ID)
builder = ClassOnePromptBuilder(tokenizer)
model = ClassOneModel.from_backbone(
base_model_name_or_path=REPO_ID,
tokenizer=tokenizer,
device="cuda",
torch_dtype=torch.float16,
)
# 2. Load the trained decision heads
heads = torch.load(hf_hub_download(REPO_ID, "classone_heads.pt"), map_location="cuda")
model.noul_head.load_state_dict(heads["noul_head"])
model.choice_head.load_state_dict(heads["choice_head"])
model.score_head.load_state_dict(heads["score_head"])
model.eval()
# 3. Pack state + questions and run a single forward pass
packed = builder.pack(
state={"customer": "Alex", "message": "I was charged twice for order #123."},
questions={
"refund": NoulQuestion(instructions="Is the user requesting a refund?"),
"dept": ChoiceQuestion(
instructions="Route to team:",
criteria={"billing": "Payment issues", "tech": "Technical bugs"}
),
"anger": ScoreQuestion(
instructions="Dissatisfaction level:",
criteria=["satisfied", "neutral", "dissatisfied", "churning"]
),
}
)
results = model.evaluate_packed(packed)
print("Refund P(true):", results["refund"].noul)
print("Department: ", results["dept"].choice, "—", results["dept"].probabilities)
print("Anger score: ", results["anger"].score)
Repository Files
| File | Description |
|---|---|
model.safetensors (sharded) |
Merged ClassOne backbone weights |
config.json |
Model configuration |
tokenizer.json, tokenizer_config.json |
Tokenizer, including ClassOne delimiter tokens |
classone_heads.pt |
Trained Noul / Choice / Score head weights + calibrated temperatures |
lora_backbone/ |
LoRA adapter (r=16, α=32) that produced the merged weights |
Citation
@misc{classone2026,
title={ClassOne: A Fast Single-Pass Decision Architecture for Language Models},
author={Thiago Gonzaga},
year={2026},
url={https://github.com/devops-thiago/class-one},
}
Attribution & Legal
- Derived from google/gemma-4-E2B-it (Google) — Apache License 2.0
- Architecture & training code: devops-thiago/class-one — Apache 2.0
- Downloads last month
- 416