Instructions to use FINAL-Bench/Darwin-27B-ZTC-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FINAL-Bench/Darwin-27B-ZTC-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="FINAL-Bench/Darwin-27B-ZTC-v2")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("FINAL-Bench/Darwin-27B-ZTC-v2") model = AutoModel.from_pretrained("FINAL-Bench/Darwin-27B-ZTC-v2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Darwin-27B-ZTC-v2
A zero-token decision engine from the Darwin family, second version. Darwin-27B-ZTC-v2 reads a piece of state and a typed question (noul yes or no, choice one of N labels, score an ordered rubric) and returns a probability for every option. It uses one forward pass per question and generates no tokens.
v2 adds control and workflow decisions (games, drones, retrieval control, entity alignment, customer and incident workflows) on top of FINAL-Bench/Darwin-27B-ZTC.
Results
S1MB (System One Mosaic Benchmark), english-v1: #1
#1 on the S1MB leaderboard under both of its rankings: the default Borda score and Task Avg. All 137 benchmarks complete, measured with the S1MB evaluator and its autojev adapter at revision a8f283a, then validated and merged by the S1MB maintainer (results).
The numbers below follow the leaderboard's own code (viewer/src/lib/borda.ts, types.ts) applied to the published results file of 2026-10-09 02:44 UTC, over the 102 models with complete results.
- Borda score (default sort): rank the models on every benchmark, give 100 points to first place and 0 to last, and average over the 137 benchmarks.
- Task Avg: mean baseline-adjusted score (0 to 100) within each task, then the mean of Noul, Choice and Score.
| Rank | Model | Borda score | Task Avg | Noul (59) | Choice (57) | Score (21) |
|---|---|---|---|---|---|---|
| 1 | Darwin-27B-ZTC-v2 | 89.58 | 66.46 | 67.66 | 71.51 | 60.21 |
| 2 | openjev/openjev | 87.50 | 62.60 | 65.22 | 67.74 | 54.83 |
| 3 | denis-pplx/AutoJev-27B | 87.07 | 60.80 | 64.96 | 68.11 | 49.34 |
| 4 | caiovicentino1/Eikos-27B | 85.43 | 59.86 | 64.14 | 67.61 | 47.83 |
| 5 | TypeSafe Jev 1.13 | 85.05 | 59.59 | 64.63 | 67.22 | 46.92 |
qasper-noul-test-v1 was run with --context-limit 32768 (12 of its decisions exceed the default 8192 tokens; nothing is truncated), recorded in the result metadata.
Training overlap (disclosed). v2 was trained on the train split of ZefanCai/Open-Jev (CC0). S1MB includes 22 benchmarks built from the test split of the same Open-Jev tasks. We used no S1MB test case: an exact-match filter over every S1MB test state, query and document removed zero training rows. v2 was not trained on any S1MB data or on Typed Decisions data.
How it was made
- Start from Darwin-27B-ZTC (v1).
- Continue full-weight training on the Open-Jev
trainsplit mixed with v1's original training data, selecting the checkpoint on held-out development rows only. - Average the weights of v1 and the continued model (50/50). The average keeps v1's skills on its original tasks and adds the new control and workflow skills.
Files and usage
The layout is the same as v1: backbone weights (BF16, about 54 GB), readout.safetensors, decision_config.json, tokenizer, and the inference code in autojev/ (from autojev, MIT, with text-only backbone support added).
ztc_server.py: aPOST /v1/systemoneserver (ZTC_MODEL=<path> PORT=8000 python ztc_server.py).ztc_engine.py: an in-process engine for the Decision Index kit.
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("FINAL-Bench/Darwin-27B-ZTC-v2")
sys.path.insert(0, path)
from autojev.model import DecisionModel
model = DecisionModel(checkpoint=path)
row = {"state": {"ticket": "Charged twice for one order."},
"question": {"type": "choice", "instructions": "What should support do?",
"criteria": {"refund": "Refund the duplicate charge.", "escalate": "Send to billing.", "close": "No action."}}}
print(model.predict([row])[0])
Citation
@misc{darwin27bztcv2,
title = {Darwin-27B-ZTC-v2: a zero-token decision engine},
author = {VIDRAFT and FINAL-Bench},
year = {2026},
url = {https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC-v2}
}
- Downloads last month
- 55