OpenThai-SystemOne — NVFP4A16

OpenThai-SystemOne is an open Thai + English System One decision model: one forward pass answers typed questions (choice over up to 255 options, ordinal score, yes/no noul) about a text / JSON state with calibrated probabilities, no text generation. It is a Qwen3.5-0.8B text tower (Thai continued pre-training) plus a 256-slot decision head. This repo is a quantization of v0.3 (commit f3709948).

What is quantized: the Linear layers of the tower. The token embeddings, the 256-slot decision head and the per-type temperatures stay in bf16. Quantization therefore only perturbs the hidden state the head reads.

Format: NVFP4 weight-only (FP4 weights, 16-bit activations). compressed-tensors (llm-compressor) checkpoint. Loaded through transformers the weights are decompressed to bf16 at load time (same speed as bf16, smaller download); native FP8 / INT8 / FP4 kernels need a runtime with this architecture (the decision head is custom, so vLLM does not serve it out of the box). NVFP4 kernels need NVIDIA Blackwell.

Size: 790 MB (bf16 original: 1,509 MB).

Usage

pip install torch transformers safetensors pydantic && pip install compressed-tensors
from transformers import AutoModel, AutoTokenizer        # trust_remote_code files are in this repo
model = AutoModel.from_pretrained("iapp/OpenThai-SystemOne-NVFP4A16", trust_remote_code=True)
# or with the pip package (git+https://github.com/iapp-technology/openthai-systemone):
from openthai_systemone import SystemOneClient
c = SystemOneClient("iapp/OpenThai-SystemOne-NVFP4A16")
r = c.system_one("ร้านนี้อาหารอร่อยมาก แต่รอนานเกือบชั่วโมง", {"sentiment": {"type": "choice", "instructions": "ความรู้สึก",
                 "criteria": {"บวก": None, "ลบ": None, "กลาง": None}}})

Accuracy vs the bf16 original (same records, single option order, first 800 per set)

Macro: public 73.6 (original 74.3), Thai 79.8 (original 80.1).

subset bf16 original this Δ
public 13-subset bench
aegis2 (noul) 83.2 82.8 -0.4
boolq (noul) 79.7 78.3 -1.3
civil_comments (noul) 79.0 76.0 -3.0
helpsteer2 (score) 41.6 41.6 +0.0
massive-de-DE (choice) 88.3 86.3 -2.0
massive-en-US (choice) 88.3 88.6 +0.3
multinli (choice) 89.0 86.6 -2.3
paws (noul) 94.0 92.4 -1.6
pubmedqa (choice) 64.0 62.0 -2.0
squad2 (noul) 89.3 87.0 -2.3
summeval-consistency (score) 75.0 78.5 +3.5
summeval-relevance (score) 21.7 25.0 +3.3
vitaminc-dev (choice) 72.5 71.6 -0.8
macro, public 13-subset bench 74.3 73.6 -0.7
Thai held-out / eval sets
banking77 (choice) 59.1 59.0 -0.1
contrastive_th (choice) 80.7 79.1 -1.7
contrastive_th (noul) 83.5 81.9 -1.6
contrastive_th (score) 78.6 80.4 +1.8
massive_th (choice) 90.6 89.5 -1.1
prachathai (choice) 98.3 98.3 +0.0
prachathai (noul) 93.4 93.7 +0.3
sib200_th (choice) 77.9 77.9 +0.0
wisesight (choice) 48.9 49.6 +0.8
wongnai (score) 64.5 64.9 +0.4
xlam_tools (choice) 99.4 99.4 +0.0
xnli_th (choice) 79.8 77.6 -2.1
xnli_th (noul) 86.8 86.4 -0.4
macro, Thai held-out / eval sets 80.1 79.8 -0.3

Notes

  • Scores are single-option-order accuracy on the first 800 records of each set (scripts/06_eval.py --limit 800), the same records for the original and the quantization. score subsets report exact level accuracy.
  • Base model, data, training and the full benchmark tables: iapp/OpenThai-SystemOne.
  • License Apache-2.0 (same as the base). Built by iApp Technology / OpenThaiGPT.
Downloads last month
-
Safetensors
Model size
0.8B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for iapp/OpenThai-SystemOne-NVFP4A16

Quantized
(14)
this model

Collection including iapp/OpenThai-SystemOne-NVFP4A16