YAML Metadata Warning:The pipeline tag "decision-model" is not in the official list: text-classification, token-classification, table-question-answering, question-answering, zero-shot-classification, translation, summarization, feature-extraction, text-generation, fill-mask, sentence-similarity, text-to-speech, text-to-audio, automatic-speech-recognition, audio-to-audio, audio-classification, audio-text-to-text, voice-activity-detection, depth-estimation, image-classification, object-detection, image-segmentation, text-to-image, image-to-text, image-to-image, image-to-video, unconditional-image-generation, video-classification, reinforcement-learning, robotics, tabular-classification, tabular-regression, tabular-to-text, table-to-text, multiple-choice, text-ranking, text-retrieval, time-series-forecasting, text-to-video, image-text-to-text, image-text-to-image, image-text-to-video, visual-question-answering, document-question-answering, zero-shot-image-classification, graph-ml, mask-generation, zero-shot-object-detection, text-to-3d, image-to-3d, image-feature-extraction, video-text-to-text, keypoint-detection, visual-document-retrieval, any-to-any, video-to-video, other

tasksource-jev-nano-v1

A 149M-parameter decision model. Given a state, a question and a set of options, it returns a probability for each option. It follows the "Jev-style" typed-decision interface (choice / noul) and is evaluated on the Decision Index.

It is LateOn (ModernBERT-base, ColBERT multi-vector), fine-tuned on Tasksource typed decisions. Options are scored by late interaction:

context  = state + "\nQuestion: " + question     → encoded once (ColBERT document side)
option_i                                          → encoded independently (ColBERT query side)
logit_i  = MaxSim(option_i, context) / T          → softmax over options      (T = 0.3168)

The context never sees the options, so you can encode a state once and score any number of options against it. This also makes the scores permutation-equivariant.

Usage

import torch
from pylate import models
from pylate.scores import colbert_scores

model = models.ColBERT("tasksource/tasksource-jev-nano-v1")
T = 0.31683406233787537  # learned temperature (decision_meta.json: init_temperature)

def decide(state, question, options):
    ctx = model.encode([f"{state}\nQuestion: {question}"], is_query=False)[0]
    opts = model.encode(options, is_query=True)
    scores = torch.stack([colbert_scores(torch.as_tensor(o)[None], torch.as_tensor(ctx)[None])[0, 0] for o in opts])
    return dict(zip(options, torch.softmax(scores / T, 0).tolist()))

decide("The movie was a total waste of two hours.", "What is the sentiment?", ["positive", "negative"])
# {'positive': 0.0088, 'negative': 0.9912}

This snippet reproduces our evaluation engine's probabilities exactly. For noul (yes/no) questions, score the two options false and true and report P(true).

When an option has both a key and a description, the engine renders it as "key: description", or as the description alone when the key is a placeholder like A or option_1.

Training

  • Init: lightonai/LateOn.
  • Data: grouped requests (one state, 1–5 questions, 2–235 options each) from tasksource-jev-typed-decisions, covering 503 Tasksource training tasks.
    • Training used a 64k-request double-firewalled manifest: every row that fingerprint-matches a Decision Index suite row was removed before training.
  • Recipe:
    • Stage 1: 4,000 steps of grouped MaxSim training with soft cross-entropy over options and a learned temperature.
    • Stage 2: continued for 400 steps with question packing at p=0.5 (several questions share one context half the time).
  • Checkpoint selection: NLL on a held-out set of unseen Tasksource tasks.

Decision Index results (public index)

These are scores on the public part of the Decision Index, run with the official kit: full suite, 100% coverage, chance-corrected skill, 0–100.

Edition Public index (skill) Raw
0.3 10.84 33.12
0.2.1 10.46 32.09

0.3 areas (skill): Knowledge & Reasoning 2.2 · Language Understanding 5.0 · Retrieval & Classification 23.5 · Tools & Automation 13.6 · Arts & Human Taste 17.7.

This is not the board's headline number. From 0.3 on, the board ranks a Full score:

Part Weight Who runs it
Public index 20% Anyone, with the kit; the table above
Private tests of the same skills 50% Maintainers
Private tasks from new domains 30% Maintainers

We do not have this model's Full score; see the board for it once it has been scored.

Limitations

  • Small model, broad but shallow. It does well on retrieval, classification and tool selection. Knowledge and reasoning are near chance: 149M parameters do not store much world knowledge.
  • Typed decisions only. It handles choice and noul. It does not generate text, and it does not natively produce scores or multilabel outputs.
  • English only.
Downloads last month
253
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tasksource/tasksource-jev-nano-v1

Base model

lightonai/LateOn
Finetuned
(14)
this model

Dataset used to train tasksource/tasksource-jev-nano-v1