The ultimate guide to multi-harness RL

LFM2.5-2.6B-opencode-SFT

A full fine-tune of LiquidAI/LFM2.5-2.6B for agentic data-analysis tasks, trained with supervised fine-tuning (TRL SFT) using OpenCode. This release is step 1,208 on main, epoch 2, scoring 47.5% pass@1 across four evaluation harnesses.

Article · Collection · Evaluation tasks · Training dashboard

Training and evaluation curves

Animated training, pass@1 and tool-use curves through step 1,208

Training curves use a trailing 50-step mean (at least 10 observations). Evaluation markers show measured checkpoints; hollow markers and dotted segments indicate incomplete coverage. Faint background lines show the full recorded trajectory; colored lines reveal the measured checkpoints. The first frame shows the completed chart before replaying, so previews also contain the full curves. Tool-call savings compare tasks solved by both the base model and checkpoint, with a different matched cohort at each checkpoint. SFT evaluations were recorded at the end of each epoch only. The animation stops at this revision's step 1,208.

Static chart · Plotted data · Interactive article

Training

TRL SFT on FineEnvs/SmolDataEnvs-multiharness-sft, using the OpenCode subset only. The data contains successful Qwen3.8-27B teacher rollouts collected through OpenEnv × Harbor in Daytona sandboxes. Up to three teacher attempts were allowed per task/harness; the first successful trace was retained.

This run uses 4,825 assistant-turn examples from 801 successful rollouts covering 801 distinct tasks. Loss is applied only to the next assistant response; user messages, tool outputs and prior history are context. Responses are tokenized for LFM, not trained against teacher token IDs or logprobs.

Setting Value
Epochs / checkpoint 2 / 1,208
Learning rate / schedule 3e-6 / constant
Effective batch size 8
Optimizer / precision paged AdamW 8-bit / bfloat16
Training harnesses OpenCode
Objective Completion-only supervised fine-tuning

Evaluation

250 fixed SmolDataEnvs test tasks (33 easy, 118 medium, 99 hard), each evaluated under four harnesses: 1,000 graded task/harness cells. Pass@1 uses the first graded attempt per cell; infrastructure retries do not turn it into pass@k. All cells below are graded.

Harness Correct / evaluated Pass@1
OpenCode 129 / 250 51.6%
Claude Code 83 / 250 33.2%
Codex 116 / 250 46.4%
Mini-SWE-Agent 147 / 250 58.8%
Overall 475 / 1,000 47.5%

Harness versions: OpenCode 1.18.31, Claude Code 2.1.270, Codex 0.154.0, Mini-SWE-Agent 2.4.6. Full scores, including difficulty breakdowns, are in eval_results.json.

These are single-run results on a specific task set and harness versions. Harness mix, training exposure and compute differ across runs; the scores do not isolate a causal effect of the harness or objective.

Load the checkpoint

The repository contains full saved model weights and the saved tokenizer/chat template, not a LoRA adapter. Training used Transformers 5.14.1; use a compatible Transformers release.

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "FineEnvs/LFM2.5-2.6B-opencode-SFT"
revision = "main"
tokenizer = AutoTokenizer.from_pretrained(model_id, revision=revision)
model = AutoModelForCausalLM.from_pretrained(
    model_id, revision=revision, dtype=torch.bfloat16, device_map="auto"
)

Reproducing the task scores requires the agent harness and tools described in the article; a plain chat prompt is not the same evaluation.

Related releases

All models, datasets, environments and the article are linked in the multi-harness RL collection.

License and provenance

Derived from LiquidAI/LFM2.5-2.6B under the LFM Open License v1.0. The base model's license is included unchanged in LICENSE. FineEnvs modified the weights by supervised fine-tuning; this is not an official Liquid AI release. See NOTICE and release_manifest.json for modification notices, the pinned base revision, checkpoint identity and file checksums. Optimizer, scheduler, RNG and trainer state are excluded.

Citation

For the experiment, methodology and interpretation, cite the main article:

@misc{kolavi2026multiharnessrl,
  author = {Adithya S Kolavi},
  title = {The ultimate guide to multi-harness RL},
  year = {2026},
  url = {https://huggingface.co/spaces/AdithyaSK/multi-harness-rl}
}
Downloads last month
-
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FineEnvs/LFM2.5-2.6B-opencode-SFT

Finetuned
(60)
this model
Quantizations
2 models

Dataset used to train FineEnvs/LFM2.5-2.6B-opencode-SFT

Collection including FineEnvs/LFM2.5-2.6B-opencode-SFT

Evaluation results