StephYang/qwen3.5-4b-swe-opd-tmax500-baseline-step105 Text Generation • 4B • Updated 3 days ago • 369
StephYang/qwen3.5-4b-swe-opd-tmax500-baseline-step105 Text Generation • 4B • Updated 3 days ago • 369
Lego-RL Collection Harness-native RL for coding agents: the trained policy and the training task index. • 6 items • Updated 6 days ago • 11
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents Paper • 2608.17393 • Published 23 days ago • 25