Running on CPU Upgrade 78 MiMo RL Environment Explorer 🧭 78 Explore the MiMo-V2.6 RL environments and run rollouts
Running 252 The ultimate guide to RL environments: building and scaling them in the LLM era 📝 252 Building and scaling RL environments for LLM training
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF Image-Text-to-Text • 27B • Updated 10 days ago • 2.13M • 1.44k
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning Paper • 2606.26790 • Published Jun 25 • 59
view article Article How we OCR'ed 30,000 papers using Codex, open OCR models and Jobs nielsr • Apr 7 • 62