Running Featured 81 QED-Nano: Teaching a Tiny Model to Prove Hard Theorems 📝 81 Who needs 1T parameters? Olympiad proofs with a 4B model
WTF GENIUS PAPERS Collection Papers that made me appreciate my major and my life a little more. obs=Observation, innov=Innovation. Most papers are abt improving tiny models. • 227 items • Updated 2 days ago • 51
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published 4 days ago • 140
SWE-Pruner Pro: The Coder LLM Already Knows What to Prune Paper • 2607.18213 • Published 7 days ago • 76
Understanding Reasoning from Pretraining to Post-Training Paper • 2607.16097 • Published 10 days ago • 28
DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation Paper • 2607.13365 • Published 12 days ago • 19
Does Your Reasoning Model Implicitly Know When to Stop Thinking? Paper • 2602.08354 • Published Feb 9 • 267
Dockerless: Environment-Free Program Verifier for Coding Agents Paper • 2606.28436 • Published Jun 26 • 112
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published 11 days ago • 201
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 11 days ago • 102
view article Article Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries +7 aminediroHF, qgallouedec, kashif, lewtun, edbeeching, albertvillanova, nouamanetazi, lvwerra, sergiopaniego • Mar 10 • 173
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published 19 days ago • 138
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Paper • 2607.07508 • Published 19 days ago • 26