4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes Paper • 2610.03715 • Published 6 days ago • 29
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 7 days ago • 274
SEBIS/code_trans_t5_large_code_documentation_generation_python_multitask Summarization • Updated Jun 23, 2021 • 59 • 8
Smaller Models, Better Rejects: Preference Distillation Scaling Paper • 2609.38987 • Published 8 days ago • 28
CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning Paper • 2609.36820 • Published 9 days ago • 39
Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States Paper • 2610.01415 • Published 7 days ago • 94
harryxi/HelpSteer3-general-code-Shift-Qwen-2.5-1.5B-Instruct-chat-formatted-generations Viewer • Updated Aug 17, 2025 • 1M • 891 • 5
DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation Paper • 2609.33711 • Published 11 days ago • 19
PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation Paper • 2609.38597 • Published 9 days ago • 32
SEBIS/code_trans_t5_base_code_documentation_generation_python Summarization • Updated Jun 23, 2021 • 116 • 19
Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing Paper • 2609.37402 • Published 9 days ago • 17