Optimizing Language Model's Reasoning Abilities with Weak Supervision Paper • 2405.04086 • Published May 7, 2024 • 2
Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for Reasoning Paper • 2403.20046 • Published Mar 29, 2024
Eliminating Reasoning via Inferring with Planning: A New Framework to Guide LLMs' Non-linear Thinking Paper • 2310.12342 • Published Oct 18, 2023
ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction Paper • 2608.13622 • Published 16 days ago • 16
Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning Paper • 2608.16554 • Published 12 days ago • 1
ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation Paper • 2310.17389 • Published Oct 26, 2023
ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction Paper • 2608.13622 • Published 16 days ago • 16
Who's Your Judge? On the Detectability of LLM-Generated Judgments Paper • 2509.25154 • Published Sep 29, 2025 • 30
Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens Paper • 2508.01191 • Published Aug 2, 2025 • 240
The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation Paper • 2505.18759 • Published May 24, 2025 • 14