Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs
Abstract
Hybrid-thinking multimodal language models suffer from response-pattern misalignment between thinking and non-thinking modes, which is addressed by a diagnostic benchmark and pattern-specific reinforcement learning penalties.
Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, their delivered responses should satisfy the same user-facing standard. Correctness alone may not characterize this response quality; we therefore evaluate task accuracy and response-pattern failures as complementary outcomes. We study this gap through response-pattern alignment: whether thinking and non-thinking interfaces preserve acceptable final-response behavior. We introduce PatternEval, a failure-enriched diagnostic benchmark comprising 2,415 multimodal prompts spanning visual perception and grounding, structured image understanding, and multimodal knowledge reasoning. PatternEval tests four recurrent failures: chain-of-thought leakage, response repetition, logical contradiction, and performative reasoning. Response-pattern failures are widespread across models from different providers, with non-thinking inference exhibiting substantially higher failure rates and thereby creating systematic misalignment between thinking and non-thinking interfaces. Motivated by this diagnosis, we develop PatternRM, a response-level reward model, and PatternRL, which introduces pattern-specific penalties during reinforcement learning. Experiments on Qwen3-VL-4B and Qwen3-VL-8B show that incorporating pattern-specific penalties into reinforcement learning can mitigate cross-mode misalignment while incurring a marginal task performance trade-off. Together, PatternEval and PatternRL provide an evaluation-and-training framework for aligning user-visible response patterns across hybrid-thinking interfaces.
Community
PatternEval reveals widespread response-pattern misalignment between thinking and non-thinking modes in hybrid-thinking MLLMs, with non-thinking inference exhibiting substantially more failures such as chain-of-thought leakage, repetition, contradiction, and performative reasoning. PatternRL mitigates this cross-mode misalignment by incorporating pattern-specific penalties during reinforcement learning, with only a marginal trade-off in task performance.
The evaluation dataset is currently going through our internal approval process. Thank you for your patience—we’ll make it available as soon as the review is complete.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment (2026)
- A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models (2026)
- SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification (2026)
- The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics (2026)
- REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning (2026)
- Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization (2026)
- Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.12781 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper