InSight-doc: Agentic Visual Perception for Long-Document Understanding Paper • 2608.10628 • Published 2 days ago • 10
CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing Paper • 2508.06937 • Published Aug 9, 2025 • 7
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs Paper • 2507.07562 • Published Jul 10, 2025 • 1
MODNet: Real-Time Trimap-Free Portrait Matting via Objective Decomposition Paper • 2011.11961 • Published Nov 24, 2020
InSight-doc: Agentic Visual Perception for Long-Document Understanding Paper • 2608.10628 • Published 2 days ago • 10
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search Paper • 2512.18745 • Published Dec 21, 2025 • 12
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search Paper • 2512.18745 • Published Dec 21, 2025 • 12
Dual Risk Minimization: Towards Next-Level Robustness in Fine-tuning Zero-Shot Models Paper • 2411.19757 • Published Nov 29, 2024 • 1
OoD-Bench: Quantifying and Understanding Two Dimensions of Out-of-Distribution Generalization Paper • 2106.03721 • Published Jun 7, 2021 • 1
CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving Paper • 2203.07724 • Published Mar 15, 2022 • 1