SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published 12 days ago • 70
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published 12 days ago • 70
Multi-scale Predictive Representations for Goal-conditioned Reinforcement Learning Paper • 2605.09364 • Published May 10
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing Paper • 2603.29852 • Published Feb 22 • 6
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval Paper • 2511.00903 • Published Nov 2, 2025
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning Paper • 2508.09804 • Published Aug 13, 2025
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published 12 days ago • 70
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks Paper • 2606.29537 • Published 25 days ago • 22
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing Paper • 2603.29852 • Published Feb 22 • 6
Mem-$π$: Adaptive Memory through Learning When and What to Generate Paper • 2605.21463 • Published May 20 • 9
Toward Open Weight Models Without Risks: Separating Public and Private Capabilities in LLMs Paper • 2606.21638 • Published Jun 19 • 8
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs Paper • 2410.13210 • Published Oct 17, 2024
Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data Paper • 2505.19274 • Published May 25, 2025
Unifying Adversarial Robustness and Training Across Text Scoring Models Paper • 2602.00857 • Published Jan 31