Anthropogenic Regional Adaptation in Multimodal Vision-Language Model Paper • 2604.11490 • Published Apr 13 • 16
Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math Paper • 2602.06291 • Published Feb 6 • 24
Aligning LLMs for Multilingual Consistency in Enterprise Applications Paper • 2509.23659 • Published Sep 28, 2025 • 20
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks Paper • 2509.23673 • Published Sep 28, 2025 • 20
PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications Paper • 2509.23879 • Published Sep 28, 2025 • 20
AccessEval: Benchmarking Disability Bias in Large Language Models Paper • 2509.22703 • Published Sep 22, 2025 • 20
M$^3$FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset Paper • 2506.02510 • Published Jun 3, 2025 • 3 • 3
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Paper • 2506.00482 • Published May 31, 2025 • 8
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Paper • 2506.00482 • Published May 31, 2025 • 8
Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy Paper • 2412.17759 • Published Dec 23, 2024