Running on Zero MCP Featured 95 Breeze TTS 2 🎙 95 Bilingual TTS with voice design, cloning, and direction
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis Paper • 2601.10945 • Published Jan 16 • 1
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis Paper • 2601.10945 • Published Jan 16 • 1
left|,circlearrowright,text{BUS},right|: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles Paper • 2511.01340 • Published Nov 3, 2025 • 13
When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs Paper • 2509.16633 • Published Sep 20, 2025 • 2
When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs Paper • 2509.16633 • Published Sep 20, 2025 • 2
When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs Paper • 2509.16633 • Published Sep 20, 2025 • 2 • 2
COFAR: Commonsense and Factual Reasoning in Image Search Paper • 2210.08554 • Published Oct 16, 2022 • 1
Answer Mining from a Pool of Images: Towards Retrieval-Based Visual Question Answering Paper • 2306.16713 • Published Jun 29, 2023 • 1
Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant Paper • 2410.19144 • Published Oct 24, 2024 • 1