SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models Paper • 2509.17664 • Published Sep 22, 2025 • 2
AddressCLIP: Empowering Vision-Language Models for City-wide Image Address Localization Paper • 2407.08156 • Published Jul 11, 2024 • 1
AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models Paper • 2508.10667 • Published Aug 14, 2025 • 1
CoL3D: Collaborative Learning of Single-view Depth and Camera Intrinsics for Metric 3D Shape Recovery Paper • 2502.08902 • Published Feb 13, 2025 • 1
AddressCLIP: Empowering Vision-Language Models for City-wide Image Address Localization Paper • 2407.08156 • Published Jul 11, 2024 • 1
AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models Paper • 2508.10667 • Published Aug 14, 2025 • 1
CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering Paper • 2602.23952 • Published Feb 27 • 4
CoL3D: Collaborative Learning of Single-view Depth and Camera Intrinsics for Metric 3D Shape Recovery Paper • 2502.08902 • Published Feb 13, 2025 • 1
TrackGS: Optimizing COLMAP-Free 3D Gaussian Splatting with Global Track Constraints Paper • 2502.19800 • Published Feb 27, 2025 • 1
Intuitive and Efficient Roof Modeling for Reconstruction and Synthesis Paper • 2109.07683 • Published Sep 16, 2021
CAGE: Continuity-Aware edGE Network Unlocks Robust Floorplan Reconstruction Paper • 2509.15459 • Published Sep 18, 2025
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs Paper • 2512.16584 • Published Dec 18, 2025 • 3
SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models Paper • 2509.17664 • Published Sep 22, 2025 • 2
Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering Paper • 2510.14605 • Published Oct 16, 2025 • 6
Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger Paper • 2506.07785 • Published Jun 9, 2025 • 1
Learning Neural Volumetric Pose Features for Camera Localization Paper • 2403.12800 • Published Mar 19, 2024 • 1
HybridGS: Decoupling Transients and Statics with 2D and 3D Gaussian Splatting Paper • 2412.03844 • Published Dec 5, 2024 • 1