MBA: Multimodal Benchmark and Agents for Real-World Business Ideation
Abstract
Researchers introduce MBA-Bench, a multimodal benchmark for business ideation agents, and propose MBA-b and MBA-k models trained with creativity and feasibility rewards via LoRA fine-tuning and group relative policy optimization, significantly outperforming text-only and multimodal baselines.
Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we automatically caption images and employ GPT-4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis. Following prior work, we evaluate agents across six business-oriented criteria using MLLM-as-a-Judge. To consider settings where criteria are hidden or disclosed, we present MBA-b and MBA-k for blind and known, respectively. We train both with two novel reward objectives---creativity and feasibility---while MBA-k further optimizes the six disclosed criteria for eight in total. Both are trained via LoRA-based supervised fine-tuning followed by group relative policy optimization with these setting-specific rewards. For extensive experiments on MBA-Bench, we set up two baselines accommodating either captions only or multimodal inputs, with the latter nearing closed-source performance on several metrics. MBA-b and MBA-k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.
Community
Business opportunities exist in the real world—not just in text.
Yet most AI-driven business ideation remains text-centric, overlooking rich visual signals from products, environments, interfaces, and everyday scenes.
We introduce MBA: Multimodal Benchmark and Agents for Real-World Business Ideation, a framework for studying how multimodal AI can turn real-world observations into actionable business ideas.
• MBA-Bench: 30K samples across 6 real-world domains
• MBA-b / MBA-k: multimodal agents optimized for creativity, feasibility, and business-oriented objectives
• +25.6% / +35.8% over open-source multimodal LLMs
Can multimodal AI move beyond understanding the world to uncovering opportunities within it?
- Paper: https://arxiv.org/abs/2608.11616
- Project: https://hchoi256.github.io/projects/mba/
- Code: https://github.com/hchoi256/MBA
- Models: https://huggingface.co/hchoi256/mba
- Dataset: https://huggingface.co/datasets/hchoi256/MBA-Bench
- Interactive Gradio Demo: https://huggingface.co/spaces/hchoi256/mba-business-ideation-agent
🙇 Special thanks to Apolinário Passos (Poli) from Hugging Face for building this interactive demo: https://huggingface.co/apolinario
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents (2026)
- DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents (2026)
- Incentivizing Vision Language Models to Search for Long Video Question Answering (2026)
- Benchmarking Agentic Newswriting via Journalistic Workflows (2025)
- MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models (2026)
- DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams (2026)
- Context-Aware RL for Agentic and Multimodal LLMs (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.11616 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash 