ExplorerArtificial IntelligenceAI
Research PaperResearchia:202607.30062

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

Weijie Wu

Abstract

With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 ...

Submitted: July 30, 2026Subjects: AI; Artificial Intelligence

Description / Details

With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the reference label. We evaluate representative open-source and proprietary AudioLLMs. Results show clear differences across evaluation dimensions, information coverage, and description reliability. We will release the MMAC benchmark and evaluation code.


Source: arXiv:2607.27109v1 - http://arxiv.org/abs/2607.27109v1 PDF: https://arxiv.org/pdf/2607.27109v1 Original Link: http://arxiv.org/abs/2607.27109v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Jul 30, 2026
Topic:
Artificial Intelligence
Area:
AI
Comments:
0
Bookmark
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning | Researchia