🔥 MMAE: A Massive Multitask Audio Editing Benchmark
📅 Published on Jun 5
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2606.07229
• PDF: https://arxiv.org/pdf/2606.07229
📊 Datasets citing this paper:
• https://huggingface.co/datasets/BoJack/MMAE
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#AudioEditingBenchmarks #MultimodalAudioProcessing #InstructionBasedAudioEditing #AudioTaxonomy #MultitaskLearningModels
💡 The paper introduces MMAE, a comprehensive benchmark for instruction-based audio editing that evaluates models across multiple modalities and complexity levels. The current evaluation infrastructure for audio editing is fragmented and limited in scope, which motivated the creation of MMAE. This benchmark extends to a broad spectrum of real-world scenarios, covering 7 distinct audio modalities, including sound, speech, music, and their mixtures, and establishes a comprehensive taxonomy spanning 6 levels of task complexity, 2 levels of granularity, and 8 distinct operation types.
The MMAE benchmark was meticulously curated through human-agent collaboration and comprises 2000 high-fidelity samples paired with a pioneering rubric-based evaluation framework. This framework enables a precise, multi-dimensional assessment of both instruction following and context consistency by decomposing free-form tasks into 17741 verifiable criteria.
The paper evaluates leading models using the MMAE benchmark and reveals significant gaps in current model capabilities. The results show that the Exact Match Rate consistently falls below 5% and plummets to 0% in complex, mixed-modality tasks, exposing critical bottlenecks in precise execution and structural robustness. The MMAE benchmark provides a clear diagnostic roadmap and establishes a standardized, long-lasting evaluation paradigm for next-generation audio editing systems, which can serve as a catalyst for future advances in the intelligent creation community.
📅 Published on Jun 5
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2606.07229
• PDF: https://arxiv.org/pdf/2606.07229
📊 Datasets citing this paper:
• https://huggingface.co/datasets/BoJack/MMAE
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#AudioEditingBenchmarks #MultimodalAudioProcessing #InstructionBasedAudioEditing #AudioTaxonomy #MultitaskLearningModels
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.