AI Briefing
KO

Big Bench Audio Benchmark Released

·2024.12.20 09:00

Key point

A new dataset called Big Bench Audio, designed to evaluate the reasoning capabilities of audio language models, has been released.

Details

Artificial Analysis has released the Big Bench Audio dataset to precisely evaluate the reasoning capabilities of audio language models. This dataset was constructed by converting questions from Big Bench Hard, a high-difficulty reasoning test, into the audio domain.

The dataset includes a total of 1,000 audio questions across 4 categories: logical reasoning (Formal Fallacies), path navigation (Navigate), object counting (Object Counting), and boolean logic (Web of Lies). The questions were generated using a variety of synthetic voices, with a focus on measuring the gap between text-based performance and audio-based performance.

Key analysis revealed a 'speech reasoning gap', a performance drop observed when models switch modalities. GPT-4o recorded a high accuracy of 92% in text-based tests, but its performance dropped sharply to 66% in the Speech-to-Speech setting.

This benchmark verifies from multiple angles how well models retain reasoning ability during speech input/output processes, through 4 settings: Speech-to-Speech, Speech-to-Text, Text-to-Speech, and Text-to-Text.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.