ON THE CHOICE OF THE OPTIMAL TEMPORAL SUPPORT FOR AUDIO CLASSIFICATION WITH PRE-TRAINED EMBEDDINGS

Current state-of-the-art audio analysis systems rely on pretrained embedding models, often used off-the-shelf as (frozen) feature extractors. Choosing the best one for a set of tasks is the subject of many recent publications. However, one aspect often overlooked in these works is the influence of the duration of audio input considered to extract an embedding, which we refer to as Temporal Support (TS). In this work, we study the influence of the TS for well-established or emerging pre-trained embeddings, chosen to represent different types of architectures and learning paradigms. We conduct this evaluation using both musical instrument and environmental sound datasets, namely OpenMIC, TAU Urban Acoustic Scenes 2020 Mobile, and ESC-50. We especially highlight that Audio Spectrogram Transformer-based systems (PaSST and BEATs) remain effective with smaller TS, which therefore allows for a drastic reduction in memory and computational cost. Moreover, we show that by choosing the optimal TS we reach competitive results across all tasks. In particular, we improve the state-of-the-art results on OpenMIC, using BEATs and PaSST without any fine-tuning.

Mots clés

audio embeddings acoustic scene classification instrument recognition temporal support transformers Representation Model

Domaines

Intelligence artificielle [cs.AI] Son [cs.SD]

Fichier principal

Pre_print_ICASSP_Paper.pdf (444.79 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Aurian Quelennec : Connectez-vous pour contacter le contributeur

https://hal.science/hal-04360221

Soumis le : jeudi 21 décembre 2023-16:45:42

Dernière modification le : mardi 27 février 2024-09:56:21

Dates et versions

hal-04360221 , version 1 (21-12-2023)

Identifiants

HAL Id : hal-04360221 , version 1
ARXIV : 2312.14005

Citer

Aurian Quelennec, Michel Olvera, Geoffroy Peeters, Slim Essid. ON THE CHOICE OF THE OPTIMAL TEMPORAL SUPPORT FOR AUDIO CLASSIFICATION WITH PRE-TRAINED EMBEDDINGS. ICASSP, IEEE, Apr 2024, Séoul, South Korea. ⟨hal-04360221⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

INSTITUT-TELECOM LTCI IDS S2A IP_PARIS

106 Consultations

21 Téléchargements