Community

AES Convention Papers Forum

Multi-Input Architecture and Disentangled Representation Learning for Multi-Dimensional Modeling of Music Similarity

In the context of music information retrieval, similarity-based approaches are useful for a variety of tasks that benefit from a query-by-example approach. Music however, naturally decomposes into a set of semantically meaningful factors of variation. Current representation learning strategies pursue the disentanglement of such factors from deep representations, and result in highly interpretable models. This allows to model the perception of music similarity, which is highly subjective and multi-dimensional. While the focus of prior work is on metadata driven similarity, we suggest to directly model the human notion of multi-dimensional music similarity. To achieve this, we propose a multi-input deep neural network architecture, which simultaneously processes mel-spectrogram, CENSchromagram and tempogram representations in order to extract informative features for different disentangled musical dimensions: genre, mood, instrument, era, tempo, and key. We evaluated the proposed music similarity approach using a triplet prediction task and found that the proposed multi-input architecture outperforms a state of the art method. Furthermore, we present a novel multi-dimensional analysis to evaluate the influence of each disentangled dimension on the perception of music similarity.

Authors: Ribecky, Sebastian; Abeßer, Jakob; Lukashevich, Hanna
Affiliation: Semantic Music Technologies Group, Fraunhofer IDMT, Ilmenau, Germany
AES Convention: 152 (May 2022) Paper Number: 10568
Publication Date: May 2, 2022
Subject: Sound Classification

Click to purchase paper as a non-member or you can login as an AES member to see more options.

No AES members have commented on this paper yet.

Subscribe to this discussion

To be notified of new comments on this paper you can subscribe to this RSS feed. Forum users should login to see additional options.

Start a discussion!

If you are not yet an AES member and have something important to say about this paper then we urge you to join the AES today and make your voice heard. You can join online today by clicking here.

Navigation

AES Convention Papers Forum

Multi-Input Architecture and Disentangled Representation Learning for Multi-Dimensional Modeling of Music Similarity

Subscribe to this discussion

Start a discussion!

ABOUT AES

Contact Us