Community

AES Convention Papers Forum

Modal Representations for Audio Deep Learning

Deep learning models for both discriminative and generative tasks have a choice of domain representation. For audio, candidates are often raw waveform data, spectral data, transformed spectral data, or perceptual features. For deep learning tasks related to modal synthesizers or processors, we propose new, modal representations for data. We experiment with representations such as an N-hot binary vector of frequencies, or learning a set of modal filterbank coefficients directly. We use these representations discriminatively–classifying cymbal model based on samples–as well as generatively. An intentionally naive application of a basic modal representation to a CVAE designed for MNIST digit images quickly yielded results, which we found surprising given less prior success when using traditional representations like a spectrogram image. We discuss applications for Generative Adversarial Networks, towards creating a modal reverberator generator.

Authors: Skare, Travis; Abel, Jonathan S.; Smith, III, Julius O.
Affiliation: CCRMA, Stanford University, Stanford, CA, USA
AES Convention: 147 (October 2019) Paper Number: 10248
Publication Date: October 8, 2019
Subject: Posters: Audio Signal Processing

Click to purchase paper as a non-member or you can login as an AES member to see more options.

No AES members have commented on this paper yet.

Subscribe to this discussion

To be notified of new comments on this paper you can subscribe to this RSS feed. Forum users should login to see additional options.

Start a discussion!

If you are not yet an AES member and have something important to say about this paper then we urge you to join the AES today and make your voice heard. You can join online today by clicking here.

Navigation

AES Convention Papers Forum

Modal Representations for Audio Deep Learning

Subscribe to this discussion

Start a discussion!

ABOUT AES

Contact Us