Hier finden Sie wissenschaftliche Publikationen aus den Fraunhofer-Instituten.

Bootstrapping a system for phoneme recognition and keyword spotting in unaccompanied singing

: Kruspe, Anna M.

Volltext (PDF; )

Mandel, M.I. ; International Society for Music Information Retrieval -ISMIR-:
17th International Society for Music Information Retrieval Conference, ISMIR 2016. Proceedings : New York City, United States, August 7-11, 2016
New York/NY, 2016
ISBN: 978-0-692-75506-8
International Society for Music Information Retrieval (ISMIR Conference) <17, 2016, New York/NY)
Konferenzbeitrag, Elektronische Publikation
Fraunhofer IDMT ()
vocal analysis; phoneme recognition; keyword spotting; lyrics alignment

Speech recognition in singing is still a largely unsolved problem. Acoustic models trained on speech usually produce unsatisfactory results when used for phoneme recognition in singing. On the flipside, there is no phonetically annotated singing data set that could be used to train more accurate acoustic models for this task. In this paper, we attempt to solve this problem using the DAMP data set which contains a large number of recordings of amateur singing in good quality. We first align them to the matching textual lyrics using an acoustic model trained on speech.
We then use the resulting phoneme alignment to train new acoustic models using only subsets of the DAMP singing data. These models are then tested for phoneme recognition and, on top of that, keyword spotting. Evaluation is performed for different subsets of DAMP and for an unrelated set of the vocal tracks of commercial pop songs.
Results are compared to those obtained with acoustic models trained on the TIMIT speech data set and on a version of TIMIT augmented for singing. Our new approach shows significant improvements over both.