Mitigating bias against non-native accents

Zhang, Y.; Zhang, Yixuan; Halpern, B.M.; Patel, T.B.; Scharenborg, O.E.

doi:10.21437/Interspeech.2022-836

Mitigating bias against non-native accents

Journal article (2022)

Authors

Y. Zhang

Yixuan Zhang Student

B.M. Halpern Nederlands Kanker Instituut - Antoni van Leeuwenhoek ziekenhuis, Universiteit van Amsterdam,

T.B. Patel

O.E. Scharenborg

DOI: https://doi.org/10.21437/Interspeech.2022-836

Speech recognition Data augmentation Voice conversion Domain adversarial training Bias mitigation

To reference this document use:

http://resolver.tudelft.nl/uuid:3968d19a-1ae4-41d4-9491-e088c389da24

More Info

expand_more

Published Date

2022

Language

English

Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Automatic speech recognition (ASR) systems have seen substantial improvements in the past decade; however, not for all speaker groups. Recent research shows that bias exists against different types of speech, including non-native accents, in state-of-the-art (SOTA) ASR systems. To attain inclusive speech recognition, i.e., ASR for everyone irrespective of how one speaks or the accent one has, bias mitigation is necessary. Here we focus on bias mitigation against non-native accents using two different approaches: data augmentation and by using more effective training methods. We used an autoencoder-based cross-lingual voice conversion (VC) model to increase the amount of non-native accented speech training data in addition to data augmentation through speed perturbation. Moreover, we investigate two training methods, i.e., fine-tuning and domain adversarial training (DAT), to see whether they can use the limited non-native accented speech data more effectively than a standard training approach. Experimental results show that VC-based data augmentation successfully mitigates the bias against non-native accents for the SOTA end-to-end (E2E) Dutch ASR system. Combining VC and speed perturbed data gave the lowest word error rate (WER) and the smallest bias against nonnative accents. Fine-tuning and DAT reduced the bias against non-native accents but at the cost of native performance.

Files

Zhang22n_interspeech.pdf

(pdf | 0.506 Mb)

- Embargo expired in 01-07-2023

Unknown license