1. Combining spectral and temporal modification techniques for speech intelligibility enhancement
- Author
-
Maria Luisa Garcia Lecumberri, Martin Cooke, Vincent Aubanel, Ikerbasque - Basque Foundation for Science, GIPSA - Perception, Contrôle, Multimodalité et Dynamiques de la parole (GIPSA-PCMD), Département Parole et Cognition (GIPSA-DPC), Grenoble Images Parole Signal Automatique (GIPSA-lab ), Institut polytechnique de Grenoble - Grenoble Institute of Technology (Grenoble INP )-Institut Polytechnique de Grenoble - Grenoble Institute of Technology-Centre National de la Recherche Scientifique (CNRS)-Université Grenoble Alpes [2016-2019] (UGA [2016-2019])-Institut polytechnique de Grenoble - Grenoble Institute of Technology (Grenoble INP )-Institut Polytechnique de Grenoble - Grenoble Institute of Technology-Centre National de la Recherche Scientifique (CNRS)-Université Grenoble Alpes [2016-2019] (UGA [2016-2019])-Grenoble Images Parole Signal Automatique (GIPSA-lab ), Institut polytechnique de Grenoble - Grenoble Institute of Technology (Grenoble INP )-Institut Polytechnique de Grenoble - Grenoble Institute of Technology-Centre National de la Recherche Scientifique (CNRS)-Université Grenoble Alpes [2016-2019] (UGA [2016-2019])-Institut polytechnique de Grenoble - Grenoble Institute of Technology (Grenoble INP )-Institut Polytechnique de Grenoble - Grenoble Institute of Technology-Centre National de la Recherche Scientifique (CNRS)-Université Grenoble Alpes [2016-2019] (UGA [2016-2019]), Language and Speech Laboratory (LasLab), Universidad del Pais Vasco / Euskal Herriko Unibertsitatea [Espagne] (UPV/EHU), European Project: 339152,EC:FP7:ERC,ERC-2013-ADG,SPEECH UNIT(E)S(2014), and Universidad del País Vasco, Vitoria, Spain
- Subjects
Computer science ,Speech recognition ,Speech sounds ,020206 networking & telecommunications ,02 engineering and technology ,Intelligibility (communication) ,[SCCO.LING]Cognitive science/Linguistics ,01 natural sciences ,[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL] ,Theoretical Computer Science ,Human-Computer Interaction ,Energetic masking ,0103 physical sciences ,0202 electrical engineering, electronic engineering, information engineering ,010301 acoustics ,Speech rate ,[SPI.SIGNAL]Engineering Sciences [physics]/Signal and Image processing ,Software ,ComputingMilieux_MISCELLANEOUS - Abstract
Modifying clean speech prior to output in noisy conditions can lead to substantial intelligibility gains. Most algorithms operate by redistributing energy across the signal, leaving the timing of the underlying speech sounds intact. Other techniques do alter the timing of speech relative to the masker. Both classes of approach – spectral and temporal – lead to a reduction in energetic masking. The current study examines how their combination affects intelligibility. Arguments can be made for both synergy and redundancy, and the presence of distortions introduced by both spectral and temporal approaches might even lead to an antagonistic combination. A cohort of native Spanish listeners identified keywords in sentences in unmodified form and following spectral, temporal and spectro-temporal modification, in the presence of a fluctuating masker. Errors in the spectro-temporal condition were substantially lower than following spectral or temporal modification alone, with a three-fold reduction compared to unmodified speech. Spectro-temporal gains were observed for all phonemes. A glimpse-based model of energetic masking incorporating speech rate changes predicts intelligibility ( r = . 96 ), and a glimpsing analysis provides further insights into the distinct mechanisms through which spectral and temporal approaches lead to a release from energetic masking.
- Published
- 2019
- Full Text
- View/download PDF