New research explores DNN denoising and Mandarin tone perception
A new study from PARC Shanghai, Shanghai Jiao Tong University, and Sonova researchers in Switzerland explores whether DNN denoising can preserve Mandarin lexical tone perception in noise.
For people who speak Mandarin Chinese, pitch is not just prosody. Changes in pitch can change the meaning of a word.
Mandarin has four lexical tones, distinguished primarily by differences in fundamental frequency (F0) contours, or the pattern of pitch change over time.
For example, a flat pitch represents Tone 1, while a rising pitch represents Tone 2. These differences can change the meaning of a word: the syllable yi, for example, can mean “one” with Tone 1 and “aunt” with Tone 2. Being able to accurately perceive these differences is therefore important for understanding spoken Mandarin.
For people with hearing loss, however, access to these cues can be affected. Previous research has shown that Mandarin speakers with mild-to-moderate sensorineural hearing loss can maintain categorical perception of lexical tones in quiet, although the precision of the boundaries between tones may be reduced.¹
Much less is known about what happens in noise, particularly for people using hearing aids. It is also unclear how DNN denoising may affect the F0 contours that Mandarin listeners rely on to distinguish lexical tones.
A collaborative study
A recently published study,¹ led by researchers at the Phonak Audiology Research Center (PARC) Shanghai, in collaboration with Professor Wang from Shanghai Jiao Tong University and Sonova researchers in Switzerland, investigated this question by examining the categorical perception of Mandarin Tones 1 and 2 in hearing aid users.
As a member of the PARC Shanghai research team and a co-author of the study, I was involved in evaluating whether deep neural network (DNN) denoising could help preserve categorical tone perception in noise. The results are encouraging, with the clearest effects emerging when individual hearing aid users were examined.
Why study categorical perception?
Although the F0 of a spoken sound can vary continuously, Mandarin listeners perceive these variations as distinct tone categories rather than a smooth gradient. This is referred to as ‘categorical perception’.
A familiar example from English:
The difference between bat and pat is partly signaled by voice onset time, or the brief interval between releasing the consonant and the onset of voicing. As that interval changes, listeners typically hear either b or p, rather than a gradual blend between the two.
In Mandarin, pitch can work in a similar way:
Mā (妈, “mother,” Tone 1) and má (麻, “hemp,” Tone 2) contain the same consonant and vowel. What changes is the F0 contour, and that change alone results in a different word meaning.
The researchers focused specifically on Tones 1 and 2. Because their F0 contours can overlap, these tones can be particularly difficult to distinguish in noise. This made them a useful test of whether hearing aid processing could preserve the acoustic information needed for tone identification.
That raises an important question for DNN denoising: can it separate speech from background noise without altering the F0 contours needed to identify Mandarin tones?
Study at a glance
- Participants: 20 adults with normal hearing and 20 adults with hearing loss were enrolled. One participant with hearing loss was excluded after not meeting the required accuracy during the Tone 1/Tone 2 training session, leaving 19 in the final analysis.
- Hearing aids: Participants with hearing loss were fitted with Phonak Audéo Sphere Infinio hearing aids and tested with DNN denoising enabled and disabled.
- Listening task: Participants identified stimuli along a nine-step continuum from a flat Tone 1 contour to a rising Tone 2 contour.
- Conditions: Testing took place in quiet and in cafeteria noise at 0 dB and −5 dB SNR.
- DNN comparison: In noise, hearing aid users were tested with DNN denoising enabled and disabled.
How were the results measured?
The two hearing aid programs were matched except for whether DNN denoising was enabled or disabled, allowing the researchers to isolate the contribution of DNN-based processing.
The researchers then examined the slope of the tone identification function. A steeper slope reflects a sharper boundary between Tone 1 and Tone 2. In other words, listeners shift more decisively from identifying a sound as Tone 1 to identifying it as Tone 2 as the F0 contour changes. A shallower slope reflects a less distinct boundary between the two categories.
What happened at the group level?
In quiet, hearing aid users with DNN off demonstrated categorical tone perception comparable to participants with normal hearing.
As the listening environment became more difficult, tone identification deteriorated. Across groups, identification slopes were steeper at 0 dB SNR than at −5 dB SNR.
In the overall analysis of the noise conditions, when DNN processing was disabled, tone identification slopes were significantly shallower than those observed in participants with normal hearing. With DNN processing enabled, slopes did not differ significantly from the normal-hearing group. However, the direct comparison between the DNN-on and DNN-off conditions was not statistically significant at the group level.
This led the researchers to look more closely at what was happening within individuals.
What happened at the individual level?
At 0 dB SNR, significantly more listeners showed categorical perception with DNN on than with it off.
- 9 participants demonstrated categorical perception with DNN both on and off.
- 8 demonstrated categorical perception with DNN enabled after not demonstrating it with DNN disabled.
- No participant showed the opposite pattern.
At the more challenging −5 dB SNR, however, the difference between DNN-on and DNN-off conditions was not statistically significant. The authors concluded that the effects of DNN processing may depend on the signal-to-noise ratio, with clearer benefits observed in the less challenging of the two noise conditions tested.
Preserving the cues that carry meaning
An important question in the study was not simply whether the DNN algorithm could reduce noise. Rather, the researchers investigated whether it could do so without compromising the F0 contours needed for Mandarin tone perception.
The authors describe this as a “do no harm” characteristic.
The authors interpret the findings as suggesting that DNN denoising facilitated perceptual separation of speech from cafeteria noise while preserving the F0 contours needed for tone identification. In doing so, DNN processing helped prevent the degradation of categorical perception observed with conventional processing in noise, without introducing harmful distortions to the cues carrying tone information.
What can we take from this study?
This study provides evidence specifically for the categorical perception of Mandarin Tones 1 and 2 under the conditions investigated. It should not be interpreted as demonstrating effects for all Mandarin speech or other tonal languages.
Within those boundaries, the findings provide new evidence that DNN denoising can support the maintenance of categorical identification of Mandarin tones for hearing aid users in noise. The study also addresses a gap identified by the authors, as categorical perception among hearing aid users, particularly in noise, had remained largely unexplored.
By examining this specific aspect of Mandarin speech perception, the study adds to our understanding of how advanced hearing aid processing interacts with acoustic cues that carry lexical meaning in Mandarin.
For hearing care professionals working with Mandarin-speaking hearing aid users, that is an important addition to the evidence base.
We invite you to read the full publication in the American Journal of Audiology.
Reference
1. Wang, Y., Tian, X., Du, X., Guan, J., Kuehnel, V., Launer, S. & Liu, C. (2026). Effects of deep neural network-enhanced hearing aids on categorical perception of Mandarin tones. American Journal of Audiology. doi:10.1044/2026_AJA-25-00294.
Co-authors

Jingjing Guan, Senior Director Innovation Center, Shanghai China
Jingjing Guan obtained her Ph.D. degree from University of Texas at Austin and worked at Texas Tech University Health Sciences Center as an assistant professor in Psychoacoustics. She was also a licensed audiologist and provided clinical services in the university clinic and public-school systems in Texas, USA. Jingjing started the new journey in Shanghai, China and has been working with an interdisciplinary team for Sonova Innovation Center since 2018. Her main research interests include technology development and clinical validation of amplification solutions for Mandarin speakers with hearing loss. She has a passion to drive innovation for people to hear better without barriers.

Lisa Bacic, Manager of Audiology Thought Leadership at Phonak HQ
Lisa is Global Manager of Audiology Thought Leadership and Editor-in-Chief of the Phonak Audiology Blog, where she works with experts across hearing care to translate evidence and professional insights into content for hearing care professionals around the world. A speech-language pathologist and audiologist by training, Lisa has more than 15 years of clinical experience and holds a Master’s degree in Speech Pathology and Audiology.
