By Topic

Cochannel speaker separation by harmonic enhancement and suppression

Sign In

Cookies must be enabled to login.After enabling cookies , please use refresh or reload or ctrl+f5 on the browser for the login options.

Formats Non-Member Member
$31 $13
Learn how you can qualify for the best price for this item!
Become an IEEE Member or Subscribe to
IEEE Xplore for exclusive pricing!
close button

puzzle piece

IEEE membership options for an individual and IEEE Xplore subscriptions for an organization offer the most affordable access to essential journal articles, conference papers, standards, eBooks, and eLearning courses.

Learn more about:

IEEE membership

IEEE Xplore subscriptions

4 Author(s)
Morgan, D. ; Signal Process. Center of Technol., Lockheed-Martin Inc., Nashua, NH, USA ; George, E.B. ; Lee, L.T. ; Kay, S.M.

This paper presents a system for separating the cochannel speech of two talkers. The proposed harmonic enhancement and suppression (HES) system is based on a frame-by-frame speaker separation algorithm that exploits the pitch estimate of the stronger talker derived from the cochannel signal. The idea behind this approach is to recover the stronger talker's speech by enhancing their harmonic frequencies and formants given a multiresolution pitch estimate. The weaker talker's speech is obtained from the residual signal created when the harmonics and formants of the stronger talker are suppressed. An automatic speaker assignment algorithm is used to place recovered frames from the target and interfering talkers in separate channels. Automatic speaker assignment performs reasonably well in most cochannel environments, including voiced-on-voiced, voiced-on-unvoiced, unvoiced-on-unvoiced, assignment after processing silence intervals, and single talker speech (no cochannel interference). The HES system has been tested at target-to-interferer ratios (TIRs) from -18 to 18 dB with widely available data bases. It has demonstrated improved performance in keyword spotting tests for TIR values of 6, 12, and 18 dB, and in human listening tests for TIR values of -6 and -18 dB

Published in:

Speech and Audio Processing, IEEE Transactions on  (Volume:5 ,  Issue: 5 )