A very low bit rate speech coder based on a recognition/synthesis paradigm

Ki-Seung Lee Ki-Seung Lee,R.V Cox

doi:10.1109/89.928913

Abstract

Previous studies have shown that a concatenative speech synthesis system with a large database produces more natural sounding speech. We apply this paradigm to the design of improved very low bit rate speech coders (sub 1000 b/s). The proposed speech coder consists of unit selection, prosody coding, prosody modification and waveform concatenation. The encoder selects the best unit sequence from a large database and compresses the prosody information. The transmitted parameters include unit indices and the prosody information. To increase naturalness as well as intelligibility, two costs are considered in the unit selection process: an acoustic target cost and a concatenation cost. A rate-distortion-based piecewise linear approximation is proposed to compress the pitch contour. The decoder concatenates the set of units, and then synthesizes the resultant sequence of speech frames using the harmonic+noise model (HNM) scheme. Before concatenating units, prosody modification which includes pitch shifting and gain modification is applied to match those of the input speech. With single speaker stimuli, a comparison category rating (CCR) test shows that the performance of the proposed coder is close to that of the 2400-b/s MELP coder at an average bit rate of about 800-b/s during talk spurts.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

A very low bit rate speech coder based on a recognition/synthesis paradigm

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Speech and Audio Processing

Lead the way for us

Journal: IEEE Transactions on Speech and Audio Processing	Publication Date: Jul 1, 2001
Citations: 71

Similar Papers

A segmental speech coder based on a concatenative TTS
Ki-Seung Lee ... Richard V Cox
Speech Communication | VOL. 38
Ki-Seung Lee, et. al.Ki-Seung Lee ... Richard V Cox
20 Feb 2002
Speech Communication | VOL. 38

Very low bit rate (VLBR) speech coding around 500 bits/sec
...
-
, et. al. ...
06 Sep 2004
06 Sep 2004

Quantisation and PCM

-

11 Feb 2014
11 Feb 2014

TTS based very low bit rate speech coder
Ki-Seung Lee ... R.V Cox
-
Ki-Seung Lee, et. al. Ki-Seung Lee ... R.V Cox
01 Jan 1998
01 Jan 1998

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A very low bit rate speech coder based on a recognition/synthesis paradigm

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Speech and Audio Processing