Source enhanced linear prediction of speech incorporating simultaneously masked spectral weighting

Authors

DOI:

https://doi.org/10.26636/jtit.2001.3.72

Keywords:

linear prediction, psychoacoustics, masking, LPC

Abstract

Linear prediction is the cornerstone of most modern speech compression algorithms. This paper proposes modifying the calculation of the linear predictor coefficients to incorporate a weighting function based on the simultaneous masking property of the ear. The resultant prediction filter better models the perceptual characteristics of the source and results in the removal of more perceptually important information from the input speech signal than a standard LP filter. When employed in a low rate speech codec the net effect is an improvement in subjective quality, with no increase in transmission rate and only a modest increase in computational complexity.

Downloads

Download data is not yet available.

References

[1] B. C. J. Moore, An Introduction to the Psychology of Hearing. Sydney: Academic Press, 1997.
View in Google Scholar

[2] H. W. Strube, "Linear prediction on a warped frequency scale", J. Acoust. Soc. Am., vol. 68, no. 4, pp. 1071-1076, 1980.
View in Google Scholar

[3] Y. Nakatoh, T. Norimatsu, A. Heng Low, and H. Matsumoto, "Low bit rate coding for speech and audio using mel linear predictive coding (MLPC) analysis", in Proc. ICSLP, 1998.
View in Google Scholar

[4] H. Hermansky, "Perceptual linear predictive analysis of speech", J. Acoust. Soc. Am., vol. 87, no. 4, pp. 1738-1753, 1990.
View in Google Scholar

[5] "MPEG4", ISO/IEC FCD 14496-3.
View in Google Scholar

[6] L. B. Rabiner and R. W. Schafer, Digital Processing of Speech Signals. New Jersey: Prentice Hall, 1978.
View in Google Scholar

[7] A. M. Kondoz, Digital Speech. New York: Wiley, 1995.
View in Google Scholar

[8] J. Makhoul and J. Wolf, "Linear prediction and the spectral analysis of speech", BBN report, no. 2304, August 1972.
View in Google Scholar

[9] J. Makhoul, "Linear prediction: a tutorial review", Proc. IEEE, vol. 63, pp. 561-580, 1975.
View in Google Scholar

[10] B. Scharf, "Critical bands", in Foundations of Modern Auditory Theory, J. Tobias, Ed. New York: Academic Press, 1970, pp. 159-202.
View in Google Scholar

[11] M. Schroeder and B. S. Atal, "Predictive coding of speech signals and subjective error criteria", in IEEE Trans. ASSP, 1979, pp. 247-254.
View in Google Scholar

[12] P. Kroon and E. F. Deprettere, "A class of analysis by synthesis predictive coders for high quality speech coding at rates between 4.8 and 16 kbits/s", IEEE J. Selec. Areas Commun., vol. 62, pp. 353-363, 1988.
View in Google Scholar

[13] D. Sen, D. H. Irving, and W. H. Holmes, "PERCELP - perceptually enhanced random codebook excited linear prediction", in Proc. IEEE W/shop Speech Cod. Telecommun., 1993, pp. 101-102.
View in Google Scholar

[14] I. S. Burnett, "Hybrid techniques for speech coding", Ph.D. thesis, University of Bath, 1992.
View in Google Scholar

[15] J. D. Johnston, "Transform coding of audio signals using perceptual noise criteria", IEEE J. Selec. Areas Commun., vol. 6, pp. 314-323, 1988.
View in Google Scholar

[16] J. Lukasiak and I. S. Burnett, "Exploiting simultaneously masked linear prediction in a WI speech coder", in Proc. IEEE W/shop Speech Cod., 2000, pp. 11-13.
View in Google Scholar

[17] J. G. Proakis and D. G. Manolakis, Digital Signal Processing. New Jersey: Prentice Hall, 1996.
View in Google Scholar

[18] National Communication System, details to assist in implementation of Federal Standard 1016 CELP, Office of the manager National Communication System, Arlington.
View in Google Scholar

[19] G. S. Kang and L. J. Fransen, "Low-bit rate speech encoders based on line spectrum frequencies (LSFs)", NRL report, no. 8857, Naval Research Lab., Washington D.C., Jan. 1985.
View in Google Scholar

[20] M. Yong, G. Davidson, and A. Gersho, "Encoding of LPC spectral parameters using switched adaptive inter frame vector prediction", in Proc. ICASSP, 1988, vol. 1, pp. 402-405.
View in Google Scholar

[21] K. K. Paliwal and B. S. Atal, "Efficient vector quantisation of LPC parameters at 24 bits/frame", IEEE Trans. Speech Audio Proc., vol. 1, no. 1, pp. 3-14, 1993.
View in Google Scholar

[22] W. B. Kleijn and J. Haagen, "A speech coder based on decomposition of characteristic waveforms", in Proc. ICASSP, 1995, vol. 1, pp. 508-511.
View in Google Scholar

[23] J. Skoglund and W. B. Kleijn, "On time frequency masking in voiced speed", IEEE Trans. Speech Audio Proc., vol. 8, no. 4, pp. 361-369, 2000.
View in Google Scholar

[24] J. Lukasiak, I. S. Burnett, J. F. Chicharo, and M. M. Thomson, "Linear prediction incorporating simultaneous masking", in Proc. ICASSP, 2000, vol. 3, pp. 1471-1474.
View in Google Scholar

Downloads

Submitted

2023-04-07

Published

2001-09-30

Issue

Section

ARTICLES FROM THIS ISSUE

How to Cite

[1]
J. Lukasiak and I. S. Burnett, “Source enhanced linear prediction of speech incorporating simultaneously masked spectral weighting”, JTIT, vol. 5, no. 3, pp. 15–23, Sep. 2001, doi: 10.26636/jtit.2001.3.72.

Most read articles by the same author(s)