DOI 10.14209/its.2002.817
Charles B. de Lima, Abraham Alcaim, José A. Apolinário Jr.
"This paper presents a performance evaluation of two classification systems for text independent speaker verification: the Gaussian Mixture Model (GMM) and the AR-Vector Model. For the GMM, 32, 16, 8 and Gaussians are evaluated. On the other hand, an order 2 model with the Itakura symmetric distance was used for the AR-Vector. Both classification systems...
DOI 10.14209/its.2002.822
Rodrigo C. de Lamare, Abraham Alcaim
"In this paper we examine the fricatives and stops encoding of a very low bit rate speech compression algorithm, based on a mixed multiband excitation system. The algorithm incorporates several improvements over previously reported coders. One of them is the use of a specific modelling and synthesis strategy for fricatives and stops at 400 b/s. The...
DOI 10.14209/its.2002.825
Rodrigo C. de Lamare, Abraham Alcaim
"In this paper we analyse postfiltering techniques for very low bit rate speech coders in tandem connections. A mixed multiband excitation (MMBE) linear predictive coding (LPC) algorithm, that encodes voiced frames at 1.75 kb/s and unvoiced frames at 0.4 kb/s, is employed to assess the performance of different postfilters in tandem connections. We...
DOI 10.14209/its.2002.829
Dirceu Gonzaga da Silva, José A. Apolinário Jr., Charles B. de Lima
"This paper presents a modification in a technique of channel normalization widely known as Cepstral Mean Subtraction (CMS). This modification is based on the introduction of language dependent phonetic modification. A careful investigation using Brazilian Portuguese was carried out showing that it is possible to improve the CMS channel identification...
DOI 10.14209/its.2002.834
Jayme G. A. Barbedo, Moisés V. Ribeiro, Amauri Lopes, João M. T. Romano
"This paper deals with the application of the Kohonen Self-Organizing Maps (KSOM) to methods of objective speech quality assessment. The performance of the objective methods so far proposed depends on many factors, among which the required mapping between the objective and subjective domains is one of the most important. The purpose of this paper is to...
DOI 10.14209/its.2002.840
Lucas M. J. Barbosa, Luís G. Meloni
"With the objective of reducing computational complexity in algebraic code-excited linear predictive (ACELP) coders, this paper describes the use of a low-complexity sequential search algorithm with signalselected pulse amplitudes. The described algorithm was inserted in both the ETSI GSM-AMR and ITU-T G.729 codecs, and when compared to some standard...
DOI 10.14209/its.2002.844
Carla Lopes, Fernando Perdigão
"This article describes a vocal tract length normalization (VTLN) procedure through pitch based frequency warping. This procedure aim to reduce de inter-speaker variability, present in speech signals. It is also described a method for coarticulation phenomena compensation, that reduce speech signal variability due to phonetic context. This procedure...
DOI 10.14209/its.2002.850
R. S. Maia, R. J. R. Cirigliano, D. Rojtenberg, F. G. V. Resende Jr.
"Vector quantization of the synthesis filter parameters in Code-Excited Linear Prediction (CELP) speech coders is a common procedure nowadays. This paper describes a CELP coder implementation and makes a comparison in terms of quality and bit rate when vector and scalar quantization of the synthesis filter parameters are employed.Usage of vector...
DOI 10.14209/its.2002.855
Fabian Vargas, Rubem D. R. Fagundes, Daniel Barros Jr.
"Hereafter, we present a new approach dealing to couple with the harmful effects of noise on speech recognition systems (SRS). This approach is oriented to hardware redundancy and it is essentially a modification of the classic software-based recovery blocks scheme. When compared to conventional approaches using Fast Fourier Transform (FFT) and Hamming...
Digital Signal Processing (DSP); Speech-Recognition Systems (SRS); Noise Immunity; On-Line Testing; Software-Based Recovery Blocks (SBRB); Area overhead; Performance Degradation
DOI 10.14209/its.2002.861
Francisco J. Fraga
"It is possible to implement a speech-to\u2014text system with unlimited vocabulary by connecting two subsystems: A phoneme recognizer, which is performed by sub-syllabic segmenting the incoming speech, and a phonologic-graphemic converter. This paper presents an automatic speech recognition system with these features. The segmentation method and the...