{"title": "Non-linear Prediction of Acoustic Vectors Using Hierarchical Mixtures of Experts", "book": "Advances in Neural Information Processing Systems", "page_first": 835, "page_last": 842, "abstract": null, "full_text": "Non-linear Prediction of Acoustic Vectors \nUsing Hierarchical Mixtures of Experts \n\nS.R.Waterhouse \nA.J.Robinson \nCambridge University Engineering Department, \nTrumpington St ., Cambridge, CB2 1PZ, England. \n\nTel: [+44] 223 332800, Fax: [+44] 223 332662, \n\nEmail: srwlO01.ajr@eng.cam.ac.uk \n\nURL: http://svr-www.eng.cam.ac.ukr srw1001 \n\nAbstract \n\nIn this paper we consider speech coding as a problem of speech \nmodelling. In particular, prediction of parameterised speech over \nshort time segments is performed using the Hierarchical Mixture of \nExperts (HME) (Jordan & Jacobs 1994). The HME gives two ad(cid:173)\nvantages over traditional non-linear function approximators such \nas the Multi-Layer Percept ron (MLP); a statistical understand(cid:173)\ning of the operation of the predictor and provision of information \nabout the performance of the predictor in the form of likelihood \ninformation and local error bars. These two issues are examined \non both toy and real world problems of regression and time series \nprediction. In the speech coding context, we extend the principle \nof combining local predictions via the HME to a Vector Quantiza(cid:173)\ntion scheme in which fixed local codebooks are combined on-line \nfor each observation. \n\n1 INTRODUCTION \n\nWe are concerned in this paper with the application of multiple models, specifi(cid:173)\ncally the Hierarchical Mixtures of Experts, to time series prediction, specifically the \nproblem of predicting acoustic vectors for use in speech coding. There have been \na number of applications of multiple models in time series prediction. A classic \nexample is the Threshold Autoregressive model (TAR) which was used by Tong & \n\n\f836 \n\nS. R. Waterhouse, A. J. Robinson \n\nLim (1980) to predict sunspot activity. More recently, Lewis, Kay and Stevens \n(in Weigend & Gershenfeld (1994)) describe the use of Multivariate and Regres(cid:173)\nsion Splines (MARS) to the prediction of future values of currency exchange rates. \nFinally, in speech prediction, Cuperman & Gersho (1985) describe the Switched \nInter-frame Vector Prediction (SIVP) method which switches between separate lin(cid:173)\near predictors trained on different statistical classes of speech. The form of time \nseries prediction we shall consider in this paper is the single step prediction fI(t) of a \nfuture quantity y(t) , by considering the previous c: samples. This may be viewed as \na regression problem over input-output pairs {x t), y(t)}~ where x(t) is the lag vec(cid:173)\ntor (y(t-I), y(t-2), ... , y(t- p\u00bb. We may perform this regression using standard linear \nmodels such as the Auto-Regressive (AR) model or via nonlinear models such as \nconnectionist feed-forward or recurrent networks. The HME overcomes a number of \nproblems associated with traditional connectionist models via its architecture and \nstatistical framework. Recently, Jordan & Jacobs (1994) and Waterhouse & Robin(cid:173)\nson (1994) have shown that via the EM algorithm and a 2nd order optimization \nscheme known as Iteratively Reweighted Least Squares (IRLS), the HME is faster \nthan standard Multilayer Perceptrons (MLP) by at least an order of magnitude on \nregression and classification tasks respectively. Jordan & Jacobs also describe var(cid:173)\nious methods to visualise the learnt structure of the HME via 'deviance trees' and \nhistograms of posterior probabilities. In this paper we provide further examples \nof the structural relationship of the trained HME and the input-output space in \nthe form of expert activation plots. In addition we describe how the HME can be \nextended to give local error bars or measures of confidence in regression and time \nseries prediction problems. Finally, we describe the extension of the HME to acous(cid:173)\ntic vector prediction, and a VQ coding scheme which utilises likelihood information \nfrom the HME. \n\n2 HIERARCHICAL MIXTURES OF EXPERTS \n\nThe HME architecture (Figure 1) is based on the principle of 'divide and conquer' \nin which a large, hard to solve problem is broken up into many, smaller, easier \nto solve problems. It consists of a series of 'expert networks' which are trained \non different parts of the input space. The outputs of the experts are combined \nby a 'gating network' which is trained to stochastically select the expert which is \nperforming best at solving a particular part of the problem. The operation of the \nHME is as follows: the gating networks receive the input vectors x(t) and produce \nas outputs probabilities P(mi/.x(t), 7'/j) for each local branch mj of assigning the current \ninput to the different branches, where T/j are the gating network parameters. The \nexpert networks sit at the leaves of the tree and each output a vector flJt) given \ninput vector x(t) and parameters Bj . These outputs are combined in a weighted sum \nby P(mjlX<t), T/j) to give the overall output vector for this region. This procedure \ncontinues recursively upwards to the root node. \nIn time series prediction, each \nexpert j is a linear single layer network with the form: \n\nflY) = B; x (t) \n\nwhere B; is matrix and x(t) is the lag vector discussed earlier, which is identical in \nform to an AR model. \n\n\fNon-Linear Prediction of Acoustic Vectors Using Hierarchical Mixtures of Experts \n\n837 \n\nx \n\nx \n\nx \n\nx \n\nFigure 1: The Hierarchical Mixture of Experts. \n\n2.1 Error bars via HME \n\nSince each expert is an AR model, it follows that the output of each expert y(t) is \nthe expected value of the observations y(t) at each time t. The conditional likelihood \nof yet) given the input and expert mj is \n\nP(y(t) I x (t), mj, Bj) = 12:Cj I exp ( - ~ (y - yy\u00bb)T Cj(y - yjt))) \n\nwhere Cj is the covariance matrix for expert mj which is updated during training as: \n\nC = _1_ \"'\" h(t)(y(r) _ y(t)l (y(t) _ y~t)) \nJ \n\n] \n\n] \n\n' \" h~t) L..J J \nL.Jt J \n\nt \n\nwhere hy) are the posterior probabilities I of each expert mj' Taking the moments of \nthe overall likelihood of the HME gives the output of the HME as the conditional \nexpected value of the target output yct), \nyet) = E(yct)lxct), 0, M) \n\n= 2: P(mjlxct), l1j)E(y(t)lx ct), ej,mj) = 2: gY)iJ/t), \n\nWhere M represents the overall HME model and e the overall set of parameters. \nTaking the second central moment of yct) gives, \n\nj \n\nj \n\nC = E\u00aby(t) - yy\u00bb)2I xct), 0, M) \n\n= 2: P(mJlx(t), l1j)E\u00aby(t) - yjt))2I x (t), ej, mj) \n= 2: gjt)(Cj + yjt). iJj(t)T), \n\nj \n\nj \n\nlSee (Jordan & Jacobs 1994) for a fuller discussion of posterior probabilities and like(cid:173)\n\nlihoods in the context of the HME. \n\n\f838 \n\nS. R. Waterhouse, A. J. Robinson \n\nwhich gives, in a direct fashion, the covariance of the output given the input and the \nmodel. If we assume that the observations are generated by an underlying model, \nwhich generates according to some function f(x(t)) and corrupted by zero mean \nnormally distributed noise n(x) with constant covariance 1:, then the covariance of \ny(t) is given by, \n\nV(y(t)) = V(to) + 1:, \n\nso that the covariance computed by the method above, V(y(t)) , takes into account \nthe modelling error as well as the uncertainty due to the noise. Weigend & Nix \n(1994) also calculate error bars using an MLP consisting of a set of tanh hidden \nunits to estimate the conditional mean and an auxiliary set of tanh hidden units \nto estimate the variance, assuming normally distributed errors. Our work differs \nin that there is no assumption of normality in the error distribution, rather that \nthe errors of the terminal experts are distributed normally, with the total error \ndistribution being a mixture of normal distributions. \n\n3 SIMULATIONS \n\nIn order to demonstrate the utility of our approach to variance estimation we con(cid:173)\nsider one toy regression problem and one time series prediction problem. \n\n3.1 Toy Problem: Computer generated data \n\n2,------..--...., \n\n0. 08,------~-..., \n\n1.5 \n\n~-0.5 \n\n-1 \n\n-1.5 \n\n-2 \n\n-2.5 \n\nN O.OS \n~ \nfO.04 \n~ \n-0.02 \n\n0.8 \nQ) go.S \n.~ \n~0.4 \n\n2 \n\nx \n\n-0.2 \n\n-0.4 \n\n-O.S \n\n-0.8 \n\no \n\n2 \n\nx \n\n2 \n\nx \n\n-3 '------'------' \n\n-1 '------'-----\" \n\no \n\n2 \n\nx \n\nFigure 2: Performance on the toy data set of a 5 level binary HME. (a) training set \n(dots) and underlying function f(x) (solid), (b) underlying function (solid) and prediction \ny(x) (dashed), (c) squared deviation of prediction from underlying function, (d) true noise \nvariance (solid) and variance of prediction (dashed). \n\nBy way of comparison, we used the same toy problem as Weigend & Nix (1994) \nwhich consists of 1000 training points and 10000 separate evaluation points from \n\n\fNon-Linear Prediction of Acoustic Vectors Using Hierarchical Mixtures of Experts \n\n839 \n\nthe function g(x) where g(x) consists of a known underlying function f(x) corrupted \nby normally distributed noise N(O, (J2(X)) , \n\nf(x) = sin(2.5x) x sin(l. 5x), \n\n(J2(x) = 0.01 + O. 25 x [1 - sin(2. 5x)f. \n\nAs can be seen by Figure 2, the HME has learnt to approximate both the underlying \nfunction and the additive noise variance. The deviation of the estimated variance \nfrom the \"true\" noise variance may be due to the actual noise variance being lower \nthan the maximum denoted by the solid line at various points. \n\n3.2 Sunspots \n\n1960 \n\n1970 \n\n1980 \n\n1920 \n\n1930 \n\n1940 \n\n8 \n\n6 -\n8.4 - -\n.-\n2 - . \u2022 \n\n\u2022 \n-\n\n)( w \n\nt \n\n\u2022 \n\n.. \n\u2022 \n\n\u2022 \n\n1950 \nYear \n\nI I \n\n.. \n.. \n\u2022 \n- l-\n-\n\u2022 \u2022 \u2022 \u2022 \n\n\u2022 \n\n\u2022 \n\nI \u2022 \n\n-.. \n--\n... \n\u2022 \n\u2022 \u2022 \n-\n-\n\n1950 \nYear \n\nOL-------~--------~--------~--------L-------~--------~ \n1920 \n\n1940 \n\n1970 \n\n1930 \n\n1960 \n\n1980 \n\nFigure 3: Performance on the Sunspots data set. (a) Actual Values (x) and predicted \nvalues (0) with error bars. (b) Activation of the expert networks; bars wide in the vertical \naxis indicate strong activation. Notice how expert 7 concentrates on the lulls in the series \nwhile expert 2 deals with the peaks. \n\nI METHOD I \n\nNMSE ' \n\nTrain \n\nTest \n\n1700-1920 \n\n1921-1955 1956-1979 \n\nMLP \nTAR \nHME \n\n0.082 \n0.097 \n0.061 \n\n0.086 \n0.097 \n0.089 \n\n0.35 \n0.28 \n0.27 \n\nTable 1: Results of single step prediction on the Sunspots data set using a mixture of 7 \nexperts (104 parameters) and a lag vector of 12 years. NMSE' is the NMSE normalised \nby the variance of the entire record 1700 to 1979. \n\n\f840 \n\nS. R. Waterhouse, A. J. Robinson \n\nThe Sunspots2 time series consists of yearly sunspot activity from 1700 to 1979 and \nwas first tackled using connectionist models by Weigend, Huberman & Rumelhart \n(1990) who used a 12-8-1 MLP (113 parameters) . Prior to this work, the TAR was \nused by Tong (1990). Our results, which were obtained using a random leave 10% \nout cross validation method, are shown in Table 1. We are considering only single \nstep prediction on this problem, which involves prediction of the next value based \non a set of previous values of the time series. Our results are evaluated in terms of \nNormalised Mean Squared Error (NMSE) (Weigend et al. 1990), which is defined \nas the ratio of the variance of the prediction on the test set to the variance of the \ntest set itself. \n\nThe HME outperforms both the TAR and the MLP on this problem, and addition(cid:173)\nally provides both information about the structure of the network after training via \nthe expert activation plot and error bars of the predictions, as shown in Figure 3. \nFurther improvements may be possible by using likelihood information during cross \nvalidation so that a joint optimisation of overall error and variance is achieved. \n\n4 SPEECH CODING USING HME \n\nIn the standard method of Linear Predictive Coding (LPC) (Makhoul 1975), speech \nis parametrised into a set of vectors of duration one frame (around 10 ms). Whilst \nsimple scalar quantization of the LPC vectors can achieve bit rates of around 2400 \nbits per second (bps), Yong, Davidson & Gersho (1988) have shown that simple \nlinear prediction of Line Spectral Pairs (LSP) (Soong & Juang 1984) vectors followed \nby Vector Quantization (VQ) (Abut, Gray & Rebolledo 1984) of the error vectors \ncan yield bit rates of around 800 bps. In this paper we describe a speech coding \nframework which uses the HME in two stages. Firstly, the HME is used to perform \nprediction of the acoustic vectors. The error vectors are then quantized efficiently \nby using a VQ scheme which utilises the likelihood information derived from the \nHME. \n\n4.1 Mixing VQ codebooks ia Gating networks \n\nIn a VQ scheme using a Euclidean distance measure, there is an implicit assumption \nthat the inputs follow a Gaussian probability density function (pdf). This is satisfied \nif we quantize the residuals from a linear predictor, but not the residuals from an \nHME which follow a mixture of Gaussians pdf. A more efficient method is therefore \nto generate separate VQ code books for each expert in the HME and combine them \nvia the priors on each expert from the gating networks. The code book for the \noverall residual vectors on the test set is then generated at each time dynamically \nby choosing the first D x gjt) codes, where D is the size of the expert codebooks and \ngY) is the prior on each expert. \n\n2 Available via anonymous \n\nDataSunspots.Yearly \n\nftp at \n\nfip.cs.colorado.edu \n\nIII \n\njpub jTime-Series as \n\n\fNon-Linear Prediction of Acoustic Vectors Using Hierarchical Mixtures of Experts \n\n841 \n\n4.2 Results of Speech Coding Evaluations \n\nInitial experiments were performed using 23 Mel scale log energy frequency bins as \nacoustic vectors and using single variances Cj = (J;I as expert network covariance \nmatrices. The results of training over 100 ,000 frames and evaluation over a further \n100,000 frames on the Resource Management (RM) corpus are shown in Table 2 \nand Figure 4 which shows the good specialisation of the HME in this problem. \n\nMETHOD \n\nLinear \n\n1 level HME \n2 level HME \n\nPrediction Gain (dB) \nTrain \n12.07 \n18.1 \n20.20 \n\nTest \n10.95 \n15.55 \n16.39 \n\nTable 2: Prediction of Acoustic Vectors using linear prediction and binary branching \nHMEs with 1 and 2 levels. Prediction gain (Cuperman & Gersho 1985) is the ratio of the \nsignal variance to prediction error variance. \n\n(a) Spectogram \n\n20 \n\n!!! \n::3 c;; 15 \nCf \n~ 10 \n8 < \n\n5 \n\n::3 \n\n. \n\no \n\n10 \n\n20 \n\n30 \n\n40 \n\n50 \n\nFrame \n\n60 \n\n70 \n\n80 \n\n90 \n\n100 \n\n(b) Gating Network Decisions \n\n-\n\n-\n\nI \n\n- - -__.-\n\nI \n71---\n6~----. . . . ~.---~--__ --------_=--~~--\u00ad\n~5~~---------1~---.I---1IHI~I~\u00b7--.... ~.~.-----\n~4~~~----------------~~~~----------\nw3~~---------~-----. . ~ __ ------_.----------------- -\n2.-.---- - --\nII \n1~--------------~~~~--------------------------------~ .. ...-\n\n- - -11-1...---11---1- --11-1 \u2022 \u2022 I \n\n\u2022 \n\nI \n\n10 \n\n20 \n\n30 \n\n40 \n\n50 \n\nFrame \n\n60 \n\n70 \n\n80 \n\n90 \n\n100 \n\nFigure 4: The behaviour of a mixture of 7 experts at predicting Mel-scale log energy \nfrequency bins over 100 16ms fram es. The top figure is a spectrogram of the speech and \nthe lower figure is an expert activation plot, showmg the gating network decisions. \n\nWe have conducted further experiments using LSPs and cepstrals as acoustic vec(cid:173)\ntors , and using diagonal expert network covariance matrices, on a very large speech \ncorpus. However, initial experiments show only a small improvement in gain over \na single linear predictor and further investigation is underway. We have also coded \nacoustic vectors using 8 bits per frame with frame lengths of 12.5 ms, passing power, \npitch and degree of voicing as side band information , without appreciable distor(cid:173)\ntion over simple LPC coding. A full system will include prediction of all acoustic \n\n\f842 \n\nS. R. Waterhouse, A. J. Robinson \n\nparameters and we anticipate further reductions on this initial figure with future \ndevelopments. \n\n5 CONCLUSION \n\nThe aim of speech coding is the efficient coding of the speech signal with little \nperceptual loss. This paper has described the use of the HME for acoustic vector \nprediction. We have shown that the HME can provide improved performance over a \nlinear predictor and in addition it provides a time varying variance for the prediction \nerror. The decomposition of the linear prediction problem into a solution via a \nmixture of experts also allows us to construct a VQ codebook on the fly by mixing \nthe codebooks of the various experts. \n\nWe expect that the direct computation of the time varying nature of the prediction \naccuracy will find many applications. Within the acoustic vector prediction problem \nwe would like to exploit this information by exploring the continuum between the \nfixed bit rate coder described here and a variable bit rate coder that produces \nconstant spectral distortion. \n\nAcknowledgements \n\nThis work was funded in part by Hewlett Packard Laboratories, UK. Steve Water(cid:173)\nhouse is supported by an EPSRC Research Students hip and Tony Robinson was \nsupported by a EPSRC Advanced Research Fellowship. \n\nReferences \n\nAbut, H., Gray, R. M. & Rebolledo, G. (1984), 'Vector quantization of speech and speech(cid:173)\n\nlike waveforms', IEEE Transactions on Acoustics, Speech, and Signal Processing. \n\nCuperman, V. & Gersho, A. (1985), 'Vector predictive coding of speech at 16 kblt/s', \n\nIEEE Transactions on Communications COM-33, 685-696. \n\nJordan, M. I. & Jacobs, R. A. (1994), 'Hierarchical Mixtures of Experts and the EM \n\nalgorithm', Neural Computation 6, 181-214. \n\n63(4) 561-580. \n\nMakhoiil, J. (1975), 'Linear prediction: A tutorial review', Proceedings of the IEEE \nSoong~ F. k. & Juang, B. H. (1984), Line spectrum pair (LSP) and speech data compres(cid:173)\nTong, H. (1990), Non-linear Time Series: a dynamical systems approach, Oxford Univer-\n\nsIon. \n\nsi~y Press. \n\nTong, H. & Lim, K. (1980), 'Threshold autoregression, limit cycles and cyclical data', \n\nJournal of Royal Statistical Society. \n\nWaterhouse, S. R. & Robinson, A. J. (1994), Classification using hierarchical mixtures of \n\nexperts, in 'IEEE Workshop on Neural Networks for SigI!al Processing'. \n\nWeigend, A. S. & Gershenfeld, N. A. (1994), Time Series Prediction: Forecasting the \n\nFuture and Understanding the Past) Addison-Wesley. \n\nWeigend, A. S. & Nix, D. A. (1994), Predictions with confidence intervals (local error bars), \nTechnical Report CU-CS-724-94, Department of Computer Science and Institute of \nCoznitive Science, University of Colorado, Boulder, CO 80309-0439. \n\nWeigend, A. S., Huberman, B. A. & Rumelhart, D. E. (1990), 'Predicting the future: a \n\nconnectionist approach', International Journal of Neural Systems 1, 193-209. \n\nYong, M., Davidson, G. & Gersho, A. (1988), Encoding of LPC spectral parameters using \nswitched-adaptive interframe vector prediction, in 'Proceedings of the IEEE Interna(cid:173)\ntional Conference on Acoustics Speech, and Signal Processing', pp. 402-405. \n\n\f", "award": [], "sourceid": 960, "authors": [{"given_name": "Steve", "family_name": "Waterhouse", "institution": null}, {"given_name": "Anthony", "family_name": "Robinson", "institution": null}]}