{"title": "Extracting Speaker-Specific Information with a Regularized Siamese Deep Network", "book": "Advances in Neural Information Processing Systems", "page_first": 298, "page_last": 306, "abstract": "Speech conveys different yet mixed information ranging from linguistic to speaker-specific components, and each of them should be exclusively used in a specific task. However, it is extremely difficult to extract a specific information component given the fact that nearly all existing acoustic representations carry all types of speech information. Thus, the use of the same representation in both speech and speaker recognition hinders a system from producing better performance due to interference of irrelevant information. In this paper, we present a deep neural architecture to extract speaker-specific information from MFCCs. As a result, a multi-objective loss function is proposed for learning speaker-specific characteristics and regularization via normalizing interference of non-speaker related information and avoiding information loss. With LDC benchmark corpora and a Chinese speech corpus, we demonstrate that a resultant speaker-specific representation is insensitive to text/languages spoken and environmental mismatches and hence outperforms MFCCs and other state-of-the-art techniques in speaker recognition. We discuss relevant issues and relate our approach to previous work.", "full_text": "Extracting Speaker-Speci\ufb01c Information with a\n\nRegularized Siamese Deep Network\n\nKe Chen and Ahmad Salman\n\nSchool of Computer Science, The University of Manchester\n\nManchester M13 9PL, United Kingdom\n\n{chen,salmana}@cs.manchester.ac.uk\n\nAbstract\n\nSpeech conveys different yet mixed information ranging from linguistic to\nspeaker-speci\ufb01c components, and each of them should be exclusively used in a\nspeci\ufb01c task. However, it is extremely dif\ufb01cult to extract a speci\ufb01c information\ncomponent given the fact that nearly all existing acoustic representations carry\nall types of speech information. Thus, the use of the same representation in both\nspeech and speaker recognition hinders a system from producing better perfor-\nmance due to interference of irrelevant information. In this paper, we present a\ndeep neural architecture to extract speaker-speci\ufb01c information from MFCCs. As\na result, a multi-objective loss function is proposed for learning speaker-speci\ufb01c\ncharacteristics and regularization via normalizing interference of non-speaker re-\nlated information and avoiding information loss. With LDC benchmark corpora\nand a Chinese speech corpus, we demonstrate that a resultant speaker-speci\ufb01c rep-\nresentation is insensitive to text/languages spoken and environmental mismatches\nand hence outperforms MFCCs and other state-of-the-art techniques in speaker\nrecognition. We discuss relevant issues and relate our approach to previous work.\n\n1 Introduction\n\nIt is well known that speech conveys various yet mixed information where there are linguistic in-\nformation, a major component, and non-verbal information such as speaker-speci\ufb01c and emotional\ncomponents [1]. For human communication, all the information components in speech turn out to\nbe very useful and exclusively used for different tasks. For example, one often recognizes a speaker\nregardless of what is spoken for speaker recognition, while it is effortless for him/her to understand\nwhat is exactly spoken by different speakers for speech recognition. In general, however, there is no\neffective way to automatically extract an information component of interest from speech signals so\nthat the same representation has to be used in different speech information tasks. The interference\nof different yet entangled speech information components in most existing acoustic representations\nhinders a speech or speaker recognition system from achieving better performance [1].\nFor speaker-speci\ufb01c information extraction, two main efforts have been made so far; one is the use\nof data component analysis [2], e.g., PCA or ICA, and the other is the use of adaptive \ufb01ltering\ntechniques [3]. However, the aforementioned techniques either fail to associate extracted data com-\nponents with speaker-speci\ufb01c information as such information is non-predominant over speech or\nobtain features over\ufb01tting to a speci\ufb01c corpus since it is unlikely that speaker-speci\ufb01c information\nis statically resided in \ufb01xed frequency bands. Hence, the problem is still unsolved in general [4].\nRecent studies suggested that learning deep architectures (DAs) provides a new way for tackling\ncomplex AI problems [5]. In particular, representations learned by DAs greatly facilitate various\nrecognition tasks and constantly lead to the improved performance in machine perception [6]-[9]. On\nthe other hand, the Siamese architecture originally proposed in [10] uses supervised yet contrastive\n\n1\n\n\fFigure 1: Regularized Siamese deep network (RSDN) architecture.\n\nlearning to explore intrinsic similarity/disimilarity underlying an unknown data space. Incorporated\nby DAs, the Siamese architecture has been successfully applied to face recognition [11] and dimen-\nsionality reduction [12]. Inspired by the aforementioned work, we present a regularized Siamese\ndeep network (RSDN) to extract speaker-speci\ufb01c information from a spectral representation, Mel\nFrequency Cepstral Coef\ufb01cients (MFCCs), commonly used in both speech and speaker recognition.\nA multi-objective loss function is proposed for learning speaker-speci\ufb01c characteristics, normalizing\ninterference of non-speaker related information and avoiding information loss. Our RSDN learning\nadopts the famous two-phase deep learning strategy [5],[13]; i.e., greedy layer-wise unsupervised\nlearning for initializing its component deep neural networks followed by global supervised learning\nbased on the proposed loss function. With LDC benchmark corpora [14] and a Chinese corpus [15],\nwe demonstrate that a generic speaker-speci\ufb01c representation learned by our RSDN is insensitive\nto text and languages spoken and, moreover, applicable to speech corpora unseen during learning.\nExperimental results in speaker recognition suggest that a representation learned by the RSDN out-\nperforms MFCCs and that by the CDBN [9] that learns a generic speech representation without\nspeaker-speci\ufb01c information extraction. To our best knowledge, the work presented in this paper is\nthe \ufb01rst attempt on speaker-speci\ufb01c information extraction with deep learning.\nIn the reminder of this paper, Sect. 2 describes our RSDN architecture and proposes a loss function.\nSect. 3 presents a two-phase learning algorithm to train the RSDN. Sect. 4 reports our experimental\nmethodology and results. The last section discusses relevant issues and relates our approach to\nprevious work in deep learning.\n\n2 Model Description\n\nIn this section, we \ufb01rst describe our RSDN architecture and then propose a multi-objective loss\nfunction used to train the RSDN for learning speaker-speci\ufb01c characteristics.\n\n2.1 Architecture\n\nAs illustrated in Figure 1, our RSDN architecture consists of two subnets, and each subnet is a fully\nconnected multi-layered perceptron of 2K+1 layers, i.e., an input layer, 2K-1 hidden layers and a\nvisible layer at the top. If we stipulate that layer 0 is input layer, there are the same number of\nneurons in layers k and 2K-k for k = 0, 1,\u00b7\u00b7\u00b7 , K. In particular, the Kth hidden layer is used as\ncode layer, and neurons in this layer are further divided into two subsets. As depicted in Figure 1,\nthose neurons in the box named CS and colored in red constitute one subset for encoding speaker-\nspeci\ufb01c information and all remaining neurons in the code layer form the other subset expected to\n\n2\n\n1tx2tx\u02c61tx\u02c62txCSCS\u000b\f\u000b\f\u000b\f,;12DCSXCSX4\fwhere\n\n\u00b5(i) =\n\nTB(cid:88)\n\nt=1\n\nTB(cid:88)\n\nt=1\n\naccommodate non-speaker related information. The input to each subnet is an MFCC representation\nof a frame after a short-term analysis that a speech segment is divided into a number of frames and\nthe MFCC representation is achieved for each frame. As depicted in Figure 1, xit is the MFCC\nfeature vector of frame t in Xi, input to subnet i (i=1,2), where Xi = {xit}TB\nt=1 collectively denotes\nMFCC feature vectors for a speech segment of TB frames.\nDuring learning, two identical subsets are coupled at their coding layers via neurons in CS with an\nincompatibility measure de\ufb01ned on two speech segments of equal length, X1 and X2, input to two\nsubnets, which will be presented in 2.2. After learning, we achieve two identical subnets and hence\ncan use either of them to produce a new representation for a speech frame. For input x to a subnet,\nonly the bottom K layers of the subnet are used and the output of neurons in CS at the code layer or\nlayer K, denoted by CS(x), is its new representation, as illustrated by the dash box in Figure 1.\n\n2.2 Loss Function\nLet CS(xit) be the output of all neurons in CS of subnet i (i=1,2) for input xit \u2208 Xi and CS(Xi) =\n{CS(xit)}TB\nt=1, which pools output of neurons in CS for TB frames in Xi, as illustrated in Figure 1.\nAs statistics of speech signals is more likely to capture speaker-speci\ufb01c information [5], we de\ufb01ne\nthe incompatibility measure based on the 1st- and 2nd-order statistics of a new representation to be\nlearned as\n\nD[CS(X1),CS(X2); \u0398] = ||\u00b5(1) \u2212 \u00b5(2)||2\n\n2 + ||\u03a3(1) \u2212 \u03a3(2)||2\nF ,\n\n(1)\n\n1\nTB\n\nCS(xit), \u03a3(i) =\n\n1\n\nTB \u2212 1\n\n[CS(xit) \u2212 \u00b5(i)][CS(xit) \u2212 \u00b5(i)]T , i = 1, 2.\n\nTB(cid:88)\n\nt=1\n\nIn Eq. (1), || \u00b7 ||2 and || \u00b7 ||F are the L2 norm and the Frobenius norm, respectively. \u0398 is a collective\nnotation of all connection weights and biases in the RSDN. Intuitively, two speech segments belong-\ning to different speakers lead to different statistics and hence their incompatibility score measured\nby (1) should be large after learning. Otherwise their score is expected to be small.\nFor a corpus of multiple speakers, we can construct a training set so that an example be in the form:\n(X1, X2;I) where I is the label de\ufb01ned as I = 1 if two speech segments, X1 and X2, are spoken\nby the same speaker or I = 0 otherwise. Using such training examples, we apply the energy-based\nmodel principle [16] to de\ufb01ne a loss function as\n\nL(X1, X2; \u0398) = \u03b1[LR(X1; \u0398) + LR(X2; \u0398)] + (1 \u2212 \u03b1)LD(X1, X2; \u0398),\n\n(2)\n\n\u2212 Dm\n\n1\nTB\n\n\u03bbm + e\n\n\u2212 DS\n\n\u03bbS ).\n\n||xit \u2212 \u02c6xit||2\n\n2 and DS = ||\u03a3(1) \u2212 \u03a3(2)||2\n\n2 (i = 1, 2), LD(X1, X2; \u0398) = ID + (1 \u2212 I)(e\n\nwhere\nLR(Xi; \u0398) =\nHere Dm = ||\u00b5(1) \u2212 \u00b5(2)||2\nF . \u03bbm and \u03bbS are the tolerance bounds\nof incompatibility scores in terms of Dm and DS, which can be estimated from a training set. In\nLD(X1, X2; \u0398), we drop explicit parameters of D[CS(X1),CS(X2); \u0398] to simplify presentation.\nEq. (2) de\ufb01nes a multi-objective loss function where \u03b1 (0 < \u03b1 < 1) is a parameter used to trade-\noff between two objectives LR(Xi; \u0398) and LD(X1, X2; \u0398). The motivation for two objectives are\nas follows. By nature, both speaker-speci\ufb01c and non-speaker related information components are\nentangled over speech [1],[5]. When we tend to extract speaker-speci\ufb01c information, the interfer-\nence of non-speaker related information is inevitable and appears in various forms. LD(X1, X2; \u0398)\nmeasures errors responsible for wrong speaker-speci\ufb01c statistics on a representation learned by a\nSiamese DA in different situations. However, using LD(X1, X2; \u0398) only to train a Siamese DA\ncannot cope with enormous variations of non-speaker related information, in particular, linguistic\ninformation (a predominant information component in speech), which often leads to over\ufb01tting to\na training corpus according to our observations. As a result, we use LR(Xi; \u0398) to measure re-\nconstruction errors to monitor information loss during speaker-speci\ufb01c information extraction. By\nminimizing reconstruction errors in two subnets, the code layer leads to a speaker-speci\ufb01c represen-\ntation with the output of neurons in CS while the remaining neurons are used to regularize various\ninterference by capturing some invariant properties underlying them for good generalization.\nIn summary, we anticipate that minimizing the multi-objective loss function de\ufb01ned in Eq. (2) will\nenable our RSDN to extract speaker-speci\ufb01c information by encoding it through a generic speaker-\nspeci\ufb01c representation.\n\n3\n\n\f3 Learning Algorithm\n\n(cid:161)\n\nIn this section, we apply the two-phase deep learning strategy [5],[13] to derive our learning algo-\nrithm, i.e., pre-training for initializing subnets and discriminative learning for learning a speaker-\nspeci\ufb01c representation.\nWe \ufb01rst present the notation system used in our algorithm. Let hkj(xit) denote the output of the\njth neuron in layer k for k=0,1,\u00b7\u00b7\u00b7 ,K,\u00b7\u00b7\u00b7 ,2K. hk(xit) =\nj=1 is a collective notation\nof the output of all neurons in layer k of subnet i (i=1,2) where |hk| is the number of neurons\nin layer k. By this notation, k=0 refers to the input layer with h0(xit) = xit, and k=2K refers\nIn the coding layer, i.e., layer K, CS(xit) =\nto the top layer producing the reconstruction \u02c6xit.\nk denote the\nhKj(xit)\nconnection weight matrix between layers k-1 and k and the bias vector of layer k in subnet i (i=1,2),\nrespectively, for k=1,\u00b7\u00b7\u00b7 ,2K. Then output of layer k is hk(xit) = \u03c3[uk(xit)] for k=1,\u00b7\u00b7\u00b7 ,2K-1,\nwhere uk(xit) = W (i)\nj=1. Note that we use the\nlinear transfer function in the top layer, i.e., layer 2K, to reconstruct the original input.\n\n(cid:162)|CS|\nj=1 is a simpli\ufb01ed notation for output of neurons in CS. Let W(i)\n(cid:162)|z|\n\nk hk\u22121(xit) + b(i)\n\n(1 + e\u2212zj )\u22121\n\nk and \u03c3(z) =\n\n(cid:162)|hk|\n\nk and b(i)\n\nhkj(xit)\n\n(cid:161)\n\n(cid:161)\n\n3.1 Pre-training\n\nFor pre-training, we employ the denoising autoencoder [17] as a building block to initialize biases\nand connection weight matrices of a subnet. A denoising autoencoder is a three-layered perceptron\nwhere the input, \u02dcx, is a distorted version of the target output, x. For a training example, (\u02dcx, x), the\noutput of the autoencoder is a restored version, \u02c6x. Since MFCCs fed to the \ufb01rst hidden layer and\nits intermediate representation input to all other hidden layers are of continuous value, we always\ndistort input, x, by adding Gaussian noise to form a distorted version, \u02dcx. The restoration learning\nis done by minimizing the MSE loss between x and \u02c6x with respect to the weight matrix and biases.\nWe apply the stochastic back-propagation (SBP) algorithm to train denoising autoencoders, and\nthe greedy layer-wise learning procedure [5],[13] leads to initial weight matrices for the \ufb01rst K\nhidden layers, as depicted in a dash box in Figure 1, i.e., W1,\u00b7\u00b7\u00b7 , WK of a subnet. Then, we set\nK\u2212k+1 for k=1,\u00b7\u00b7\u00b7 ,K to initialize WK+1,\u00b7\u00b7\u00b7 , W2K of the subnet. Finally, the second\nWK+k = W T\nsubnet is created by simply duplicating the pre-trained one.\n\n3.2 Discriminative Learning\n\nFor discriminative learning, we minimizing the loss function in Eq. (2) based on pre-trained subnets\nfor speaker-speci\ufb01c information extraction. Given our loss function is de\ufb01ned on statistics of TB\nframes in a speech segment, we cannot update parameters until we have TB output of neurons in\nCS at the code layer. Fortunately, the SBP algorithm perfectly meets our requirement; In the SBP\nalgorithm, we always set the batch size to the number of frames in a speech segment. To simplify the\npresentation, we shall drop explicit parameters in our derivation if doing so causes no ambiguities.\nIn terms of the reconstruction loss, LR(Xi; \u0398), we have the following gradients. For layer k = 2K,\n\nFor all hidden layers, k=2K-1,\u00b7\u00b7\u00b7 ,1, applying the chain rule and (3) leads to\n\n\u2202LR\n\n\u2202u2K(xit)\n\n= 2(\u02c6xit \u2212 xit), i = 1, 2.\n\n(cid:181)\n\n\u2202LR\n\n=\n\n\u2202LR\n\nhkj(xit)[1\u2212hkj(xit)]\n\n\u2202LR\n\n,\n\n=\n\nW (i)\nk+1\n\n\u2202hkj(xit)\n\n\u2202uk(xit)\nAs the contrastive loss, LD(X1, X2; \u0398), de\ufb01ned on neurons in CS at code layers of two subnets, its\ngradients are determined only by parameters related to K hidden layers in two subnets, as depicted\nby dash boxes in Figure 1. For layer k=K and subnet i=1, 2, after a derivation (see the appendix for\ndetails), we obtain\n\n\u2202uk+1(xit)\n\n\u2202hk(xit)\n\nj=1\n\n(cid:163)\n\n(cid:164)T\n\n\u2202LR\n\n(3)\n\n. (4)\n\n(cid:182)|hk|\n\n4\n\n(cid:179)(cid:161)\n(cid:179)(cid:161)\n\n\u2202LD\n\n\u2202uK(xit)\n\n=\n\n[I \u2212 \u03bb\u22121\n[I \u2212 \u03bb\u22121\n\nm (1 \u2212 I)e\nS (1 \u2212 I)e\n\n\u2212 Dm\n\n\u03bbm ]\u03c8j(xit)\n\n\u2212 DS\n\n\u03bbS ]\u03bej(xit)\n\n(cid:161)\n(cid:161)\n\n(cid:162)|hK|\n(cid:162)|hK|\n\n0\n\n(cid:162)|CS|\n(cid:162)|CS|\n\nj=1,\n\nj=1,\n\n0\n\nj=|CS|+1\n\nj=|CS|+1\n\n(cid:180)\n(cid:180)\n\n.\n\n+\n\n(5)\n\n\fj\n\n(cid:164)\n\nj\n\nand \u03bej(xit) = qj(xit)\n\n(cid:161)CS(xit)\n(cid:162)\n\n1\u2212(cid:161)CS(xit)\n(cid:163)\n(cid:162)\n\nHere, \u03c8j(xit) = p(i)\nj\nwhere p(i) = 2\nTB\nand\n\n(cid:161)CS(xit)\n\n(cid:162)\nsign(1.5\u2212i)(\u00b5(1)\u2212\u00b5(2)), q(xit) = 4\n(cid:181)\nj is output of the jth neuron in CS for input xit. For layers k=K-1, \u00b7\u00b7\u00b7 ,1, we have\n\n1\u2212(cid:161)CS(xit)\n(cid:163)\n(cid:162)\n(cid:164)\nTB\u22121 sign(1.5\u2212i)(\u03a3(1)\u2212\u03a3(2))[CS(xit)\u2212\u00b5(i)]\n(cid:182)|hk|\nt=1;I(cid:162)\n\n\u2202hkj(xit)\nj=1\nGiven a training example,\n, we use gradients achieved from Eqs. (3)-(6) to\nupdate all the parameters in the RSDN. For layers k=K+1, \u00b7\u00b7\u00b7 , 2K, their parameters are updated by\n\n(cid:161)CS(xit)\n(cid:162)\n(cid:164)T\n(cid:163)\n2(cid:88)\nTB(cid:88)\n\n(cid:161){x1t}TB\n2(cid:88)\n\nhkj(xit)[1\u2212hkj(xit)]\n\nt=1,{x2t}TB\n\nTB(cid:88)\n\n\u2202uk+1(xit)\n\n\u2202uk(xit)\n\n\u2202hk(xit)\n\nW (i)\nk+1\n\n. (6)\n\n\u2202LD\n\n\u2202LD\n\n\u2202LD\n\n\u2202LR\n\n,\n\nj\n\n=\n\n=\n\n,\n\nj\n\nk \u2212 \u0001\u03b1\nTB\n\nk \u2212 \u0001\u03b1\nk \u2190 W (i)\nW (i)\nTB\n(cid:180)\nFor layers k=1, \u00b7\u00b7\u00b7 , K, their weight matrices and biases are updated with\n\n[hk\u22121(xrt)]T , b(i)\n\nk \u2190 b(i)\n\n\u2202uk(xrt)\n\n\u2202LR\n\nr=1\n\nt=1\n\nTB(cid:88)\n\n\u2202LR\n\n\u2202uk(xrt)\n\n. (7)\n\nt=1\n\nr=1\n\n(cid:179)\n2(cid:88)\nTB(cid:88)\n\nr=1\n\n\u2202LR\n\n\u2202uk(xrt)\n\n\u03b1\n\n2(cid:88)\n\n(cid:179)\n\nk \u2190 W (i)\nW (i)\n\nk \u2212 \u0001\nTB\nk \u2190 b(i)\nk \u2212 \u0001\nb(i)\nTB\n\nt=1\n\n+(1 \u2212 \u03b1)\n\n\u2202LD\n\n\u2202uk(xrt)\n\n[hk\u22121(xrt)]T ,\n\n(cid:180)\n\n.\n\n(8a)\n\n(8b)\n\n\u03b1\n\n\u2202LR\n\n\u2202uk(xrt)\n\n+(1 \u2212 \u03b1)\n\n\u2202LD\n\n\u2202uk(xrt)\n\nt=1\n\nr=1\n\nIn Eqs. (7) and (8), \u0001 is a learning rate. Here we emphasize that using sum of gradients caused by\ntwo subnets in update rules guarantees that two subsets are always kept identical during learning.\n\n4 Experiment\n\nIn this section, we describe our experimental methodology and report experiments results in visual-\nization of vowel distributions, speaker comparison and speaker segmentation.\nWe employ two LDC benchmark corpora [14], KING and TIMIT, and a Chinese speech corpus [15],\nCHN, in our experiments. KING, including wide-band and narrow-band sets, consists of 51 speakers\nwhose utterances were recorded in 10 sessions. By convention, its narrow-band set is called NKING\nwhile KING itself is often referred to its wide-band set. There are 630 speakers in TIMIT and 59\nspeakers in CHN of three sessions, respectively. All corpora were collected especially for evaluating\na speaker recognition system. The same feature extraction procedure is applied to all three corpora;\ni.e., after a short-term analysis suggested in [18], including silence removal with an energy-based\nmethod, pre-emphasis with the \ufb01lter H(z) = 1\u22120.95z\u22121 as well as Hamming windowing with the\nsize of 20 ms and 10 ms shift, we extract 19-order MFCCs [1] for each frame.\nFor the RSDN learning, we use utterances of all 49 speakers recorded in sessions 1 and 2 in KING.\nFurthermore, we distort all the utterances by the additive white noise channel with SNR of 10dB\nand the Rayleigh fading channel with 5 Hz Doppler shift [19] to simulate channel effects. Thus\nour training set consists of clean utterances and their corrupted versions. We randomly divide all\nutterances into speech segments of a length TB (1 sec\u2264 TB \u2264 2 sec) and then exhaustively combine\nthem to form training examples as described in Sect. 2.2. With a validation set of all the utterances\nrecorded in session 3 in KING, we select a structure of K=4 (100, 100, 100 and 200 neurons in\nlayers 1-4 and |CS|=100 in the code layer or layer 4) from candidate models of 2<K<5 and 50-\n1000 neurons in a hidden layer. Parameters used in our learning are as follows: Gaussian noise of\nN (0, 0.1\u03c3) used in denoising autoencoder, \u03b1=0.2, \u03bbm=100 and \u03bbS=2.5 in the loss function de\ufb01ned\nin Eq. (2), and learning rates \u0001=0.01 and 0.001 for pre-training and discriminative learning. After\nlearning, the RSDN is used to yield a 100-dimensional representation, CS, from 19-order MFCCs.\nFor any speaker recognition tasks, speaker modeling (SM) is inevitable. In our experiments, we use\nthe 1st- and 2nd-order statistics of a speech segment based on a representation, SM = {\u00b5, \u03a3}, for\nSM. Furthermore, we employ a speaker distance metric: d(SM1,SM2) = tr[(\u03a3\u22121\n2 )(\u00b51 \u2212\n\u00b52)(\u00b51 \u2212 \u00b52)T ], where SMi = {\u00b5i, \u03a3i} (i = 1, 2) are two speaker models (SMs). This distance\nmetric is derived from the divergence metric for two normal distributions [20] by dropping the term\nconcerning only covariance matrices based on our observation that covariance matrices often vary\nconsiderably for short segments and the original divergence metric often leads to poor performance\nfor various representations including MFCCs and ours. In contrast, the one de\ufb01ned above is stable\nirrespective of utterance lengths and results in good performance for different representations.\n\n1 + \u03a3\u22121\n\n5\n\n\fFigure 2: Visualization of all 20 vowels. (a) CS representation. (b) CS representation. (c) MFCCs.\n\n(a)\n\n(b)\n\n(c)\n\n4.1 Visualization\n\nVowels have been recognized to be a main carrier of speaker-speci\ufb01c information [1],[4],[18],[20].\nTIMIT [14] provides phonetic transcription of all 10 utterances containing all 20 vowels in English\nfor every speaker. As all the vowels may appear in 10 different utterances, up to 200 vowel segments\nin length of 0.1-0.5 sec are available for a speaker, which enables us to investigate vowel distributions\nin a representation space for different speakers. Here, we merely visualize mean feature vectors of\nup to 200 segments for a speaker in terms of a speci\ufb01c representation with the t-SNE method [21],\nwhich is likely to re\ufb02ect intrinsic manifolds, by projecting them onto a two-dimensional plane.\nIn the code layer of our RSDN, output of neurons 1-100 forms a speaker-speci\ufb01c representation, CS,\nand that of remaining 100 neurons becomes a non-speaker related representation, dubbed CS. For a\nnoticeable effect, we randomly choose only \ufb01ve speakers (four females and one male) and visualize\ntheir vowel distributions in Figure 2 in terms of CS, CS and MFCC representations, respectively,\nwhere a maker/color corresponds to a speaker. It is evident from Figure 2(a) that, by using the CS\nrepresentation, most vowels spoken by a speaker are tightly grouped together while vowels spoken\nby different speakers are well separated. For the CS representation, close inspection on Figure\n2(b) reveals that the same vowels spoken by different speakers are, to a great extent, co-located.\nMoreover, most of phonetically correlated vowels, as circled and labeled, are closely located in\ndense regions independent of speakers and genders. For comparison, we also visualize the same by\nusing their original MFCCs in Figure 2(c) and observe that most of phonetically correlated vowels\nare also co-located, as circled and labeled, whilst others scatter across the plane and their positions\nare determined mainly by vowels but affected by speakers. In particular, most of vowels spoken\nby the male, marked by (cid:164) and colored by green, are grouped tightly but isolated from those by all\nfemales. Thus, visualization in Figure 2 demonstrates how our RSDN learning works and could lend\nan evidence to justi\ufb01cation on why MFCCs can be used in both speech and speaker recognition [1].\n\n4.2 Speaker Comparison\n\nSpeaker comparison (SC) is an essential process involved in any speaker recognition tasks by com-\nparing two speaker models to collect evidence for decision-making, which provides a direct way to\nevaluate representations/speaker modeling without addressing decision-making issues [22]. In our\nSC experiments, we employ NKING [14], a narrow-band corpus, of many variabilities. During data\ncollection, there was a \u201cgreat divide\u201d between sessions 1-5 and 6-10; both recording device and en-\nvironments changed, which alters spectral features of 26 speakers and leads to 10dB SNR reduction\non average. As suggested in [18], we conduct two experiments: within-divide where SMs built\non utterances in session 1 are compared to SMs on those in sessions 2-5 and cross-divide where\nSMs built on utterances in session 1 are compared with those in sessions 6-10. As short utterances\nposes a greater challenge for speaker recognition [4],[18],[20], utterances are partitioned into short\nsegments of a certain length and SMs built on segments of the same length are always used for SC.\nFor a thorough evaluation, we apply the SM technique in question to our representation, MFCCs,\nand a representation (i.e., the better one of those yielded by two layers) learned by the CDBN [9]\non all 10 sessions in NKING, and name them SM-RSDN, SM-MFCC and SM-CDBN hereinafter. In\naddition, we also compare them to GMMs trained on MFCCs (GMM-MFCC), a state-of-the-art SM\ntechnique that provides the baseline performance [4],[20], where for each speaker a GMM-based\nSM consisting of 32 Gaussian components is trained on his/her utterances of 60 sec in sessions 1-2\nwith the EM algorithm [18]. For the CDBN learning [9] and the GMM training [18], we strictly\nfollow their suggested parameter settings in our experiments (see [9],[18] for details).\n\n6\n\n/aa/, /iy/, /aw/, /ay//ae/, /aw/, /iy/, /ix//iy/, /ih/, /eh/, /ix//ae/, /aa/, /aw/, /ay/\f(a)\n\n(b)\n\n(c)\n\nFigure 3: Performance of speaker comparison (DET) in the within-divide (upper row) and the cross-\ndivide (lower row) experiments for different segment lengths. (a) 1 sec. (b) 3 sec. (d) 5 sec.\nTable 1: Performance (mean\u00b1std)% of speaker segmentation on TIMIT and CHN audio streams.\nIndex\n\nTIMIT Audio Stream\n\nCHN Audio Stream\n\nBIC-MFCC Dist-MFCC Dist-RSDN BIC-MFCC Dist-MFCC Dist-RSDN\n\nFAR\nMDR\nF1\n\n26\u00b109\n26\u00b114\n67\u00b112\n\n22\u00b111\n22\u00b112\n74\u00b111\n\n18\u00b111\n18\u00b110\n79\u00b109\n\n46\u00b104\n46\u00b110\n44\u00b108\n\n27\u00b111\n27\u00b117\n68\u00b117\n\n24\u00b111\n24\u00b117\n72\u00b117\n\nWe use Detection Error Trade-off (DET) curves as the performance index in SC. From Figure 3, it is\nevident that SM-RSDN outperforms SM-MFCC, SM-CDBN and GMM-MFCC, a baseline system\ntrained on much longer utterances, as it always yields a smaller operating region, i.e., all possible\nerrors, in all the settings. In contrast, SM-MFCC performs better in within-divide settings while\nSM-CDBN is always inferior to the baseline system. Relevant issues will be discussed later on.\n\n4.3 Speaker Segmentation\n\nSpeaker segmentation (SS) is a task of detecting speaker change points in an audio stream to split\nit into acoustically homogeneous segments so that every segment contains only one speaker [23].\nFollowing the same protocol used in previous work [23], we utilize utterances in TIMIT and CHN\ncorpora to simulate audio conversations. As a result, we randomly select 250 speakers from TIMIT\nto create 25 audio streams where the duration of speakers ranges from 1.6 to 7.0 sec and 50 speakers\nfrom CHN to create 15 audio streams where the duration of speakers is from 3.0 to 8.3 sec. In the\nabsence of prior knowledge, the distance-based and the BIC techniques are two main approaches\nto SS [23]. In our simulations, we apply the distance-based method [23] to our representation and\nMFCCs, dubbed Dist-RSDN and Dist-MFCC, where the same parameters, including sliding window\nof 1.5 sec and tolerance level of 0.5 sec, are used. In addition, we also apply the BIC method [23] to\nMFCCs (BIC-MFCC). Note that the BIC method is inapplicable to our representation since it uses\nonly covariance information but the high dimensionality of our representation and the use of a small\nsliding window in the BIC result in unstable performance, as pointed out early in this section.\nFor evaluation, we use three common indexes [23], i.e., False Alarm Rate (FAR), Miss Detection\nRate (MDR) and F1 measure de\ufb01ned based on both precision and recall rates. Moreover, we only\nreport results as FAR equals MDR to avoid addressing decision-making issues [23]. Table 1 tabulates\nSS performance where, as boldfaced, results by our representation are superior to those by MFCCs\nregardless of SS methods and corpora for creating audio streams used in our simulations.\nIn summary, visualization of vowels and results in SC and SS suggest that our RSDN successfully\nextracts speaker-speci\ufb01c information; its resultant representation can be generalized to unseen cor-\npora during learning and is insensitive to text and languages spoken and environmental changes.\n\n7\n\n0.10.20.40.60.810.10.20.40.60.81False Alarm ProbabilityMiss ProbabilitySM\u2212MFCCGMM\u2212MFCCSM\u2212CDBNSM\u2212RSDN0.10.20.40.60.810.10.20.40.60.81False Alarm ProbabilityMiss ProbabilitySM\u2212MFCCGMM\u2212MFCCSM\u2212CDBNSM\u2212RSDN0.10.20.40.60.810.10.20.40.60.81False Alarm ProbabilityMiss ProbabilitySM\u2212MFCCGMM\u2212MFCCSM\u2212CDBNSM\u2212RSDN0.10.20.40.60.810.10.20.40.60.81False Alarm ProbabilityMiss ProbabilitySM\u2212MFCCGMM\u2212MFCCSM\u2212CDBNSM\u2212RSDN0.10.20.40.60.810.10.20.40.60.81False Alarm ProbabilityMiss ProbabilitySM\u2212MFCCGMM\u2212MFCCSM\u2212CDBNSM\u2212RSDN0.10.20.40.60.810.10.20.40.60.81False Alarm ProbabilityMiss ProbabilitySM\u2212MFCCGMM\u2212MFCCSM\u2212CDBNSM\u2212RSDN\f5 Discussion\n\nAs pointed out earlier, speech carries different yet mixed information and speaker-speci\ufb01c informa-\ntion is minor in comparison to predominant linguistic information. Our empirical studies suggest\nthat our success in extracting speaker-speci\ufb01c information is attributed to both unsupervised pre-\ntraining and supervised discriminative learning with a contrastive loss. In particular, the use of data\nregularization in discriminative learning and distorted data in two learning phases plays a critical\nrole in capturing intrinsic speaker-speci\ufb01c characteristics and variations caused by miscellaneous\nmismatches. Our results not reported here, due to limited space, indicate that without the pre-\ntraining in Sect. 3.1, a randomly initialized RSDN leads to unstable performance often considerably\nworse than that of using the pre-training in general. Without discriminative learning, a DA working\non unsupervised learning only, e.g., the CDBN [9], tends to yield a new representation that redis-\ntributes different information but does not highlight minor speaker-speci\ufb01c information given the\nfact that the CDBN trained on all 10 sessions in NKING leads to a representation that fails to yield\nsatisfactory SC performance on the same corpus but works well for various audio classi\ufb01cation tasks\n[9]. If we do not use the regularization term, LR(Xi; \u0398), in the loss function in Eq. (2), our RSDN\nis boiled down to a standard Siamese architecture [10]. Our results not reported here show that such\nan architecture learns a representation often over\ufb01tting to the training corpus due to interference\nof predominant non-speaker related information, which is not a problem in predominant informa-\ntion extraction. The previous work in face recognition [11] could lend an evidence to support our\nargument where a Siamese DA without regularization successfully captures predominant identity\ncharacteristics from facial images as, we believe, facial expression and other non-identity informa-\ntion are minor in this situation. While the use of distorted data in pre-training is in the same spirit of\nself-taught learning [24], our results including those not reported here reveal that the use of distorted\ndata in pre-training but not in discriminative learning yields results worse than the baseline perfor-\nmance in the cross-divide SC experiment. Hence, suf\ufb01cient training data re\ufb02ecting mismatches are\nalso required in discriminative learning for speaker-speci\ufb01c information extraction.\nOur RSDN architecture resembles the one proposed in [12] for dimensionality reduction of hand-\nwritten digits via learning a nonlinear embedding. However, ours distinguishes from theirs in the\nuse of different building blocks in our DAs, loss functions and motivations. The DA in [12] uses\nthe RBM [13] as a building block to construct a deep belief subnet in their Siamese DA and the\nNCA [25] as their contrastive loss function to minimize the intra-class variability. However, the\nNCA does not meet our requirements as there are so many examples in one class. Instead we pro-\npose a contrastive loss to minimize both intra- and inter-class variabilities simultaneously. On the\nother hand, intrinsic topological structures of a handwritten digit convey predominant information\ngiven the fact that without using the NCA loss a deep belief autoencoder already yields a good rep-\nresentation [7],[12],[13],[26]. Thus, the use of the NCA in [12] simply reinforces the topological\ninvariance by minimizing other variabilities with a small amount of labeled data [12]. In our work,\nhowever, speaker-speci\ufb01c information is non-predominant in speech and hence a large amount of la-\nbeled data re\ufb02ecting miscellaneous variabilities are required during discriminative learning despite\nthe pre-training. Finally, our code layer yields an overcomplete representation to facilitate non-\npredominant information extraction. In contrast, a parsimonious representation seems more suitable\nfor extracting predominant information since dimensionality reduction is likely to discover \u201cprinci-\npal\u201d components that often associate with predominant information, as are evident in [11],[12].\nTo conclude, we propose a deep neural architecture for speaker-speci\ufb01c information extraction and\ndemonstrate that its resultant speaker-speci\ufb01c representation outperforms the state-of-the-art tech-\nniques. It should also be stated that our work presented here is limited to speech corpora available\nat present. In our ongoing work, we are employing richer training data towards learning a univer-\nsal speaker-speci\ufb01c representation. In a broader sense, our work presented in this paper suggests\nthat speech information component analysis (ICA) becomes critical in various speech information\nprocessing tasks; the use of proper speech ICA techniques would result in task-speci\ufb01c speech rep-\nresentations to improve their performance radically. Our work demonstrates that speech ICA is\nfeasible via learning. Moreover, deep learning could be a promising methodology for speech ICA.\n\nAcknowledgments\n\nAuthors would like to thank H. Lee for providing their CDBN code [9] and L. Wang for offering\ntheir SIAT Chinese speech corpus [15] to us; both of which were used in our experiments.\n\n8\n\n\fReferences\n\n[1] Huang, X., Acero, A. & Hon, H. (2001) Spoken Language Processing. New York: Prentice Hall.\n[2] Jang, G., Lee, T. & Oh, Y. (2001) Learning statistically ef\ufb01cient feature for speaker recognition. Proc.\nICASSP, pp. I427-I440, IEEE Press.\n[3] Mammone, R., Zhang, X. & Ramachandran, R. (1996) Robust speaker recognition: a feature-based ap-\nproach. IEEE Signal Processing Magazine, 13(1): 58-71.\n[4] Reynold, D. & Campbell, W. (2008) Text-independent speaker recognition. In J. Benesty, M. Sondhi and\nY. Huang (Eds.), Handbook of Speech Processing, pp. 763-781, Berlin: Springer.\n[5] Bengio, Y. (2009) Learning deep architectures for AI. Foundation and Trends in Machine Learning 2(1):\n1-127.\n[6] Hinton, G. (2007) Learning multiple layers of representation. Trends in Cognitive Science 11(10): 428-434.\n[7] Larochelle, H., Bengio, Y., Louradour, J. & Lamblin, P. (2009) Exploring strategies for training deep neural\nnetworks. Journal of Machine Learning Research 10(1): pp. 1-40.\n[8] Boureau, Y., Bach, F., LeCun, Y. & Ponce, J. (2010) Learning mid-level features for recognition. Proc.\nCVPR, IEEE Press.\n[9] Lee, H., Largman, Y., Pham, P. & Ng, A. (2009) Unsupervised feature learning for audio classi\ufb01cation using\nconvolutional deep belief networks. In Advances in Neural Information Processing Systems 22, Cambridge,\nMA: MIT Press.\n[10] Bromley, J., Guyon, I., LeCun, Y., Sackinger, E. & Shah, R. (1994) Signature veri\ufb01cation using a Siamese\ntime delay neural network. In Advances in Neural Information Processing Systems 5, Morgan Kaufmann.\n[11] Chopra, S., Hadsell, R. & LeCun, Y. (2005) Learning a similarity metric discriminatively, with application\nto face veri\ufb01cation. In Proc. CVPR, IEEE Press.\n[12] Salakhutdinov, R. & Hinton, G. (2007) Learning a non-linear embedding by preserving class neighborhood\nstructure. In Proc. AISTATS, Cambridge, MA: MIT Press.\n[13] Hinton, G., Osindero, S. & Teh, Y. (2006) A fast learning algorithm for deep belief nets. Neural Compu-\ntation 18(7): 1527-1554.\n[14] Linguistic Data Consortium (LDC). [online] www.ldc.upenn.edu\n[15] Wang, L. (2008) A Chinese speech corpus for speaker recognition. Tech. Report, SIAT-CAS, China.\n[16] LeCun, Y., Chopra, S. Hadsell, R., Ranzato, M. & Huang, F. (2007) Energy-based models. In Predicting\nStructured Outputs, pp. 191-246, Cambridge, MA: MIT Press.\n[17] Vincent, P., Bengio, Y. & Manzagol, P. (2008) Extracting and composing robust features with denoising\nautoencoders. Proc. ICML, pp. 1096-1102, ACM Press.\n[18] Reynolds, D. (1995) Speaker Identi\ufb01cation and veri\ufb01cation using Gaussian mixture speaker models.\nSpeech Communication 17(1): 91-108.\n[19] Proakis, J. (2001) Digital Communications (4th Edition). New York: McGraw-Hill.\n[20] Campbell, J. (1997) Speaker recognition: A tutorial. Proceedings of The IEEE 85(10): 1437-1462.\n[21] van der Maaten, L. & Hinton, G. (2008) Visualizing data using t-SNE. Journal of Machine Learning\nResearch 9: 2579-2605.\n[22] Campbell, W. & Karam, Z. (2009) Speaker comparison with inner product discriminant functions. In\nAdvances in Neural Information Processing Systems 22, Cambridge, MA: MIT Press.\n[23] Kotti, M., Moschou, V. & Kotropoulos, C. (2008) Speaker segmentation and clustering. Signal Processing\n88(8): 1091-1124.\n[24] Raina, R., Battle, A., Lee, H., Packer, B. & Ng, A. (2007) Self-taught learning: transfer learning from\nunlabeled data. Proc. ICML, ACM press.\n[25] Goldberger, J., Roweis, S., Hinton, G. & Salakhutdinov, R., (2005) Neighbourhood component analysis.\nIn Advances in Neural Information Processing Systems 17, Cambridge, MA: MIT Press.\n[26] Hinton, G. & Salakhutdinov, R. (2006) Reducing the dimensionality of data with neural networks. Science\n313: 504-507.\n\n9\n\n\f", "award": [], "sourceid": 216, "authors": [{"given_name": "Ke", "family_name": "Chen", "institution": null}, {"given_name": "Ahmad", "family_name": "Salman", "institution": null}]}