{"title": "Fast Transformation-Invariant Factor Analysis", "book": "Advances in Neural Information Processing Systems", "page_first": 1287, "page_last": 1294, "abstract": "", "full_text": "Fast Transformation-Invariant Factor Analysis\n\nAnitha Kannan\n\nNebojsa Jojic\u0001\n\nBrendan Frey\n\n University of Toronto, Toronto, Canada\n\u0002 anitha, frey\n@psi.utoronto.ca\n\u0001 Microsoft Research, Redmond, WA, USA\n\njojic@microsoft.com\n\nAbstract\n\nDimensionality reduction techniques such as principal component analy-\nsis and factor analysis are used to discover a linear mapping between high\ndimensional data samples and points in a lower dimensional subspace.\nIn [6], Jojic and Frey introduced mixture of transformation-invariant\ncomponent analyzers (MTCA) that can account for global transforma-\ntions such as translations and rotations, perform clustering and learn lo-\ncal appearance deformations by dimensionality reduction. However, due\nto enormous computational requirements of the EM algorithm for learn-\ning the model, O(\nis the dimensionality of a data sample,\nMTCA was not practical for most applications. In this paper, we demon-\nstrate how fast Fourier transforms can reduce the computation to the or-\n. With this speedup, we show the effectiveness of MTCA\nder of\nin various applications - tracking, video textures, clustering video se-\nquences, object recognition, and object detection in images.\n\n\u0001 ) where\n\nlog\n\n1 Introduction\nDimensionality reduction techniques such as principal component analysis [7] and factor\nanalysis [1] linearly map high dimensional data samples onto points in a lower dimensional\nsubspace. In factor analysis, this mapping is de\ufb01ned by subspace origin\nand the subspace\n. A mixture of factor analyzers\nbases stored in the columns of the factor loading matrix,\nlearn to place the data into several learned subspaces. In computer vision, this approach\nhas been widely used in face modeling for learning facial expressions (e.g. [2] , [12] ).\n\nWhen the variability in the data is due, in part, to small transformations such as translations,\nscales and rotations, factor analyzer learns a linearized transformation manifold which is\noften suf\ufb01cient ( [4], [11]). However, for large transformations present in the data, linear\napproximation is insuf\ufb01cient. For instance, a factor analyzer trained on a video sequence\nof a person walking tries to capture a linearized model of large translations (\ufb01g. 2a.) as\nopposed to learning local deformations such as motion of legs and hands (\ufb01g. 2c.).\n\n, enables clustering\nIn [6], it was shown that a discrete hidden transformation variable\nand learning subspaces within clusters, invariant to global transformations. However, ex-\nperiments were done on images of very low resolution due to enormous computational cost\nof EM algorithm used for learning the model. It is known that fast Fourier transform(FFT)\n\n\u0003\n\u0004\n\u0004\n\u0004\n\u0004\n\u0005\n\u0006\n\u0007\n\fFigure 1: Mixture of transformed component analyzers (MTCA). (a) The generative model with\ncluster index c, subspace coordinates\nand\n+noise; (b) An example of the generative process, where subspace\ngenerated \ufb01nal image\ncoordinates\n\r\u0012\u0011\n\nare inferred from a captured video sequence\n\n\u0001\u0003\u0002\u0005\u0004\u0007\u0006\t\b\u000b\n\f\n\n+noise; transformation\n\n, and image position\n\n\u000e\u000f\u0002\u0010\r\u000b\u0001\n\n, latent image\n\n\u0012\u0013\n\n,\n\nis very useful in dealing with transformations in images ( [3], [13]). The main purpose of\nthis work is to show that under very mild assumptions, we can have an effective implemen-\ntation of MTCA that reduces the complexity from\nis\nthe number of factors, N is the size of input, \u0017\nis the set of all possible transformations.\nThis means that for 256x256 images, the current implementation will be 4000 times faster.\n\n\u0017\u001a\u0015 log\n\n, where\n\n\u0017\u0003\u0015\n\n\u0001\u0016\u0015\n\n\u0001\u0019\u0015\n\nto\n\n(MTCA)\n\ntransformation-invariant component analyzers\n\nWe present experimental results showing the effectiveness of MTCA in various applications\n- tracking, video textures, clustering video sequences, object recognition and detection.\n2 Fast inference and learning in MTCA\nMixture of\nis used for\ntransformation-invariant clustering and learning a subspace representation within each\ncluster. The set of transformations, \u0017\n, to which the model is invariant is speci\ufb01ed a priori.\ndimensional Gaus-\nFig. 1a. shows a generative model for MTCA. The vector\nsian \u0018\u001d\u001c 0\nrandom variable. Cluster index, c is a C-valued discrete random variable with\n)latent image,\nprobability distribution,\nhas mean,\n \"!\nis the factor loading matrix\n\u0006*!+\u001b\n!)(\n, (with dis-\nfor class c. An observation\n. Fig. 1b\ntribution\nillustrates this generative process for a one class MTCA. The subspace coordinates\n, are\n(without noise), and the horizontal and vertical image\nused to generate a latent image,\nposition\n. In fact,\n\u000798\n\u0007.6\nshown in the \ufb01gure are actually inferred from the captured video sequence\n(see sec. 3).\nThe joint distribution over all variables is [6],\n\ndimensional (\n,-!\nand adding independent Gaussian noise,\n\nare used to shift the latent image to obtain\n\nis obtained by applying a transformation\n\n, and diagonal covariance,\n\n\u0004$#%#&\u0014\nmatrix\n\u0006.!\n\n) on the latent image\n\n. The\n\n\u000710\n\n; the\n\nis a\n\nand\n\nand\n\n\u000798\n\nI\n\n243\n\n\u000776\n\nx\n\n,\n\n\u001b;\u001e=<>\u001e\f'?\u001e\n\n\u0007@\u001e\f/A\u001fCB\n\n0\n\n\u001b;\u001e\f<E\u001f\nI\n\u0018\u001d\u001c\n\n\u001bD\u001f\n\u0018\u001d\u001c\n\u001b;G\n\n'?G\n\n'?\u001e\n!)(\n\n\u00077\u001f+F\n\u00077\u001f+F\n\u0006.!\f\u001b;\u001eH,-!H\u001f\n\n<E\u001f\n\u0018\u001d\u001c\n\n/IG\n\n\u00079'?\u001eH5J\u001fK2\n\n L!\n\n\n\n\n\u000e\n\u0014\n\u0018\n\u0014\n\u0004\n\u0014\n\u001b\n\u0014\n\u001e\n\u001f\n\u0004\n'\n\u0005\n\u0004\n\u0014\n/\n\u0017\n'\n5\n\u001b\n'\n/\n\u001b\n/\n:\n\u001c\n:\n\u001c\n:\n\u001c\n'\n\u0015\n:\n\u001c\n/\n\u0015\n\u001c\n\u001c\nB\n\u001e\n\u001f\n\u0005\n3\n\fFigure 2: Means\nlearned using (a) FA, (b) FA applied on data normalized\nusing a correlation tracker, and (c) transformed component analysis (TCA) applied directly on data.\n\nand components in\n\nPerforming inference on transformations and class,\n\nrequires evaluating the joint,\n\n/A\u001f\n\n<>\u001e\n\n\u0006*!\n\n!\u0003(\n,.!H\u001f\n\n\u0006.!\n\n,.!H\u001f\n\n\u001fK2\n\n(1)\n\n(2)\n\nThe likelihood of\n\n/I\u001e=<>\u001e\n\n\u0007*\u001f\n\n\u00077\u001f+F\n\n<E\u001f+F\n\n\u0018\u001d\u001c\n\n/IG\n\n\u00077\u001f\n\u0018\u001d\u001c\n\n/IG\n\nis\n/A\u001f\n\n!\u0005\u0004\n\n3\u0007\u0006\t\b\n\n\u000b\u0010\u000f\n\n/A\u001f\n\n/\u000e\n\n\u001e=<>\u001e\n\n\f\u000b\n\n\u001b;\u001e='\n\n, the number of clusters,\n\nand M step in which it updates the parameters.\n\nThe parameters of the model are learned from i.i.d training examples by maximizing their\n) using an exact EM algorithm. The only inputs to the EM are the\nlikelihood (\ntraining examples, the number of factors,\n, and the set of all\npossible transformations, \u0017\n. Starting at a random initialization, EM algorithm for MTCA\niterates between E step, where it probabilistically \ufb01lls in for hidden variables by \ufb01nding the\nexact posterior :\nThe likelihood of the data (eqn. 2) requires summing over all possible transformations and\nis very expensive. In fact, each of inference and update equations in [6] has a complexity\n. In this section, we show how these equations can be derived and evaluated in\nof\nFourier domain at a considerably lower computational cost of\n. We focus on\nas examples for ef\ufb01cient computation.\ninferring the means of :\nSimilar mathematical manipulations will result in the inferences provided in the appendix.\nWe assume that data is represented in a coordinate system in which transformations are\ndiscrete shifts with wrap-around. For translations, it is 2D rectangular grid, while for ro-\ntations and scales it is shifts in a radial grid (c.f. [3] [13]). We also assume that the post-\ntransformation noise is isotropic,\n\u001e+/\u0018\u0017\n\u0007@\u001e\f<\nand COV\n, it is possible to\n, to small\npreset\nvalue, if the actual value in the data is larger, it can be accounted for in\n\n(in our experiments we set it to .001). By presetting the sensor noise,\n\nB\u0013\u0012\u0015\u0014\nbecome independent of\n\n, so that covariances matrices,COV\n\n. In fact, for isotropic\n\n\u0007@\u001e\f<>\u001e\f/\u0018\u0017\n\n\u0017\u001a\u0015 log\n\nand :\n\n\u001e\f<>\u001e\f/A\u001f\n\n\u0007@\u001e\f<\n\n\u001e+/A\u001f\n\n\u0001\u0016\u0015\n\n.\n\n\u0001 ),\nare de\ufb01ned this way. For diagonal matrices such as\n\nFirst, we describe the notation that simpli\ufb01es expressions for transformation that corre-\nto be an integer vector in the coordinate\nsponds to a shift in input coordinates. De\ufb01ne\nsystem in which input is measured. For 2D nxn image, (\nelement\nwhere\n. Vectors in the input coordinate system\nsuch as\nde\ufb01nes the ele-\n. This notation enables treating transformations\nment corresponding to pixel at coordinate\ncorresponding to a shift be represented as a vector\nin the same system, so that a shift of\n\u001c\u001d\u001f\n\u0007\u0005B\n\n/I\u001e\nis represented as\n\nB#\"%$&$'$(\u001aI\u001e\n\nB\u0013\"%$'$&$(\u001a\n\nsuch that\n\nis the\n\nB\u001b\u001a\n\nmod\n\nmod\n\n\u001c\u001d\u001f\n\u001e\f'\n\n\u001f! \n\n\u00077\u001f\n\n\u001aI\u001e\n\n\u0019\u001d\u001c\u001d\u001e\n\n\u001aA\u001f\n\nby\n\n,\n\n.\n\n(*)\n\n(*)\n\n\u0004\n\b\nF\n\u001c\n\u0007\n\u0015\nF\n\u001c\nB\nF\n\u001c\n/\n\u0015\n<\n\u001e\n\u001c\n\u001c\nB\n\u0007\n\u0005\n!\n\u001e\n\u0007\n\u001c\n\u0006\n\u0001\n\u0001\n(\n5\n\u001f\n2\n3\n \n!\n/\n:\n\u001c\nB\n\u0002\n\u0003\n\n\u0003\n\u0007\n\u0005\n!\n\u001e\n\u0007\n\u001c\n\u0006\n\u0001\n!\n(\n\u0001\n(\n5\n3\n \n!\n:\n\u001c\n\u001f\n\u0014\n\u0011\n\u001c\n\u0007\n\u0015\n\u0014\n\u0001\n\u0015\n\u0017\n\u0015\n\u0004\n\u0014\n\u0004\n\u001c\n\u001b\n\u0015\n\u001c\n'\n\u0015\n5\n\u0016\n'\n\u0015\n\u0016\n\u001b\n\u0015\n\u0007\n5\n\u0012\n\u0012\n,\n\u0019\n\u0004\n/\n\u001c\n\u0019\n\u001f\n\u0019\n0\n\u0002\n\n\u001e\n\u001f\n\u0001\n\u001f\n\n\u001f\n\u0001\n\u0003\n\u0005\n,\n,\n\u001c\n\u0019\n\u001f\n\u0019\n\u0007\n'\n\u0007\n'\n\u001c\n\u0019\n(\n\u0019\n(\n\n\n\u001f\n\u0001\n\u0001\n\fFigure 3: Transformation invariant clustering with and without a subspace model: (a) Parameters\nof a three-cluster TMG [6], and a three-cluster MTCA (b) Frames from the video sequence\n, cor-\nin the corresponding subspace of\nresponding TMG mean\nMTCA; (c) An illustration of the role of components for the \ufb01rst class. Factor\ntends to model\nlighting variation and\n\ntends to model small out-of-plane rotations\n\nand the object appearance\n\n\u0001\n\n\u0003\u0002\n\nIn the appendix, we show that all expensive operations in inference and learning in-\n\n\u0017\u001a\u0015\u000e\r\u0010\u000f\u0012\u0011\n\nfor all shifts in frequency domain, while it is \f\n\nvolve computing correlation (\u0004\u0006\u0005\b\u0007 ), or convolution(\u0004\n\t\u000b\u0007 ). These operations is only\n\u001c\f\u0015\nthe pixel domain. For notational ease, we represent \u0013\n\u0014J\u001f\n\r\u0016\u0015\u0018\u0017\nand \u0004\u001d\u001c\u001e\u0007 de\ufb01nes an element wise product between \u0004 and \u0007\nIn principal component analysis (PCA), where there is no noise, the data is projected to\nsubspace through the projection matrix. Similarly, in MTCA, we can derive that when\n\nin\n, by\n) extracts the diagonal elements of matrix\n\n\u0010\u000f\u0012\u0011\nrespectively. Also, diag(\u0014\n\n\u0017\u001a\u0015\n\u001c\f\u0015\ncolumn and row of a matrix, \u0014\n\n, i.e. \f\n\u0014J\u001f\n\nand \u001c\n\n\u001a\u0019\u001b\u0017\n\n\u001c\u001d\u001e\n\n.\n\nvariances. %\n\nof subspace for a given\nimage\n\nand\n\nB#\"$ \n\n\u00067!&%! and it accounts for\n\nis the inverse of noise variances in input space, and \"\n\nis the inverse of the noise variances in the projected subspace. The mean\nis obtained by subtracting the mean of the latent\n\u000b\u0010\u000f and applying the projection matrix:\n\n, the projection matrix is \u001f! \n5\u001a\u001f('\n\u000b\u0010\u000f , c, and\n\u000b\u0010\u000f/.\n\n\u0014E\u001f)'+*\nfrom the transformation-normalized\n\u0007@\u001e\f<\n\u000b\u0010\u000f\n\n. For each factor,\n\n, it reduces to\n\n\u000b\u0010\u000f\n\n\u000b\u0010\u000f\n\n\u000b\u0010\u000f\n\n\u001f\u0005\u0017\n\n\u0017\"B\nAs the summation over\nat the same time in the frequency domain in\n\nfor all\n\n\u0019\u0001\u0017\n\n\u00077\u001f\n\nlog\n\ntime.\n\n\u001f+\u001f;B\n\n\u0019\u0001\u0017\n\nis a correlation, it can be ef\ufb01ciently computed for all\n\n\u0017\"B-\u001f\n\u0007@\u001e\f<\n\nThe inference on the latent image\n\nis given by its expected value:\n\n\u000b\u0010\u000f\n\n<>\u001e+/\nCOV\n\nB-2.!\n\n\u0006-!\n\n,.!=\u001f\n\n2-!\nis independent of\n\n!)(\n\nwhere 2\n\nthe model can be easily computed.\nwith the probability map\n\u000b\u0010\u000f\n\n\u0007@\u001e=<>\u001e+/\n\n\u000b\u0010\u000f\n\n<>\u001e\f/\n\n\u001e+/\n3\u0007\u0006\t\b\nde\ufb01ned for all\n\n\u000b\u0010\u000f\n\n\u001e+/\n\n\u000b\u0010\u000f\n\n3%\u0006\t\b\n. The \ufb01rst term, dictated by\n\u000b\u0010\u000f and\n\u000b\u0010\u000f\n\u000b\u0010\u000f\n\u000b\u0010\u000f\n, as a particular element in the sum\n\nis a convolution of\n\n\u000e\n\u0004\n\n\u0004\n\n\u0006\n\b\n\n\n\f\n\u0004\n\u001f\n\u001c\n\u0004\n\u0004\n\u001f\n\u0004\n\u001f\n\u001c\n\u0019\n\u000f\n\u0015\n\u000f\n\u0014\n\u0007\n\u0007\n\u0001\nB\n\u0014\n5\nB\n\u0012\n\u0014\n \nB\n\u001c\n,\n!\n(\n\n \nB\n\u001c\n\u0006\n\u0001\n!\n%\n \n\u0006\n!\n(\n/\n\n\u0007\n\u0005\n!\n/\n\n,\n\u0016\n\u001b\n\u0015\n/\n\n\u001e\n \n\u001c\n\u0007\n\u0001\n/\n\n\u0005\n!\n\u001f\n\u001b\n\u0019\n,\n\u0016\n\u001b\n\u0019\n\u0015\n/\n\n\u001e\n0\n\u0003\n1\n\u0004\n\n\u001c\n\u001f\n \n\u001f\n1\n\u001c\n/\n\n\u001c\n\u0019\n.\n.\n\u0005\n!\n\u001c\n\u0019\n\u0016\n\u001c\n\u001f\n \n\u001f\n\u0015\n\u0005\n\u001c\n/\n\n.\n\u0005\n!\n\u0019\n\u0007\n\u0007\n\u0004\n\u0004\n'\n,\n\u0016\n'\n\u0015\n\n\u0017\n\u001c\n\u0006\n\u0001\n!\n(\n'\n\n\u0005\n5\n'\n\n\u0003\nF\n\u001c\n\u0007\n\u0015\n<\n\n\u001f\n\u0007\n\u0001\n/\n\n!\nB\n\u0016\n'\n\u0015\n\n\u0017\n/\n\n\u0007\n3\nF\n\u001c\n\u0007\n\u0015\n<\n\n\u001f\n\u0007\n\u0001\n/\n\n/\n\nF\n\u001c\n\u0007\n\u0015\n\n\u001f\n\u0007\n\fFigure 4: Comparison of FA applied on data normalized for translations using correlation tracker\nand TCA. (a)Frames from sequence. (b) shift normalized frames, using correlation-based tracker and\nfor the TCA model.\n\nobtained through factor analysis model. (c)\u0002\u0001\n\nand \u0002\u0001\n\n\u0001\u0006\u0005\n\n\u000e\u0007\u0003\n\n\u0006\u0007\b\n\n\u0004\u0003\n\n\u0002\u0001\n\n\u0004\u001a\u0006\u0007\b\u000b\u0004\u0003\n\nFigure 5: Simulated walk sequence synthesized after training an AR model on the subspace and\nimage motion parameters. The sequence enlarged for better viewing of translations. The \ufb01rst row\ncontains a few frames from the sequence simulated directly from the model. The second row contains\na few frames from the video texture generated by picking the frames in the original sequence for\nwhich the recent subspace trajectory was similar to the one generated\n\nby the AR model.\n\nis 3\n\n3\u0007\u0006\t\b\n\n\u000b\u0010\u000f\n\n<>\u001e+/\n<>\u001e\f/\n\n\u000b\u0010\u000f\n\n\u000b\u0010\u000f\n2-!\n\n\u00077\u001f\n!\u0003(\nx\n\n. We can ef\ufb01ciently compute this sum for all\n\n:\n\n\t\b\n\n\u000b\u0010\u000f\n\n\u000b\u0010\u000f\u000b\n\n\t\t/\n\n2-!\n\n\u0017\"B\n\n\u001e+/\n\n!D(\n\n,.!H\u001f\nmatrices above can be done ef\ufb01ciently by factorizing\n\n\u0006.!\nNote that multiplication with\nthem and applying a sequence of vector multiplication from right to left.\n3 Experimental Results\nClustering face poses. In Fig. 3b the \ufb01rst column shows examples from a video sequence\nof one person with different facial poses walking across cluttered background. We trained\na transformation-invariant mixture of Gaussians (TMG) [6] with 3 clusters that learned\nmeans shown in Fig. 3b. TMG captures the three generic poses in the data.However, due\nto presence of lighting variations and small out-of-plane rotations in addition to big pose\nchanges, it is dif\ufb01cult for TMG to learn these variations without many more classes.\n\nWe trained a MTCA model with 3 classes and 2 factors, initializing the parameters to those\nlearned by TMG. Fig. 3a compares TMG means and components to those learned using\nMTCA. The MTCA model learns to capture small variations around the cluster means.\nFor example, for the \ufb01rst cluster, the two subspace coordinates tend to model out-of-plane\n,\nrotations and illumination changes (Fig. 3c).\n/\u0018\u0017\n(\n), of TMG and MTCA for various training examples, illustrating\nbetter tracking and appearance modelling of MTCA.\n\nIn Fig. 3b, we compare ,\n\nB\r\f\u000f\u000e\u0011\u0010\u0013\u0012\u0014\f\u0016\u0015\n\n!=F\n\n/A\u001f\n\n\u0004\n\nF\n\u001c\n\u0007\n\u0015\n\n\u001f\n/\n\n\u001c\n\u0019\n.\n\u0019\n,\n\u0016\n'\n\u0015\n\n\u001c\n\u0006\n\u0001\n'\n\n\u0005\n5\n'\nF\n\u001c\n\u0007\n\u0015\n<\n\n\u001f\n\n\u0004\n\u0004\n\u0016\n\u0005\n!\n(\n\u0006\n!\n\u001b\n\u0015\n<\n\u001c\n<\n\u0015\n\fFigure 6: Clustering faces extracted from a personal database prepared using face detector.\nTraining examples (b) Means, variances and components for two classes learned using MTCA. (c)\ncolumn contains several photos in which the detector [8] failed to \ufb01nd the face.\n\n\u0007 column contains central 100x100 portion of\n\n(a)\n\u0002\u0001\n\u0007 column contains\n.\n\n\u0004\u0006\u0005\n\n\u0006-\b\u000b\n\ncentral 100x100 portion of\n\n.\n\n\b\n\t\n\n\u0001\u0006\u0005\n\nModeling a walking person. Fig. 4a. shows three 165x285 frames from a video sequence\nof a person walking. For effective summarization, we need to learn a compact representa-\ntion for the dynamically and periodically changing hand and leg movements.\n\nA regular PCA or FA will learn a representation that focuses more on learning linearized\nshifts, and less on the more interesting motion of hands and legs (Fig. 2a.). The traditional\napproach is to track the object using, for example,a correlation tracker and then learn the\nsubspace model on normalized images. The parameters learned in this fashion are shown in\nFig. 2b. Without previously learning a good model, the tracker fails to provide the perfect\ntracking necessary for precise subspace modelling of limb motion and thus the inferred\nsubspace projection is blurred. (Fig. 2b).\n\n\u001c\f\u000b\n\n/\u0018\u0017\n\n.\n\n/\u0018\u0017\n\n\u00067\u001b\n\nand image position\n\nand infers a much cleaner projection ,\n\nAs TCA performs tracking and learns appearance model at the same time, not only does it\navoids the tracker initialization that plagues the \u201dtracking \ufb01rst\u201d approaches, but also pro-\nvides perfectly aligned ,\nThe TCA model can be used to create video textures based on frame reshuf\ufb02ing similar\nto [10]. However, instead of shuf\ufb02ing frames based directly on pixel similarity, we use\nthe subspace position\n\u000f generated from an AR process [9],\nfor which the window\nand for each t \ufb01nd the best frame u in the original video\nis the most similar to\n\u000f . Then, gen-\n\u001e\u0002\u000f\u000e\u000f\u0010\u000f\n\u001e\u0012\u000f\u0010\u000f\u0010\u000f\n\u001e\f\u001b\nis applied on the normalized image ,\n. The result is shown\nerated transformation\n\r\u000e\r\nin \ufb01g. 5b and contains a bit sharper images than the ones simulated directly from the gen-\nerative model, \ufb01g. 5a. We let the simulated walk last longer than in the original sequence\nletting MTCA live on twice as wide frames.\nClustering and face recognition We used a standard face detector [8] to obtain 85 32x32\nimages of faces of 2 persons, from a personal photo database of a mother and her daughter.\nIn \ufb01g. 6a. we present examples from the training set.\n\n\u000e\n\n\u000e\n\n\u000e\n\n\u000e\n\n\u0003\n\u0001\n\u000e\n\u0003\n\u0004\n\n\n\u0001\n\n\u0005\n\u000e\n\u0003\n\u0016\n'\n\u0015\n\u0016\n\u0005\n(\n\u0015\n\u001b\n\u001f\n\u0007\n\n\u001c\n/\n\u000f\n,\n\u0016\n\u001b\n\u0015\n/\n\u000f\n\u0017\n\u001e\n,\n\u0016\n\u001b\n\u0015\n/\n'\n\n\u000f\n\u0017\n,\n\u0016\n\u001b\n\u0015\n/\n'\n\u0011\n\u000f\n\u0017\n\u001b\n\n\u001c\n\u000f\n\n\u001c\n'\n\u0011\n\u0007\n\u0016\n'\n\u0015\n/\n\u000f\n\u0017\n\fWe learned a MTCA model with 2 classes and 4 factors. To model global lighting variation,\nwe preset one of the factors to be uniform at .01 (see \ufb01g. 6b.). This handles linearized\nversion of ambient lighting condition. We also preset another factor to be smoothly varying\nin brightness (see \ufb01g. 6b.) to capture side illumination. The other two components are\nlearned and they model slight appearance deformation such as facial expressions. The\n\nmodel learned to cluster faces in the training example with\u0002\u0001\n\nAn interesting application is to use the learned representation of the faces to detect and\nrecognize faces in the original photos. For instance, the face detector did not recognize\nfaces in many photographs (for eg.,\ufb01g. 6c), which we were able to using the learned model\n(\ufb01g. 6c). We increased the resolution of model parameters\nto match the res-\nolution of photos (640x480), padding around the original parameters with uniform mean,\nzero factors and high variance. Then, we performed inference, inferring most likely class,c,\nmost likely,\n. We also incorporated 3 rotations and 4 scales\nas possible transformations, in addition to all possible shifts. In \ufb01g. 6c , we present three\nexamples which were not in the training set and the face detector we used failed. In all\nthree cases MTCA detected and recognized the face correctly as belonging to class 2.\n\nfor that class and ,\n\n\"\u0004\u0003\u0006\u0005\n\naccuracy.\n\n\u0006.!E\u001e\n\n/I\u001e=<\n\n,7!\n\n4 Conclusion\n\nMixture of transformation-invariant component analyzers is a technique for modeling vi-\nsual data that \ufb01nds major clusters and major transformations in the data and learns subspace\nmodels of appearance. In this paper, we have described how a fast implementation of learn-\ning the model is possible through ef\ufb01cient use of identities from matrix algebra and the use\nof fast Fourier transforms.\nAppendix: EM updates Before performing inference in Estep, we pre-compute,for all classes the\nquantities that are independent of training examples:\n\nFor each training example\n\nComputing posteriors over\n\n\u0007\t\b\n\n\u0002\u000b\n\r\f\n\u0002\"\u000e\u0019\n\n\u000e*)\n+',\n\n\"\u0006\u000f\u000e\u0011\u0010\u0013\u0012\n\u0007\t\b$#%\u0007\u0011\b\n\r@012\n+',\n\u0007\u0011\b\n\nGFH0K\rI\u0010\n+',\n2S0K\u000e\n\n+\u0013,\n\n\u0002\u0001\n\nOQP\n\nEQR\n\nThe summation over\n\n\u0002\u0001\n\n\u0001\u0006\u0005\n\n+',\n\n0(2\n\ndistance between input and the transformed latent image. The determinant is simpli\ufb01ed to\n\ndistribution, we require the determinant of covariance matrix, COV\u0001\nCOV\u0001\n\nThe Mahalanobis distance between\n\nis evaluated and saved.\n\n(eqn. 1). To compute this\nand the Mahalanobis\n\n\u001d\u001e\b\n\b\u0016\u0015\n\u000e*)\n\u0006\u0018\u000e\n\n\u0004E\u0004\n\nOQP\nOQP\n\nEQR\nEQR\n\n\b\u000b\n\n\u0002\u0001\n\n\u0002\t\b\n\n\u0002\t\u0004\n\nI\u0010\n\n\b\u001b\u0015\n\n\u0007\t\b\nI\u0010\u0013\u0012\n\u0006\u0018\u000e\u0019\u0012\n\u0002\u000b\n\n\b\u0016\u0015\n\u0002\u000b\n\r\f\t\u0012\n\u0010\u0013\u0012\n\u0002\u000b\n\n\u0007\t\b\n\u0002\u000b\n\n\f\t\u0012\n\f\t\u0012\n\n'\u0010(\u0015\n+\u0013, , \u0002\u0001\n+\u0013,\n\u000e/)\n91:;\n\n\u00101-87\n\u000e/)\n.-\u0016\u0005\n\u0002546\n\n+',10+\r\t032\n\u00106<\n\r@012A\u0010\nand c, requires evaluating=>\n\n\u000e?)\n+\u0013,\n+\u0013,\n\r\t032\n\u0006B\u000e\u001e\u0010\n\n\f\n\u0006DC\n+', and the latent image is\n\u000e?)\n\u0007\u0011\b\n\u0007\t\b\n\r@0\n\r@0\n+',\n+\u0013,\n+',\n+\u0013,\n\u0010\u00139\n\u0006\u000f\u000e\n\nGL?M\n98:\n\u00101-87\n\n1\n\n9N:J\u000e\n\u00101-N7\nOQP\nEQR\nEQR\n\u0010UTWV\n\u0010\u00139\n-[T\\\n\n\u0010\u00139\n-ZY\n+',\n+',\n\u0010U`a\n\n280\n\u0010]T\\\nG^_\n\n+\u0013,\n-]g\u0004\nG^_\n\n\u0010UTcb\n\u0004ed\n\u0010\u00139\n\nGf\n\nGFH0\n\r\u001b\u00101\u0010]`J\u000e\n-'Y\n+\u0013,\n\r\u001b\u00106<on\n\nGm60\n\nGFH0K\r\u001b\u0010\n^_\n\n\u0010\u00139\n\nGf\n-/4_V\n\u0010\u00139\nM?jlk\ntime, afterE\n\nGFo0K\r\u001b\u0010\u0006q\u0016F and\nis computed inr\nlogs\n+\u0013,\n+',\n+\u0013,\n\u0010U`J\u000e\n280\n^_\n\n\u0010U`\u0007\u000e\n2S0\n^_\n\nKJ\nOQP\n\n-'Y\n\n+\u0013,\n\nGf\n\ntakes only\n\n\u0007\u0011\b\n\n+',\n\n+',6h\n\ntime.\n\n\u000f\n\u0005\n!\n\u001e\n\u0007\n\u0016\n'\n\u0015\n\u0017\n\n\u0014\n\b\n\n\b\n\n\u0006\n\n\u0017\n\b\n\n\n\n\n\u001a\n\b\n\u0002\n\u0014\n\b\n\n\u001c\n\b\n\u0002\n\u0017\n\b\n\f\n\u0012\n\n\n\b\n\n\u0014\n\b\n\b\n\u0015\n\n\f\n\u0012\n\n\n\b\n\n\u0006\n\u0012\n\n\u001f\n\b\n\u0015\n\n\f\n\u0012\n\n\n\b\n\n \n\b\n\u0015\n\n\f\n\u0012\n\n\n\u0017\n\b\n!\n\b\n\b\n\n\u0014\n\b\n\n\u0010\n\u0004\n\n&\n\b\n\u0017\n\b\n\n\n\u0006\n\u001c\n\b\n \n\b\n\n\n\b\n\u0003\n\u001a\n\b\n#\n\u0004\n\n\u0005\n\u000e\n)\n\u0005\n\u0003\n\u0005\n\u0005\n\u0003\n\u0005\n\u0002\n\u0005\n\f\n\n\u0005\n\u0005\n\b\n\u0015\n\n\u0012\n\n\b\n\n\u0005\n\n\u0015\n\u000e\n)\n\u0010\n\u0015\n\n\u0015\n\u000e\n)\n#\n\u0015\n\n\u0015\n\u000e\n)\n\u0006\n\u0004\n\u0015\n\n\u0004\n\n#\n\n\u0005\n\u000e\n)\n\u0003\n\u0015\n\u0014\n\b\n\u0012\n\n\n\u0001\n\n\u0005\n\u000e\n)\n\u0003\nE\n\u0002\n\b\n\u000e\n)\n7\n\b\n7\n-\n\u0002\n!\n\b\n\u0015\n\f\n\u0012\n\n\n\b\n\n\u0012\n\n&\n\b\n)\n\u0010\n\n\u0001\n\u0001\n\u0015\n\u0005\n)\n\u0010\n\u0003\n\u0002\n\n\u0017\n\b\n\u0010\n#\n\n\u0017\n\b\n\u0017\n\b\n\f\n\u0012\n\n\n\f\n\u0012\n\n\nX\n\n\n\b\n\n7\n\b\n\n\u0014\n\b\n7\n-\n\u0006\n\n\u0017\n\b\n\u0002\n\u000e\n\u0012\n\u0002\n\n\u0005\n\u000e\n)\n\u000e\n)\n\u0010\n\u0002\n\u0010\n\u0006\n\n\u0017\n\b\n\u0002\n\f\n\u0012\n\u0002\n\n\u000e\n\u0012\n\u0002\n#\nV\nX\n\n\b\n7\n\n\u0005\n\u000e\n)\n\u0010\nE\n)\nV\nX\n\n\b\n7\nX\ni\nY\n\n\b\n7\ni\nX\n\n\u0005\n\u000e\n)\n\u0010\nE\nE\n\n\u0005\np\n\u0005\n\n\u0005\np\n\u0005\n\u000e\n)\n\u0003\n\u0002\n!\n\b\n\u0006\n\u0017\n\b\n\u000e\n\u0012\n\n4\n\n\u0005\n\u000e\n)\n)\n<\n\u0006\n\u001c\n\b\ng\n \n\b\n\u000e\n\u0012\n\n4\n\n\u0005\n\u000e\n)\n)\n<\nh\n\f+',\n\n0(2\n\n\u0002\u0001\n\n\u0010\u0001\nEQR\n\n\u0002\u0001\n\nOQP\n\n2S0K\u000e\n\n+',\n\n\u0003K\f\n\nDe\ufb01ning, \u0004\u0006\u0005\b\u0007\n\n280\n\n\u000f\u000e\u0011\u0010\u0006\u0012\n\r\u0019\u000e\u001c\u0010\n\n\u0002\u0001\n\u0002\n\t\f\u000b\n\n\u001e\u001d\n\n\u0010'\n\n\u0010\u0001\n\n-'Y\n\n\u0004E\u0004\n\n+',\nOQP\n\n\u0015\u0017\u0016\n\n\u0014\u0013\n\n\u0014\u0013\n\u0002\u0001\n\u0001\u0006\u0005\n\u0002\u0001\n\u0002\u0001\n\n\u00034\u0002\n\r\u0019\u0018\n\r\u0019\u0018\n\n,\u001b\u001a\n2S0K\u000e\nEQR\n\n\u0002\u0001\n\n+\u0013,\n\n2S0\n\n+',\n\n\u0002\u0001\n\n\u0003>\u0006\n\n\u0003\u0002T\n+',\n\n\u0002\u0001\n\n2S0K\u000e\n-[`\u0007\u000e\n\n\u0010\u00139\n\n+\u0013,\nTI\u0004\n\n0+\r@0(2\n+',\n\n+\u0013,\n\n\u0002\u0001\n\nM?jlk\n\u0010\u00139\n-ZY\n280K\u000e\n\n\u0002\u0001\n\n\b\u000b\n\n\u0001\u0006\u0005\n\nG^_\n\nT@g\u0004\n\n2S0K\u000e\n^_\n\n+',\n2S0K\u000e\n>\u0010'\n\n\u0002\u0001\n\n280\n+\u0013,\n+',\n\n012\n280\n\n\u0003\u0019\u0006\n\n\u0002\u0001\n\n-ZY\n\n\b\u000b\n\n+\u0013,\n280\n\n\u0002\u0001\n\n\u0002\u000b\u0005\n\n+',\n\n\u0010UT\n\n\b\u0012\n\n\u0002\u0001\n\nOQP\n\u0010'\nGL?M\n^_\n\n\u0002\u0001\n\n\u0001\u0006\u0005\n\nEQR\n+',\n\n0(2\n2S0\n\u0010\u00139\n2S0\n2S0K\u000e\n>\u0010\n\n?\n\n\u0007\u001f\u0004\n\n\u0002\u0001\n\n+',\n\n+',\n\n0K\r@012\n280K\u000e\n+\u0013,\n\u0010.T\n\u0010'\nGL?M\n\n+',\n\u0010\u00139\n+\u0013,\n+\u0013,\n\n+',\n280K\u000e\n\n+\u0013,\n\n, in the Mstep, parameters are updated according to\n\nReferences\n[1] Everitt, B.S. An Introduction to Latent variable models Chapman and Hall, New York NY 1984\n[2] Frey, B.J. , Colmenarez, A. & Huang, T.S. Mixtures of local linear subspaces for face recognition.\nIn Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1998 IEEE\nComputer Society Press: Los Alamitos, CA.\n\n[3] Frey, B.J. & Jojic, N. Fast, large-scale transformation-invariant clustering. In Advances in Neural\n\nInformation Processing Systems 14. Cambridge, MA: MIT Press 2002\n\n[4] Ghahramani, Z. & Hinton, G. The EM Algorithm for Mixtures of Factor Analyzers University\n\nof Toronto Technical Report CRG-TR-96-1, 1996\n\n[5] Hinton, G., Dayan, P. & Revow, M. Modeling the manifolds of images of handwritten digits In\n\nIEEE Transactions on Neural Networks 1997\n\n[6] Jojic, N. & Frey, B.J. Topographic transformation as a discrete latent variable In Advances in\n\nNeural Information Processing Systems 13. Cambridge, MA: MIT Press 1999\n\n[7] Jolliffe,I.T. Principal Component Analysis Springer-Verlag, New York NY, 1986.\n[8] Li, S.Z., Zhu.L , Zhang, Z.Q. & Zhang,H.J. Learning to Detect Multi-View Faces in Real-Time\nIn Proceedings of the 2nd International Conference on Development and Learning, June, 2002.\n[9] Neumaier, A. & Schneider,T. Estimation of parameters and eigenmodes of multivariate autore-\n\ngressive models In ACM Transactions on Math Software 2001.\n\n[10] Schdl,A. Szeliski,R.,Salesin,D.& Irfan Essa Video textures In Proceedings of SIGGRAPH2000\n[11] Simard, P. , LeCun, Y. & Denker, J. Efficient pattern recognition using a new transformation\n\ndistance In Advances in Neural Information Processing Systems 1993\n\n[12] Turk, M. & Pentland, A. Face recognition using eigenfaces In Proceedings of IEEE Conference\n\non Computer Vision and Pattern Recognition Maui, Hawaii, 1991\n\n[13] Wolberg,G. & Zokai,S. Robust image registration using log-polar transform In Proceedings\n\nIEEE Intl. Conference on Image Processing, Canada 2000.\n\n\n\n\u0015\n7\n+\n\u0005\n\u000e\n)\n\u0003\n\u0002\n\n\u0014\n\b\n7\n+\n\u0006\nX\n\n\u0005\n\u000e\n)\n\u000e\n)\n\n+\n\u0005\n\u000e\n)\n\u0003\n\u0010\n\n\u0001\n#\n\u0004\n\n#\n\b\n\n\n\u0001\n#\n\u0004\n\n#\n\b\n\n\n\u0010\n\u0015\n\u0005\n\u000e\n)\n\u0003\n\u0002\n\n\u0001\n\n\u0001\n\u0001\n\u0015\n\u0005\n)\n\u0010\n\u0004\n\n\u0006\nV\nX\n\n\n\b\n\n7\n-\n\b\n\n\n\n\u0015\n\u0005\n\u000e\n)\n\u0003\n7\n-\n#\n\u0004\n\u0001\n\n\u0015\n\u0005\n)\n\u0003\nh\n#\n\nT\n)\n\u0004\n\n\n\u0005\n\u000e\n)\n\u0003\n\u0004\n\n\n\u0001\n\u0001\n\u0001\n\u0015\n\u0005\n)\n\u0012\n\n\n\b\n\n\u0002\n&\n\b\n\u0006\n!\n\b\nV\nX\n\n\n\u0005\n\u000e\n)\n7\n\b\n7\n-\n\u0006\n\n\u0017\n\b\n\u0006\n\u001c\n\b\n \n\b\n\u0010\n\u0003\n\u0012\n\ng\nV\nX\n\n\n\u0005\n\u000e\n)\n7\n\b\n7\n)\nh\n\u0001\n\n\u0015\n\u0005\n\u000e\n)\n4\n\u0001\n\u0001\n\u0015\n\u0005\n)\n\u0003\n\f\n\u0012\n\n\n#\n)\n\u0003\n\u0004\n\u0015\n\n\f\n\u0012\n\n\n\b\n\n<\n\u001d\n\b\n)\n\t\n\u000b\n\u0012\n)\n\u0015\n\u0016\n,\n\u0004\n\u0004\n)\n\u0003\n#\n\n\u0005\n)\n\u0003\n\u0007\n\f\n\n\u001d\n\u0004\n\n\u0001\n#\n\u0004\n\n#\n\b\n\n\u0001\n#\n\u0004\n\n#\n\b\n\n\u0015\n\u0005\n\u000e\n)\n0\n2\n\u0003\n\u0007\n\b\n\n\u001d\n\u0004\n\u0001\n\n\u0015\n\u0005\n\u000e\n)\n\u0003\n#\n\u0004\n\n\n\u0015\n\u0005\n\u000e\n)\n\u0003\n\u0015\n\u0005\n)\n\u0003\n\u0007\n\u0012\n\n\f", "award": [], "sourceid": 2212, "authors": [{"given_name": "Anitha", "family_name": "Kannan", "institution": null}, {"given_name": "Nebojsa", "family_name": "Jojic", "institution": null}, {"given_name": "Brendan", "family_name": "Frey", "institution": null}]}