{"title": "The Infinite Hidden Markov Model", "book": "Advances in Neural Information Processing Systems", "page_first": 577, "page_last": 584, "abstract": null, "full_text": "The In\ufb01nite Hidden Markov Model\n\nMatthew J. Beal\n\nZoubin Ghahramani\n\nCarl Edward Rasmussen\n\nGatsby Computational Neuroscience Unit\n\nUniversity College London\n\n17 Queen Square, London WC1N 3AR, England\n\nhttp://www.gatsby.ucl.ac.uk\n\n m.beal,zoubin,edward\u0001 @gatsby.ucl.ac.uk\n\nAbstract\n\nWe show that it is possible to extend hidden Markov models to have\na countably in\ufb01nite number of hidden states. By using the theory of\nDirichlet processes we can implicitly integrate out the in\ufb01nitely many\ntransition parameters, leaving only three hyperparameters which can be\nlearned from data. These three hyperparameters de\ufb01ne a hierarchical\nDirichlet process capable of capturing a rich set of transition dynamics.\nThe three hyperparameters control the time scale of the dynamics, the\nsparsity of the underlying state-transition matrix, and the expected num-\nber of distinct hidden states in a \ufb01nite sequence. In this framework it\nis also natural to allow the alphabet of emitted symbols to be in\ufb01nite\u2014\nconsider, for example, symbols being possible words appearing in En-\nglish text.\n\n\u0005\u0010\u000f\n\n\u000f*12\u0007\n\n1 Introduction\nHidden Markov models (HMMs) are one of the most popular methods in machine\nlearning and statistics for modelling sequences such as speech and proteins. An\n\n\u000f , \u0018\u0006\u001c\n\n\t\u0006\u000b\f\u000b\f\u000b\u001a\t\u001b\u0018\n\n\t\f\u000b\f\u000b\u0006\u000b\u0012\t\u001b\u0018\n\n\u00033)54\n\nHMM de\ufb01nes a probability distribution over sequences of observations (symbols) \u0002\u0004\u0003\n\u0006\u0005\b\u0007\n\t\f\u000b\u0006\u000b\f\u000b\r\t\u000e\u0005\u0010\u000f\u0011\t\f\u000b\f\u000b\u0006\u000b\u0012\t\u0013\u0005\u0015\u0014\n\u0001 by invoking another sequence of unobserved, or hidden, discrete\n\u0019\u0018\n\u0001 . The basic idea in an HMM is that the se-\nstate variables \u0016\u0017\u0003\nquence of hidden states has Markov dynamics\u2014i.e. given \u0018\nis independent of \u0018\u001e\u001d\nfor all \u001f! #\"$ &% \u2014and that the observations \u0005'\u000f are independent of all other variables\ngiven \u0018\n\u000f . The model is de\ufb01ned in terms of two sets of parameters, the transition matrix\nwhose (*)\u0015+-, element is .0/\n\u00036(87 and the emission matrix whose (:9\u0019+-, element\n\u0018\u0006\u000f\nis .0/\n\u0003\u0004(=7 . The usual procedure for estimating the parameters of an HMM is\n\u0003;9<4\nmatrices > and ?\n\nthe Baum-Welch algorithm, a special case of EM, which estimates expected values of two\ncorresponding to counts of transitions and emissions respectively, where\n\nthe expectation is taken over the posterior probability of hidden state sequences [6].\nBoth the standard estimation procedure and the model de\ufb01nition for HMMs suffer from\nimportant limitations. First, maximum likelihood estimation procedures do not consider\nthe complexity of the model, making it hard to avoid over or under\ufb01tting. Second, the\nmodel structure has to be speci\ufb01ed in advance. Motivated in part by these problems there\nhave been attempts to approximate a full Bayesian analysis of HMMs which integrates over,\nrather than optimises, the parameters. It has been proposed to approximate such Bayesian\nintegration both using variational methods [3] and by conditioning on a single most likely\nhidden state sequence [8].\n\n\u0007\n\u000f\n\u0014\n\u0018\n\u0018\n\u000f\n\fIn this paper we start from the point of view that the basic modelling assumption of\nHMMs\u2014that the data was generated by some discrete state variable which can take on\none of several values\u2014is unreasonable for most real-world problems. Instead we formu-\nlate the idea of HMMs with a countably in\ufb01nite number of hidden states. In principle,\nsuch models have in\ufb01nitely many parameters in the state transition matrix. Obviously it\nwould not be sensible to optimise these parameters; instead we use the theory of Dirichlet\nprocesses (DPs) [2, 1] to implicitly integrate them out, leaving just three hyperparameters\nde\ufb01ning the prior over transition dynamics.\nThe idea of using DPs to de\ufb01ne mixture models with in\ufb01nite number of components has\nbeen previously explored in [5] and [7]. This simple form of the DP turns out to be inade-\nquate for HMMs.1 Because of this we have extended the notion of a DP to a two-stage hi-\nerarchical process which couples transitions between different states. It should be stressed\nthat Dirichlet distributions have been used extensively both as priors for mixing propor-\ntions and to smooth n-gram models over \ufb01nite alphabets [4], which differs considerably\nfrom the model presented here. To our knowledge no one has studied inference in discrete\nin\ufb01nite-state HMMs.\nWe begin with a review of Dirichlet processes in section 2 which we will use as the basis\nfor the notion of a hierarchical Dirichlet process (HDP) described in section 3. We explore\nproperties of the HDP prior, showing that it can generate interesting hidden state sequences\nand that it can also be used as an emission model for an in\ufb01nite alphabet of symbols. This\nin\ufb01nite emission model is controlled by two additional hyperparameters. In section 4 we\ndescribe the procedures for inference (Gibbs sampling the hidden states), learning (op-\ntimising the hyperparameters), and likelihood evaluation (in\ufb01nite-state particle \ufb01ltering).\nWe present experimental results in section 5 and conclude in section 6.\n\n. The transition probabilities\nrow of the transition matrix can be interpreted as mixing proportions for\n\n2 Properties of the Dirichlet Process\nLet us examine in detail the statistics of hidden state transitions from a particular state \u0018\u0010\u000f\n\u000f*12\u0007 , with the number of hidden states \ufb01nite and equal to \u0001\nto \u0018\ngiven in the (\n+-,\n\u0018\u0006\u000f*12\u0007\nthat we call\u0002\nImagine drawing >\non values \r\f\u0015\t\u0006\u000b\f\u000b\f\u000b\f\t\n.0/\n\n\u0004\u0003\u0012\u0007\u0019\t\f\u000b\f\u000b\u0006\u000b\f\t\u0005\u0003\u0007\u0006\n\u0001 .\nsamples \t\b\n\u0007\u001e\t\u0006\u000b\f\u000b\u0006\u000b\n\u0001 with proportions given by \u0002\n\t\u0006\u000b\f\u000b\f\u000b\u0006\t\u0005\b\n\nfrom a discrete indicator variable which can take\n. The joint distribution of these indicators\n\nis multinomial\n\n\t\u0005\b\u000b\n\n\u0012\u0011\n\nwith\n\n\u00148\t\n\n(1)\n\n\u000f\u0005\u0010\nthat \u0018\n\n)\b7\n\u001a , and \u001c otherwise)\nwhere we have used the Kronecker-delta function (\u0016\nto count the number of times >\n\u00030) has been drawn. Let us see what happens to\nthe distribution of these indicators when we integrate out the mixing proportions \u0002 under a\nconcentration hyperparameter\u001d\n\u001d27 \u001f\"!$#&%'#)(\u001b*\u0017+&,.-\n/\u001e\u001d0/\n\nconjugate prior. We give the mixing proportions a symmetric Dirichlet prior with positive\n\n\u0007\u0017\u0016\niff \u0018\n\n\t\u0006\u000b\f\u000b\u0006\u000b\n\n.0/\u001e\u0002\n\n\u000343\r5\n\n\u000676\u0012\u0007\n\n\u000f*12\u0007\n\n\u0015\u0014\n\n(2)\n\n\t\u001b\u001a\n\n7\u0012\u0003\n\n/\u0019\u0018\n\nis restricted to be on the simplex of mixing proportions that sum to 1. We can\n\n7$\u0003\nwhere \u0002\nanalytically integrate out \u0002 under this prior to yield:\n\n\u001d0/\n\n/2\u001d\n/\u001e\u001d0/\n\n\u000f\u0005\u0010\n\n1That is, if we only applied the mechanism described in section 2, then state trajectories under the\nprior would never visit the same state twice; since each new state will have no previous transitions\nfrom it, the DP would choose randomly between all in\ufb01nitely many states, therefore transitioning to\nanother new state with probability 1.\n\n\u0003\n(\n\u0003\n\u0001\n\u0001\n\b\n\u0007\n\n4\n\u0002\n7\n\u0003\n\u0006\n\u000e\n\u0007\n\u0003\n\u000f\n\t\n>\n\u000f\n\u0003\n\n\u0013\n\u0010\n/\n\b\n\n\f\n\u0003\n\u000f\n4\n\u0001\n\t\n\u0001\n1\n7\n1\n\u0001\n7\n\u0006\n\u0006\n\u000e\n\u0007\n\u000f\n\t\n\f/*>\n\n\u000f\u0005\u0010\n\n)54\n\n\u001d27\n\n\u001d0/\n\n.0/\n\n.0/\n\n.0/\n\n/\u001e\u001d0/\n\n78.0/2\u0002\n\n/\u001e\u001d27\n/->\u0005\u0004\n\nThus the probability of a particular sequence of indicators is only a function of the counts\n\nindicator removed. Note the self-reinforcing\nis more likely to choose an already popular state. A key property of DPs,\nwhich is at the very heart of the model in this paper, is the expression for (4) when we take\n\n\u0003\u0001\u0003\u0002\n\t\f\u000b\f\u000b\u0006\u000b\f\t'\b\n\t\f\u000b\u0006\u000b\f\u000b\f\t'\b\n\u0001 . The conditional probability of an indicator \b\u0007\u0006 given the setting of all other\n6\t\u0006 ) is given by\n\u0007\u0019\t\f\u000b\u0006\u000b\f\u000b\nindicators (denoted\b\n6\n\u0006\n\u001d27\n6\n\u0006\f\u000b\nis the counts as in (1) with the\u0002'+-,\nproperty of (4): \b\f\u0006\nwhere>\n\n\u0014\u0013\u0016\u0015\u0018\u0017\nthe limit as the number of hidden states \u0001\n)\u001a\u0019\n6\n\u0006\n6\u0012\u0007\u00131\n\u0003\u0010\u000f\n6\u0012\u0007\u00131\nwhere\u001b\n\u001c ), which cannot\nis \ufb01nite. \u001d\n\u0001 , i.e. the strength of belief in the symmetric prior. 2 In the in\ufb01nite limit\n\n)54\nis the number of represented states (i.e. for which >\nbe in\ufb01nite since >\n\t.\f\n\t\u0006\u000b\f\u000b\u0006\u000b\n\nacts as an \u201cinnovation\u201d parameter, controlling the tendency for the model to populate a\n\nfor all unrepresented) , combined\n\n\u001d27\n6\t\u0006\f\u000b\n>\u000e\r\n\f'\t\f\u000b\u0006\u000b\f\u000b\n\ncan be interpreted as the number of pseudo-observations of\n\ntends to in\ufb01nity:\n\ni.e. represented\n\n6\t\u0006\f\u000b\n\n\u000f\u001e\u001d\n\n\t\u001c\u001b\n\n.0/\n\n\u001d0/\n\n\u001d27\n\n(3)\n\n(4)\n\n(5)\n\npreviously unrepresented state.\n\n3 Hierarchical Dirichlet Process (HDP)\nWe now consider modelling each row of the transition and emission matrices of an HMM as\na DP. Two key results from the previous section form the basis of the HDP model for in\ufb01nite\nHMMs. The \ufb01rst is that we can integrate out the in\ufb01nite number of transition parameters,\nand represent the process with a \ufb01nite number of indicator variables. The second is that\nunder a DP there is a natural tendency to use existing transitions in proportion to their\nprevious usage, which gives rise to typical trajectories. In sections 3.1 and 3.2 we describe\nin detail the HDP model for transitions and emissions for an in\ufb01nite-state HMM.\n\n3.1 Hidden state transition mechanism\n\n(87\n\n\u00148\t\n\n)54\n\n.0/\n\n\t\u001c\u001b\n\n\u0018\u0006\u000f*12\u0007\n\n\"$#\n\n, i.e. we prefer to\n\nreuse transitions we have used before and follow typical trajectories (see Figure 1):\n\nNote that the above probabilities do not sum to 1\u2014under the DP there is a \ufb01nite probability\n\nImagine we have generated a hidden state sequence up to and including time \" , building\nto ) , i.e. >!\u001f\na table of counts > \u001f\n6\u0012\u0007\n( , we impose on state \u0018\n\u000f*12\u0007 a DP\n(5) with parameter \u001d whose counts are those entries in the (\n+-,\n)\u0005\u0019\n\t\u001c\u001b\n\t\u001c\u001b\n\n\u001d27\n7 of not selecting one of these transitions. In this case, the model defaults\n\u001d0/</\n\u000f*12\u0007 with parameter& whose counts are given by a vector\nto a second different DP (5) on \u0018\n> '\n\u000f . We refer to the default DP and its associated counts as the oracle. Given that we have\n\n*)\n\u0014/.\u00140\n+-,\n\u000f*12\u0007\n\u0014/.\u00140\n+-,\nresented states available, each of which have in\ufb01nitesimal mass proportional to4\n\nfor transitions that have occured so far from state (\n)\b7 . Given that we are in state \u0018\nrow of >\n\f\u0015\t\u0006\u000b\f\u000b\u0006\u000b\n\u0018\u0006\u000f\n>%\u001f\n)\u001a\u0019\n)32\n\n2Under the in\ufb01nite model, at any time, there are an in\ufb01nite number of (indistinguishable) unrep-\n\ndefaulted to the oracle DP, the probabilities of transitioning now become\n\nrepresented \t\nis a new state \u000b\n\n\f'\t\f\u000b\f\u000b\u0006\u000b\n\f'\t\f\u000b\f\u000b\u0006\u000b\n\n\u0003(\u000f\n\ni.e.)\ni.e.)\n\n1\t1\n1\t1\n\n&\r7\n\n.0/\n\n)54\n\n(6)\n\n.\n\n(7)\n\n\b\n\u0007\n\n4\n\u001d\n7\n\u0002\n\b\n\u0007\n\n4\n\u0002\n4\n\u0003\n1\n1\n\u0006\n\u000e\n\u0007\n1\n\u000f\n\u0004\n\u0001\n7\n1\n\u0001\n7\n\u000b\n\n>\n\t\n>\n\u0006\n\b\n\u0006\n\u0003\n\b\n\t\n\u0003\n>\n\u000f\n\u0004\n\u0001\n\f\n\u0004\n\u001d\n\t\n\u000f\n\b\n\u0006\n\u0003\n\b\n\t\n\u0011\n\u0012\n\u0011\n\n3\n\n\u0001\n3\n\n3\n\u0002\n\u0003\n\n\f\n/\n\u0001\n/\n\u0001\n\u001d\n\u000f\n\u000f\n\u0003\n\"\n\u000f\n\u000f\n\u0014\n\u0010\n\u0007\n\u0016\n/\n\u0018\n\u000f\n\u0016\n/\n\u0018\n\u000f\n\u0014\n1\n\u0007\n\t\n\u000f\n\u0003\n\u0003\n\u0003\n(\n\t\n>\n\t\n\u0003\n>\n\u001f\n\u000f\n\u000f\n\u0014\n\u0010\n\u0007\n\u000f\n\u0014\n\u0004\n\u001d\n\n\u0001\n\u000b\n\"\n\u000f\n\u0014\n>\n\u001f\n\u000f\n\u0014\n\u0004\n\u001d\n\u0018\n\u0003\n\u0018\n\u000f\n\u0003\n(\n\t\n>\n'\n\t\n\u0011\n\u0012\n\u0011\n\u0011\n\n)\n\u0011\n\u0014\n\n\u0001\n1\n\u0011\n\n)\n\u0011\n\u0014\n\u0019\n\n\u0001\n\fa)\n\nc)\n\nnii + a\nnij + b + a\n\nS\nj\n\nself(cid:13)\n\ntransition\n\nnij\n\nb\n\nS\nj\n\nnij + b + a\nexisting(cid:13)\ntransition\n\nj=i\n\nS\nj\n\nnij + b + a\noracle\n\no\nnj\no + g\n\nS\nnj\nj\nexisting(cid:13)\n\nstate\n\ng\n\nS\nj\n\no + g\n\nnj\nnew(cid:13)\nstate\n\nb)\n\nd)\n\n(left) State transition generative mechanism. (right a-d) Sampled state trajectories\n(time along horizontal axis) from the HDP: we give examples of four modes of\n, explores many states with a sparse transition matrix. (b)\n,\n, has strict left-to-right transition\n\n, retraces multiple interacting trajectory segments. (c)\n\n\b\t\u0001\n\u0007\f\u000b\u000e\r\u0010\u000f\n\n\b\u0011\u0001$\r\u0010\u000f\n\n\u0002\u0001\u0004\u0003\u0006\u0005\u0006\u0007\n\nFigure 1:\nof length\nbehaviour. (a)\n\b\u0019\u0001\u0012\u0007\f\u000f\nswitches between a few different states. (d)\ndynamics with long linger time.\n\n4\u0011\u0001\u0012\r\u0013\u0007\u0010\u0007\u0006\u0007\f\u000f\u0015\u0014\u0016\u0001\u0017\r\u0018\u0007\u0010\u0007\n4\u001a\u0001\u0012\u0007\f\u000b\u000e\r\u0006\u000f\u001b\u0014\u001c\u0001\u001d\r\u0018\u0007\u0010\u0007\nUnder the oracle, with probability proportional to&\n>%\u001f\nus. After each transition we set > \u001f\n\u000f)(\noracle DP just described then in addition we set >!'\n' will increase.\n\nstate then the size of > and >\n\n\b\u001e\u0001 \u001f!\u000f\n\n4\u001a\u0001\u0002\u0003\"\u000f#\u0014\u001c\u0001\u0002\u0003\n4%\u0001&\r\u0006\u000f#\u0014'\u0001&\r\u0013\u0007\u0010\u0007\u0006\u0007\u0010\u0007\n\f and, if we transitioned to the state) via the\n> '\n\f . If we transitioned to a new\n\nan entirely new state is transitioned\nto. This is the only mechanism for visiting new states from the in\ufb01nitely many available to\n\nSelf-transitions are special because their probability de\ufb01nes a time scale over which the\ndynamics of the hidden state evolves. We assign a \ufb01nite prior mass\nto self transitions for\neach state; this is the third hyperparameter in our model. Therefore, when \ufb01rst visited (via\n\nin the HDP), its self-transition count is initialised to\n\n.\n\nin\ufb02uences the tendency to explore new transitions, correspond-\n\nThe full hidden state transition mechanism is a two-level DP hierarchy shown in decision\ntree form in Figure 1. Alongside are shown typical state trajectories under the prior with\ndifferent hyperparameters. We can see that, with just three hyperparameters, there are a\n\nwealth of types of possible trajectories. Note that& controls the expected number of repre-\nsented hidden states, and\u001d\n\ning to the size and density respectively of the resulting transition count matrix. Finally\ncontrols the prior tendency to linger in a state.\nThe role of the oracle is two-fold. First it serves to couple the transition DPs from different\nhidden states. Since a newly visited state has no previous transitions to existing states,\nwithout an oracle (which necessarily has knowledge of all represented states as it created\nthem) it would transition to itself or yet another new state with probability 1. By consulting\nthe oracle, new states can have \ufb01nite probability of transitioning to represented states. The\nsecond role of the oracle is to allow some states to be more in\ufb02uential (more commonly\ntransitioned to) than others.\n\n3.2 Emission mechanism\n\n\u000f,+\n\nin every\nrespect except that there is no concept analogous to a self-transition. Therefore we need\nfor the emission HDP. Like for state\n\nis identical to the transition process \u0018\n\nThe emission process \u0018\nonly introduce two further hyperparameters\u001d/.\n\u001f10\ntransitions we keep a table of counts ?\nof times before \"\n\n\u0018\u0006\u000f2\u00148\t\nthat state ( has emitted symbol 9 , and ?\n\nand&).\n6\u0012\u0007\n\nis the number of times symbol\n\n9\u00107 which is the number\n\n\u0005\u0010\u000f2\u00148\t\n\n\u000f*12\u0007\n\n\u000f-+\n\n(=7\n\n\u000f\n\u0004\n\u000f\n(\n\u000f\n\u0004\n*\n&\n*\n*\n\u0005\n\u000f\n\u0018\n\u0003\n\"\n\u000f\n\u000f\n\u0014\n\u0010\n\u0007\n\u0016\n/\n\u0016\n/\n'\n0\n\fmiq\nS\nmiq + be\nq\nexisting(cid:13)\nemission\n\nbe\n\nqS\n\nmiq + be\noracle\n\nmq\n\nge\n\nqS\n\no + ge\n\nmq\nexisting(cid:13)\nsymbol\n\no + ge\n\nS\nq\n\nmq\nnew(cid:13)\nsymbol\n\n2500\n\n2000\n\n1500\n\n1000\n\n500\n\n0\n0\n\n102\n\n101\n\n0.5\n\n1\n\n1.5\n\n2\n\n2.5\n\nx 104\n\n100\n0\n\n20\n\n40\n\n60\n\n80\n\n100\n\nFigure 2:\n(left) State emission generative mechanism. (middle) Word occurence for entire Alice\nnovel: each word is assigned a unique integer identity as it appears. Word identity (vertical) is plotted\nagainst the word position (horizontal) in the text. (right) (Exp 1) Evolution of number of represented\n(vertical), plotted against iterations of Gibbs sweeps (horizontal) during learning of the\nstates\nascending-descending sequence which requires exactly 10 states to model the data perfectly. Each\nline represents initialising the hidden state to a random sequence containing\n\u0003!\u000f\u0004\u0003\f\u000f\u0018\u000b\u0013\u000b\u0013\u000b\u0013\u000f\u0013\r\u0018\u0003\u0006\u001f\u0006\u0005\ndistinct represented states. (Hyperparameters are not optimised.)\n\n\u0001\u0002\u0001\"\r\u0006\u000f\n\n9 has been emitted using the emission oracle.\n\nFor some applications the training sequence is not expected to contain all possible obser-\nvation symbols. Consider the occurence of words in natural text e.g. as shown in Figure 2\n(middle) for the Alice novel. The upper envelope demonstrates that new words continue to\nappear in the novel. A property of the DP is that the expected number of distinct symbols\n(i.e. words here) increases as the logarithm of the sequence length. The combination of\nan HDP for both hidden states and emissions may well be able to capture the somewhat\nsuper-logarithmic word generation found in Alice.\n\n\u001e\u0018\u0010\u0007\u001e\t\f\u000b\f\u000b\u0006\u000b\n\n\t\u001b\u0018\n\u0014\n\n4 Inference, learning and likelihoods\nGiven a sequence of observations, there are two sets of unknowns in the in\ufb01nite HMM:\n\nde\ufb01ning the transition and emission HDPs. Note that by using HDPs for both states and\nobservations, we have implicitly integrated out the in\ufb01nitely many transition and emission\nparameters. Making an analogy with non-parametric models such as Gaussian Processes,\n\nthe hidden state sequence \u0016\u0012\u0003\nwe de\ufb01ne a learned model as a set of counts \n\n\u0001 , and the \ufb01ve hyperparameters \n>!'\n\n?3'\u0019\u0001 and optimised hyperparameters\n\nWe \ufb01rst describe an approximate Gibbs sampling procedure for inferring the posterior over\nthe hidden state sequence. We then describe hyperparameter optimisation. Lastly, for cal-\nculating the likelihood we introduce an in\ufb01nite-state particle \ufb01lter. The following algorithm\nsummarises the learning procedure:\n\n.\u0006\u0001 .\n\nInstantiate a random hidden state sequence\n\n1.\n2. For\n\r\u0006\u000f\u0013\u000b\u0013\u000b\u0013\u000b\u0018\u000f#\n- Gibbs sample\n- Update count matrices to re\ufb02ect new\n\n\u0007\u0011\u0010\n\nsented hidden states.\n\n.\n\n\u0001\b\u0007\n\t\u0013\u000f\u0018\u000b\u0013\u000b\u0013\u000b\u0013\u000f\u000b\u0007\r\f\u000e\u0005\n\ngiven hyperparameter settings, count matrices, and observations.\n, the number of repre-\n\n; this may change\n\n\u000f#\u0014\u0013\u0012\r\u0005\n\n\u0007\r\u0010\n\n4\u000e\u0012\n' and\n\n3. End\n4. Update hyperparameters\n5. Goto step 2.\n\n\u0001\u0018\b\n\n\u000f#\u0014\n\n4.1 Gibbs sampling the hidden state sequence\n\ngiven hidden state statistics.\n\nDe\ufb01ne\n\ncontributed by \u0018\u0006\u000f . De\ufb01ne similar items\n\nas the results of removing from> and?\n\n> and\n\nthe transition and emission counts\nrelated to the transition and emission\n\n\n\n*\n\t\n\u001d\n\t\n&\n\t\n\u001d\n.\n\t\n&\n.\n\u0001\n>\n\t\n\t\n?\n\t\n\n*\n\t\n\u001d\n\t\n&\n\t\n\u001d\n.\n\t\n&\n\u000f\n\u0001\n\n\u000f\n\u000f\n4\n\u000f\n\u0014\n\u0014\n?\n\u0014\n>\n\u0014\n?\n'\n\f\f\u0015\t\f\u000b\u0006\u000b\f\u000b\n\u0004\u0003\noracle vectors. An exact Gibbs sweep of the hidden state from \"\noperations, since under the HDP generative process changing \u0018\u0019\u000f affects the probability of\nall subsequent hidden state transitions and emissions. 3 However this computation can be\n\u000f only on the state of\nreasonably approximated in\nits neigbours \u001e\u0018\n\u000f\nsampler, we also sample a set of auxiliary indicator variables \u0006\u0005\u0015\u000f\u001b\t\u0007\u0005\u001e\u000f*12\u0007\u001e\t\u0007\u0005\n\u0001 alongside\n\u0018\u0006\u000f ; each of these is a binary variable denoting whether the oracle was used to generate\n\t\u001b\u0018\n\u001e\u0018\n\n7 , by basing the Gibbs update for \u0018\n?3' .4\n\nIn order to facilitate hyperparameter learning and improve the mixing time of the Gibbs\n\n\u0001 and the total counts\n\n6\u0012\u0007\u001e\t\u000e\u0005\u0010\u000f\u0011\t\u001b\u0018\u0006\u000f*12\u0007\n\nrespectively.\n\ntakes\n\n\u000f*12\u0007\n\n\t\u0001\n\n\t\u0013\u0005\n\n4.2 Hyperparameter optimisation\n\n6\u001e\u0007\n\n\t\u001b\u001a\n\n.0/\n\nare exact:\n\n,\u001d\n/*>\n\n*\u000e\u0004\n\n#\u0010\u000f\u0012\u0011\u0014\u0013\n\nfor large&\n\nWe place vague Gamma priors5 on the hyperparameters \n\n&).\u0006\u0001 . We derive an\n, and\u001d'.\n*\u000e\u0004\n>%\u001f\n\napproximate form for the hyperparameter posteriors from (3) by treating each level of the\nHDPs separately. The following expressions for the posterior for\nare accurate\n\n, while the expressions for& and&\n/\u0019\u0018\n7\u000e\n\n\u0016\u001e7\t\b\u000b\n\n\t\u001b\u001a\n3\u0018\u0017\n3\u0018\u0017\n/\u001e\u0018\n\u000227\u0015\b\u0016\n\n&\r7\n&\u00127\n\n/\u0019\u0018\r\f\n\u001d27\n.0/2\u001d\n&).\n\t'\u001a\n/\u001e\u0018\n\u0016\u001e7\u0015\b\u000b\n\nwhere\u001b\u001c\u001b\ning itself); similarly\u001b\nis the number of represented states that are transitioned to from state ( (includ-\n' and\ncalculated from the indicator variables \u0006\u0005\u0019\u000f\u0011\t\u0007\u0005\n\u0001 . We solve for the maximum a posteriori\n(MAP) setting for each hyperparameter; for example \u001d/.MAP is obtained as the solution to\n\u001f\u001e\u001d\n?3\u001f\n.MAP\u0004\u001c \n\nis the number of possible emissions from state ( .\n\nare the number of times the oracle has been used for the transition and emission processes,\n\nfollowing equation using gradient following techniques such as Newton-Raphson:\n\n\u001d27\n\u000f\u0019\u0011\u0014\u0013\n\u000227\t\b\u001a\n\n.MAP7!\r! \n\n.MAP7#\"\n\n.\f7\n\t'\u001a\n\n.MAP \u0003\n\n\u0007\u0004\u001f\n\n7'/\t\u001d\n\n.0/\n\n.0/\n\n/2\u001d\n\n/\u001e\u0018\n\n\t'\u001a\n\n\u001f\u001e\u001d\n\n\u001f\u001e\u001d\n\n/\t\u001d\n\n/\u001e\u001d\n\n/\u0019\u0018\n\n4.3 In\ufb01nite-state particle \ufb01lter\n\nThe likelihood for a particular observable sequence of symbols involves intractable sums\nover the possible hidden state trajectories. Integrating out the parameters in any HMM\ninduces long range dependencies between states. In particular, in the DP, making the tran-\n\nstandard tricks like dynamic programming. Furthermore, the number of distinct states can\ndis-\n\ngrow with the sequence length as new states are generated. If the chain starts with\u001b\n\n) makes that transition more likely later on in the sequence, so we cannot use\n\" possible distinct states making the total number\n\nsition (\ntinct states, at time \"\nof trajectories over the entire length of the sequence /\n\n3Although the hidden states in an HMM satisfy the Markov condition, integrating out the param-\n\nthere could be\u001b\n\n.\n\n7%$\n\n4This approximation can be motivated in the following way. Consider sampling parameters\n\n, and can therefore be computed without considering its effect on future states.\n\nand\n\nthe shape and inverse-scale parameters.\n\n\u0007\u0011\u001010\n\n2\u0006\u0010\n\nof parameter matrices, which will depend on the count\nand\n\n, the probability of\n\nonly depends on\n\n,\n\neters induces these long-range dependencies.\n\nfrom the posterior distribution\n\n\u000f.-%/\nmatrices. By the Markov property, for a given&\n\u001043\n55\u00046\u001c7\n0GF1H , with\n\n\u000f:9;/)\u0001<9\u0001=?>A@B(48C/+D\n\n'\u0004()&+*\n\t.E\n\n(48\n\n\u0003\n\u0002\n/\n7\n\u0002\n/\n\n\u0014\n>\n\t\n\u0014\n?\n\t\n\u0014\n>\n'\n\t\n\u0014\n.\n\u000f\n\u000f\n\u000f\n\u0001\n*\n\t\n\u001d\n\t\n&\n\t\n\u001d\n.\n\t\n*\n.\n*\n\t\n\u001d\n4\n\f\n3\n3\n7\n#\n\u000e\n\u001f\n\u0010\n\u0007\n\u001d\n1\n/\n1\n/\n*\n7\n1\n\u001f\n\u001f\n\u0004\n*\n7\n1\n/\n\"\n\u000f\n\u000f\n\u0004\n\t\n.\n4\n\u0016\n\t\n7\n#\n\u000e\n\u001f\n\u0010\n\u0007\n\u001d\n.\n#\n\u0017\n1\n1\n/\n\"\n0\n?\n\u001f\n0\n\u0004\n\u001d\n.\n7\n\t\n&\n4\n1\n1\n7\n&\n#\n1\n/\n1\n/\n\n'\n\u0004\n\t\n&\n.\n4\n\u0016\n\t\n1\n\u0017\n1\n\u0017\n7\n&\n#\n\u0017\n1\n/\n7\n1\n/\n\n'\n.\n\u0004\n&\n.\n7\n.\n\u001b\n\n'\n.\n.\n\u000f\n\"\n#\n\u001f\n\u0010\n\u001b\n.\n\u001b\n/\n\"\n0\n0\n\u0004\n\u001d\n\n\u001a\n3\n\u0017\n\u0004\n3\n\u0017\n\n\f\n\u001c\n\u000b\n+\n\u0004\n\u001b\n\u0004\n\n/\n\u001b\n$\n&\n,\n\u0007\n\u0010\n\t\n\u0007\n\t\n5\n=\n0\n8\n9\n\fWe propose estimating the likelihood of a test sequence given a learned model using particle\n\n\t\u0006\u000b\f\u000b\f\u000b\u001a\t\u001b\u0018\u0002\u0001\n\nthe recursive procedure is as speci\ufb01ed\n:\n\n\ufb01ltering. The idea is to start with some number of particles  distributed on the represented\nhidden states according to the \ufb01nal state marginal from the training sequence (some of the \nmay fall onto new states).6 Starting from the set of particles \u001e\u0018\n\u0001 , the tables from\n?3'\n\u0001 , and \"\nthe training sequences \n\"\u0006\u0005\n\t\u0006\u000b\f\u000b\u0006\u000b\u0012\t\u0013\u0005\nbelow, where .0/\nfor each particle \u000b .\n1. Compute \u0007\t\b\n\u0001<'\u0004(42\n2. Calculate \u0007\n/\u000f\u000e\n>\r\f\n\u000f:2\n\u001010\n'\u0004(42\u0006\u0010\u000e*\n(#\r\n\u000f\u0013\u000b\u0013\u000b\u0013\u000b\n3. Resample \f particles\n/\u000f\u000e\n\u0007\t\b\u0014\u0013\n(#\r%>\u0012\u000e\n\u0007\u0016\b\u0015\u0013\n\u0010\u0018\u0017\n\b\u0014\u0013\n\b\u0015\u0013\n\b , \u001b\n4. Update transition and emission tables \u001a\n5. For each \u000b sample forward dynamics:\n\u0007\n\b\n'\u0004(\nThe log likelihood of the test sequence is computed as\"\n\ncause particles to land on novel states. Update \u001a\u001e\b and \u001b\u001f\b .\n\u000f! \n+\u0016#%$\u0011&\n\nstate space, with much of the probability mass concentrated on the represented states, it is\n\n\u000f . Since it is a discrete\n\n7\u0004\u0003\n\u0007\n\b\n\u0010\u0011\u0010\n\nfor each particle.\n\n\u0007\u0019\b\n\n\u000f\u0014\u001a\u001c\b\u0010\u000f\u0015\u001b\u001d\b\n\n6. If\n\n, Goto 1 with\n\n.\n\n\u000f\u001e\"\n\n6\r\u0007\n\n\u0007\n\b\n\n; this may\n\n\t\u0011\u0018\n\n.\n\n.\n\n\u0007\u0019\b\u0014\u0013\n\u001043\n\nfeasible to use '\n7 particles.\n5 Synthetic experiments\n\nand 11.\n\ncatenated copies of (*),+.-0/213/4-5+.)\n\nExp 1: Discovering the number of hidden states We applied the in\ufb01nite HMM infer-\nence algorithm to the ascending-descending observation sequence consisting of 30 con-\n. The most parsimonious HMM which models this\ndata perfectly has exactly 10 hidden states. The in\ufb01nite HMM was initialised with a ran-\ndistinct represented states. In Figure 2 (right) we\nshow how the number of represented states evolves with successive Gibbs sweeps, starting\nconverges to 10, while occasionally exploring 9\n\ndom hidden state sequence, containing\u001b\n. In all cases\u001b\nfrom a variety of initial\u001b\nExp 2: Expansive A sequence of length\nExp 3: Compressive A sequence of length\nwith\u001b\n\n\u001c was generated from a 4-state 8-symbol\n\u001c was generated from a 4-state 3-symbol\n\u0003:9\u0015\u001c distinct states. Figure 3 shows that, over successive Gibbs sweeps and hyper-\n\nHMM with the transition and emission probabilities as shown in Figure 3 (bottom left).\nIn both Exp 2 and Exp 3 the in\ufb01nite HMM was initialised with a hidden state sequence\n\nparameter learning, the count matrices for the in\ufb01nite HMM converge to resemble the true\nprobability matrices as shown on the far left.\n\nHMM with the transition and emission probabilities as shown in Figure 3 (top left).\n\n\u000376\n\u000386\n\n6 Discussion\n\nWe have shown how a two-level Hierarchical Dirichlet Process can be used to de\ufb01ne a non-\nparametric Bayesian HMM. The HDP implicity integrates out the transition and emission\nparameters of the HMM. An advantage of this is that it is no longer necessary to constrain\nthe HMM to have \ufb01nitely many states and observation symbols. The prior over hidden state\ntransitions de\ufb01ned by the HDP is capable of producing a wealth of interesting trajectories\nby varying the three hyperparameters that control it.\nWe have presented the necessary tools for using the in\ufb01nite HMM, namely a linear-time\napproximate Gibbs sampler for inference, equations for hyperparameter learning, and a\nparticle \ufb01lter for likelihood evaluation.\n\n6Different particle initialisations apply if we do not assume that the test sequence immediately\n\nfollows the training sequence.\n\n\u0007\n\u000f\n\u000f\n>\n\t\n>\n'\n\t\n?\n\t\n\u0003\n\f\n\u0018\n\u000f\n4\n\u0005\n\u0007\n\u000f\n\u0007\n\u0001\n\u0016\n/\n\u0018\n\u000f\n\u0005\n\u000f\n7\n\u0010\n\u0010\n*\n\u0007\n\u0010\n\u0001\n\u0010\n/\n\u0010\n\u0001\n\b\n\u0007\n\b\n2\n\t\n\t\n/\n\u0010\n6\n\u0010\n(\n\u0007\n\u0010\n\u000f\n\u0010\n/\n\b\n\u0010\n3\n\t\n6\n\u0007\n\t\n*\n\u0007\n\u0010\n\u0001\n\u0010\n/\n\u000f\n\u0001\n\n\u000f\n/\n\u001b\n\u001c\n\u001c\n\fTrue transition and\nemission probability\n\nmatrices used for Exp 2\n\nTrue transition and\nemission probability\n\nmatrices used for Exp 3\n\n\u0002\u0004\u0003\n\n\u0001\u0015\u000f\n\n\u0002\u0004\u0003\n\n\u0001\u0015\u000f\n\n\u0001\u0015\u000f\n\n\u0005\u0004\u0003\n\n\u0005\u0004\u0003\n\n\u0001\u0015\u000f\n\n\u0001\u0015\u000f\n\n\u0003\u0004\u0003\n\n\u0003\u0006\u0003\n\n\u0001\u0015\u000f\n\nFigure 3:\nThe far left pair of Hinton diagrams represent the true transition and emission prob-\nabilities used to generate the data for each experiment 2 and 3 (up to a permutation of the hidden\nstates; lighter boxes correspond to higher values). (top row) Exp 2: Expansive HMM. Count matrix\nsweeps of Gibbs sampling. (bottom row) Exp 3:\npairs\nCompressive HMM. Similar to top row displaying count matrices after\nsweeps of\nGibbs sampling. In both rows the display after a single Gibbs sweep has been reduced in size for\nclarity.\n\nare displayed after\n\n\u0001\"\r\u0006\u000f\u0018\r\u0018\u0007\u0010\u0007!\u000f\u0018\n\n\u0010\u000f\n\n\u001f\u0010\u0007!\u000f\u0018\r\u0010\r\u0018\u0005!\u000f\u0018\n\n\u0001\u0019\u001a)\u000f\u0015\u001b\n\n\u0005!\u000f\n\n\u0003\r\f\u0006\u0007\n\n\u0007\u0004\u0005\n\n\u0007\u0004\u0005\n\n\u0001\u0015\u000f\n\n\u000f\t\b\u0004\n\n\u0001\u0015\u000f\t\b\u0004\n\nOn synthetic data we have shown that the in\ufb01nite HMM discovers both the appropriate\nnumber of states required to model the data and the structure of the emission and transition\nmatrices. It is important to emphasise that although the count matrices found by the in\ufb01nite\nHMM resemble point estimates of HMM parameters (e.g. Figure 3), they are better thought\nof as the suf\ufb01cient statistics for the HDP posterior distribution over parameters.\nWe believe that for many problems the in\ufb01nite HMM\u2019s \ufb02exibile nature and its ability to\nautomatically determine the required number of hidden states make it superior to the con-\nventional treatment of HMMs with its associated dif\ufb01cult model selection problem. While\nthe results in this paper are promising, they are limited to synthetic data; in future we hope\nto explore the potential of this model on real-world problems.\n\nAcknowledgements\nThe authors would like to thank David Mackay for suggesting the use of an oracle, and\nQuaid Morris for his Perl expertise.\n\nReferences\n[1] C. E. Antoniak. Mixtures of Dirichlet processes with applications to Bayesian nonparametric\n\nproblems. Annals of Statistics, 2(6):1152\u20131174, 1974.\n\n[2] T. S. Ferguson. A Bayesian analysis of some nonparametric problems. Annals of Statistics,\n\n1(2):209\u2013230, March 1973.\n\n[3] D. J. C. MacKay. Ensemble learning for hidden Markov models. Technical report, Cavendish\n\nLaboratory, University of Cambridge, 1997.\n\n[4] D. J. C. MacKay and L. C. Peto. A hierarchical Dirichlet language model. Natural Language\n\nEngineering, 1(3):1\u201319, 1995.\n\n[5] R. M. Neal. Markov chain sampling methods for Dirichlet process mixture models. Technical\n\nReport 9815, Dept. of Statistics, University of Toronto, 1998.\n\n[6] L. R. Rabiner and B. H. Juang. An introduction to hidden Markov models. IEEE Acoustics,\n\nSpeech & Signal Processing Magazine, 3:4\u201316, 1986.\n\n[7] C. E. Rasmussen. The in\ufb01nite Gaussian mixture model.\nProcessing Systems 12, Cambridge, MA, 2000. MIT Press.\n\nIn Advances in Neural Information\n\n[8] A. Stolcke and S. Omohundro. Hidden Markov model induction by Bayesian model merging. In\nS. J. Hanson, J. D. Cowan, and C. L. Giles, editors, Advances in Neural Information Processing\nSystems 5, pages 11\u201318, San Francisco, CA, 1993. Morgan Kaufmann.\n\n\n\u000f\n0\n\u0013\n0\n\u0013\n\n\u000f\n\u0013\n\u0013\n\n\u000f\n0\n0\n\u0005\n\u0013\n0\n0\n\u0005\n\u0013\n\n\u000f\n0\n\u0013\n0\n\u0013\n\n\u000f\n0\n\u0013\n0\n\u0013\n\n\u000f\n0\n\u0013\n0\n\u0013\n\n\u000f\n0\n\u0013\n0\n\u0013\n\n\u0003\n\u0013\n\u0003\n\u0013\n\u0005\n\u0001\n\u0005\n\u0007\n\u0005\n\u000b\n\u0005\n\f", "award": [], "sourceid": 1956, "authors": [{"given_name": "Matthew", "family_name": "Beal", "institution": null}, {"given_name": "Zoubin", "family_name": "Ghahramani", "institution": null}, {"given_name": "Carl", "family_name": "Rasmussen", "institution": null}]}