{"title": "Mixed Membership Stochastic Blockmodels", "book": "Advances in Neural Information Processing Systems", "page_first": 33, "page_last": 40, "abstract": "Observations consisting of measurements on relationships for pairs of objects arise in many settings, such as protein interaction and gene regulatory networks, collections of author-recipient email, and social networks. Analyzing such data with probabilisic models can be delicate because the simple exchangeability assumptions underlying many boilerplate models no longer hold. In this paper, we describe a class of latent variable models of such data called Mixed Membership Stochastic Blockmodels. This model extends blockmodels for relational data to ones which capture mixed membership latent relational structure, thus providing an object-specific low-dimensional representation. We develop a general variational inference algorithm for fast approximate posterior inference. We explore applications to social networks and protein interaction networks.", "full_text": "Mixed Membership Stochastic Blockmodels\n\nEdoardo M. Airoldi 1,2, David M. Blei 1, Stephen E. Fienberg 3,4 & Eric P. Xing 4\u2217\n\neairoldi@Princeton.EDU\n\n1 Department of Computer Science, 2 Lewis-Sigler Institute, Princeton University\n\n3 Department of Statistics, 4 School of Computer Science, Carnegie Mellon University\n\nAbstract\n\nIn many settings, such as protein interactions and gene regulatory networks, col-\nlections of author-recipient email, and social networks, the data consist of pair-\nwise measurements, e.g., presence or absence of links between pairs of objects.\nAnalyzing such data with probabilistic models requires non-standard assumptions,\nsince the usual independence or exchangeability assumptions no longer hold. In\nthis paper, we introduce a class of latent variable models for pairwise measure-\nments: mixed membership stochastic blockmodels. Models in this class combine\na global model of dense patches of connectivity (blockmodel) with a local model\nto instantiate node-speci\ufb01c variability in the connections (mixed membership).\nWe develop a general variational inference algorithm for fast approximate poste-\nrior inference. We demonstrate the advantages of mixed membership stochastic\nblockmodel with applications to social networks and protein interaction networks.\n\n1 Introduction\n\nThe problem of modeling relational information among objects, such as pairwise relations repre-\nsented as graphs, arises in a number of settings in machine learning. For example, scienti\ufb01c liter-\nature connects papers by citation, the Web connects pages by links, and protein-protein interaction\ndata connect proteins by physical interaction records. In these settings, we often wish to infer hidden\nattributes of the objects from the pairwise observations. For example, we might want to compute\na clustering of the web-pages, predict the functions of a protein, or assess the degree of relevance\nof a scienti\ufb01c abstract to a scholar\u2019s query. Unlike traditional attribute data measured over indi-\nvidual objects, relational data violate the classical independence or exchangeability assumptions\nmade in machine learning and statistics. The objects are dependent by their very nature, and this\ninterdependence suggests that a different set of assumptions is more appropriate.\nRecently proposed models aim at resolving relational information into a collection of connectivity\nmotifs. Such models are based on assumptions that often ignore useful technical necessities, or im-\nportant empirical regularities. For instance, exponential random graph models [11] summarize the\nvariability in a collection of paired measurements with a set of relational motifs, but do not provide a\nrepresentation useful for making unit-speci\ufb01c predictions. Latent space models [4] project individ-\nual units of analysis into a low-dimensional latent space, but do not provide a group structure into\nsuch space useful for clustering. Stochastic blockmodels [8, 6] resolve paired measurements into\ngroups and connectivity between pairs of groups, but constrain each unit to instantiate the connec-\ntivity patterns of a single group as observed in most applications. Mixed membership models, such\nas latent Dirichlet allocation [1], have emerged in recent years as a \ufb02exible modeling tool for data\nwhere the single group assumption is violated by the heterogeneity within a unit of analysis\u2014e.g., a\ndocument, or a node in a graph. They have been successfully applied in many domains, such as doc-\nument analysis [1], image processing [7], and population genetics [9]. Mixed membership models\nassociate each unit of analysis with multiple groups rather than a single groups, via a membership\n\n\u2217A longer version of this work is available online, at http://jmlr.csail.mit.edu/papers/v9/airoldi08a.html\n\n1\n\n\fprobability-like vector. The concurrent membership of a data in different groups can capture its dif-\nferent aspects, such as different underlying topics for words constituting each document. The mixed\nmembership formalism is a particularly natural idea for relational data, where the objects can bear\nmultiple latent roles or cluster-memberships that in\ufb02uence their relationships to others. Existing\nmixed membership models, however, are not appropriate for relational data because they assume\nthat the data are conditionally independent given their latent membership vectors. Conditional inde-\npendence assumptions that technically instantiate mixed membership in recent work, however, are\ninappropriate for the relational data settings. In such settings, an objects is described by its rela-\ntionships to others. Thus assuming that the ensemble of mixed membership vectors help govern the\nrelationships of each object would be more appropriate.\nHere we develop mixed membership models for relational data and we describe a fast variational\ninference algorithm for inference and estimation. Our model captures the multiple roles that ob-\njects exhibit in interaction with others, and the relationships between those roles in determining the\nobserved interaction matrix. We apply our model to protein interaction and social networks.\n\n2 The Basic Mixed Membership Blockmodel\nObservations consist of pairwise measurements, represented as a graph G = (N , Y ), where Y (p, q)\ndenotes the measurement taken on the pair of nodes (p, q). In this section we consider observations\nconsisting of a single binary matrix, where Y (p, q) \u2208 {0, 1}, i.e., the data can be represented with a\ndirected graph. The model generalizes to two important settings, however, as we discuss below\u2014a\ncollection of matrices and/or other types of measurements. We summarize a collection of pairwise\nmeasurements with a mapping from nodes to sets of nodes, called blocks, and pairwise relations\namong the blocks themselves. Intuitively, the inference process aims at identifying nodes that are\nsimilar to one another in terms of their connectivity to blocks of nodes. Similar nodes are mapped\nto the same block.\nIndividual nodes are allowed to instantiate connectivity patterns of multiple\nblocks. Thus, the goal of the analysis with a Mixed Membership Blockmodel (MMB) is to identify\n(i) the mixed membership mapping of nodes, i.e., the units of analysis, to a \ufb01xed number of blocks,\nK, and (ii) the pairwise relations among the blocks. Pairwise measurements among N nodes are\nthen generated according to latent distributions of block-membership for each node and a matrix of\nblock-to-block interaction strength. Latent per-node distributions are speci\ufb01ed by simplicial vectors.\nEach node is associated with a randomly drawn vector, say ~\u03c0i for node i, where \u03c0i,g denotes the\nprobability of node i belonging to group g. In this fractional sense, each node can belong to multiple\ngroups with different degrees of membership. The probabilities of interactions between different\ngroups are de\ufb01ned by a matrix of Bernoulli rates B(K\u00d7K), where B(g, h) represents the probability\nof having a connection from a node in group g to a node in group h. The indicator vector ~zp\u2192q\ndenotes the speci\ufb01c block membership of node p when it connects to node q, while ~zp\u2190q denotes\nthe speci\ufb01c block membership of node q when it is connected from node p. The complete generative\nprocess for a graph G = (N , Y ) is as follows:\n\n\u2022 For each node p \u2208 N :\n\n\u2022 For each pair of nodes (p, q) \u2208 N \u00d7 N :\n\n\u2013 Draw a K dimensional mixed membership vector ~\u03c0p \u223c Dirichlet(cid:0) ~\u03b1(cid:1).\n\u2013 Draw membership indicator for the initiator, ~zp\u2192q \u223c Multinomial(cid:0) ~\u03c0p\n\u2013 Draw membership indicator for the receiver, ~zq\u2192p \u223c Multinomial(cid:0) ~\u03c0q\n\u2013 Sample the value of their interaction, Y (p, q) \u223c Bernoulli(cid:0) ~z >\n\n(cid:1).\n(cid:1).\n(cid:1).\n\np\u2192qB ~zp\u2190q\n\nNote that the group membership of each node is context dependent, i.e., each node may assume\ndifferent membership when interacting with different peers. Statistically, each node is an admixture\nof group-speci\ufb01c interactions. The two sets of latent group indicators are denoted by {~zp\u2192q : p, q \u2208\nN} =: Z\u2192 and {~zp\u2190q : p, q \u2208 N} =: Z\u2190. Further, the pairs of group memberships that underlie\ninteractions, e.g., (~zp\u2192q, ~zp\u2190q) for Y (p, q), need not be equal; this fact is useful for characterizing\nasymmetric interaction networks. Equality may be enforced when modeling symmetric interactions.\nThe joint probability of the data Y and the latent variables {~\u03c01:N , Z\u2192, Z\u2190} sampled according to\nthe MMB is:\nP (~\u03c0p|~\u03b1).\n\nP (Y (p, q)|~zp\u2192q, ~zp\u2190q, B)P (~zp\u2192q|~\u03c0p)P (~zp\u2190q|~\u03c0q)Y\n\np(Y, ~\u03c01:N , Z\u2192, Z\u2190|~\u03b1, B) =Y\n\np,q\n\np\n\n2\n\n\f\u03b1\n\n\u03c0\n1\n\n\u03c0\n2\n\n\u03c0\n3\n\n.\n.\n.\n\n\u03c0\nn\n\nz 1\u21901\n\nz 1\u21902\n\nz 1\u21903\n\nz 1\u21921\n\ny 11\n\nz 1\u21922\n\ny 12\n\nz 1\u21923\n\ny 13\n\nz 2\u21901\n\nz 2\u21902\n\nz 2\u21903\n\nz 1\u2190N\n\nz 1\u2192N\n\ny 1N\n\nz 2\u2190N\n\n. . .\n\n. . .\n\nz 2\u21921\n\ny 21\n\nz 2\u21922\n\ny 22\n\nz 2\u21923\n\ny 23\n\nz 2\u2192N\n\ny 2N\n\nB\n\nFigure 1: The graphical model\nof the mixed membership block-\nmodel (MMB). We did not draw all\nthe arrows out of the block model\nB for clarity. All the pairwise mea-\nsurements, Y (p, q), depend on it.\n\nz 3\u21901\n\nz 3\u21902\n\nz 3\u21903\n\nz 3\u21921\n\ny 31\n\nz 3\u21922\n\ny 32\n\nz 3\u21923\n\ny 33\n\n.\n.\n.\n\n.\n.\n.\n\n.\n.\n.\n\nz N\u21901\n\nz N\u21902\n\nz N\u21903\n\nz 3\u2190N\n\nz 3\u2192N\n\ny 3N\n\n.\n.\n.\n\nz N\u2190N\n\n. . .\n\n.\n\n.\n\n.\n\n. . .\n\nz N\u21921\n\ny N1\n\nz N\u21922\n\ny N2\n\nz 1\u21921\n\ny N3\n\nz N\u2192N\n\ny NN\n\nIntroducing Sparsity. Adjacency matrices encoding binary pairwise measurements often contain\na large amount of zeros, or non- interactions; they are sparse. It is useful to distinguish two sources\nof non- interaction: they may be the result of the rarity of interactions in general, or they may be\nan indication that the pair of relevant blocks rarely interact. In applications to social sciences, for\ninstance, nodes may represent people and blocks may represent social communities. In this setting,\nit is reasonable to expect that a large portion of the non- interactions is due to limited opportunities\nof contact between people in a large population, or by design of the questionnaire, rather than due to\ndeliberate choices, the structure of which the blockmodel is trying to estimate. It is useful to account\nfor these two sources of sparsity at the model level. A good estimate of the portion of zeros that\nshould not be explained by the blockmodel B reduces the bias of the estimates of B\u2019s elements.\nWe introduce a sparsity parameter \u03c1 \u2208 [0, 1] in the model above to characterize the source of non-\ninteraction. Instead of sampling a relation Y (p.q) directly the Bernoulli with parameter speci\ufb01ed as\nabove, we down- weight the probability of successful interaction to (1 \u2212 \u03c1) \u00b7 ~z >\np\u2192qB ~zp\u2190q. This is\nthe result of assuming that the probability of a non- interaction comes from a mixture, 1 \u2212 \u03c3pq =\n(1 \u2212 \u03c1) \u00b7 ~z >\np\u2192q(1 \u2212 B) ~zp\u2190q + \u03c1, where the weight \u03c1 capture the portion zeros that should not be\nexplained by the blockmodel B. A large value of \u03c1 will cause the interactions in the matrix to be\nweighted more than non- interactions, in determining plausible values for {~\u03b1, B, ~\u03c01:N}.\nRecall that {~\u03b1, B} are constant quantities to be estimated, while {~\u03c01:N , Z\u2192, Z\u2190} are unknown vari-\nable quantities whose posterior distribution needs to be determined. Below, we detail the variational\nexpectation- maximization (EM) procedure to carry out approximate estimation and inference.\n\n2.1 Variational E- Step\n\nDuring the E- step, we update the posterior distribution over the unknown variable quantities\n{~\u03c01:N , Z\u2192, Z\u2190}. The normalizing constant of the posterior is the marginal probability of the data,\nwhich requires an intractable integral over the simplicial vectors ~\u03c0p,\n\nZ\n\nX\n\np(Y | ~\u03b1, B) =\n\np(Y, ~\u03c01:N , Z\u2192, Z\u2190|~\u03b1, B).\n\n(1)\n\n~\u03c01:N\n\nzp\u2190q,zp\u2192q\n\nWe appeal to mean- \ufb01eld variational methods [5] to approximate the posterior of interest. The main\nidea behind variational methods is to posit a simple distribution of the latent variables with free\nparameters, which are \ufb01t to make the approximation close in Kullback- Leibler divergence to the\ntrue posterior of interest. The log of the marginal probability in Equation 1 can be bound as follows,\n\n(cid:2) log p(Y, ~\u03c01:N , Z\u2192, Z\u2190|\u03b1, B)(cid:3) \u2212Eq\n\n(cid:2) log q(~\u03c01:N , Z\u2192, Z\u2190)(cid:3),\n\nlog p(Y | \u03b1, B) \u2265 Eq\n\n(2)\n\nby introducing a distribution of the latent variables q that depends on a set of free parameters.\nWe specify q as the mean- \ufb01eld fully- factorized family, q(~\u03c01:N , Z\u2192, Z\u2190 | ~\u03b31:N , \u03a6\u2192, \u03a6\u2190), where\n{~\u03b31:N , \u03a6\u2192, \u03a6\u2190} is the set of free variational parameters that must be set to tighten the bound. We\n\n3\n\n\f\u02c6\u03c6p\u2192q,g \u221d e\n\nEq\n\n\u02c6\u03c6p\u2190q,h \u221d e\n\nEq\n\ntighten the bound with respect to the variational parameters, to minimize the KL divergence between\nq and the true posterior. The update for the variational multinomial parameters is\n\nh\n\n(cid:18)\n(cid:3) \u00b7Y\n(cid:2)log \u03c0p,g\n(cid:18)\n(cid:2)log \u03c0q,h\n(cid:3) \u00b7Y\n\u02c6\u03b3p,k = \u03b1k +X\n\nB(g, h)Y (p,q)\u00b7(cid:0) 1 \u2212 B(g, h)(cid:1)1\u2212Y (p,q)(cid:19)\u03c6p\u2190q,h\nB(g, h)Y (p,q)\u00b7(cid:0) 1 \u2212 B(g, h)(cid:1)1\u2212Y (p,q)(cid:19)\u03c6p\u2192q,g\n\u03c6p\u2192q,k +X\n\n\u03c6p\u2190q,k,\n\ng\n\n(3)\n\n,\n\n(4)\n\n(5)\n\nfor g, h = 1, . . . , K. The update for the variational Dirichlet parameters \u03b3p,k is\n\nq\nfor all nodes p = 1, . . . , N and k = 1, . . . , K.\n\nq\n\nNested Variational Inference. To improve convergence, we developed a nested variational infer-\nence scheme based on an alternative schedule of updates to the traditional ordering [5]. In a na\u00a8\u0131ve\niteration scheme for variational inference, one initializes the variational Dirichlet parameters ~\u03b31:N\nand the variational multinomial parameters (~\u03c6p\u2192q, ~\u03c6p\u2190q) to non-informative values, and then it-\nerates until convergence the following two steps: (i) update ~\u03c6p\u2192q and \u03c6p\u2190q for all edges (p, q),\nand (ii) update ~\u03b3p for all nodes p \u2208 N . At each variational inference cycle one needs to allocate\nN K + 2N 2K scalars. In our experiments, the na\u00a8\u0131ve variational algorithm often failed to converge,\nor converged only after many iterations. We attribute this behavior to dependence between ~\u03b31:N and\nB in the model, which is not accounted for by the na\u00a8\u0131ve algorithm. The nested variational inference\nalgorithm retains portion of this dependence across iterations by following a particular path to con-\nvergence. We keep the block of free parameters (~\u03c6p\u2192q, ~\u03c6p\u2190q) at their optimal values conditionally\non the other variational parameters. These parameters are involved in the updates of parameters\nin ~\u03b31:N and in B, thus effectively providing a channel to maintain some dependence among them.\nFrom a computational perspective, the nested algorithm trades time for space thus allowing us to\ndeal with large graphs. At each variational cycle we allocate N K + 2K scalars only. The algorithm\ncan be parallelized, and, empirically, leads to a better likelihood bound per unit of running time.\n\n2.2 M-Step\n\nDuring the M-step, we maximize the lower bound in Equation 2, used as a surrogate for the like-\nlihood, with respect to the unknown constants {~\u03b1, B}. In other words, we compute the empirical\nBayes estimates of the hyper-parameters. The M-step is equivalent to \ufb01nding the MLE using ex-\npected suf\ufb01cient statistics under the variational distribution. We consider the maximization step for\neach parameter in turn. A closed form solution for the approximate maximum likelihood estimate\nof ~\u03b1 does not exist. We used linear-time Newton-Raphson, with gradient and Hessian\n\n(cid:18)\n(cid:18)\n\n(cid:1) \u2212\u03c8(\u03b1k)\n\n(cid:19)\n+X\n\u03c8(cid:0)X\nI(k1=k2) \u00b7 \u03c80(\u03b1k1) \u2212 \u03c80(cid:0)X\n\n\u03b1k\n\nk\n\np\n\n(cid:18)\n\u03c8(\u03b3p,k) \u2212 \u03c8(cid:0)X\n(cid:1)(cid:19)\n\nk\n\n\u03b1k\n\n,\n\n= N\n\n= N\n\n(cid:1)(cid:19)\n\n\u03b3p,k\n\n, and\n\n\u2202L~\u03b1\n\u2202\u03b1k\n\u2202L~\u03b1\n\u2202\u03b1k1\u03b1k2\n\nto \ufb01nd optimal values for ~\u03b1, numerically. The approximate MLE of B is\n\n\u02c6B(g, h) =\n\nk\n\nP\nP\np,q Y (p, q) \u00b7 \u03c6p\u2192qg \u03c6p\u2190qh\nP\n\n(cid:0) 1 \u2212 Y (p, q)(cid:1) \u00b7(cid:0)P\nP\n\np,q \u03c6p\u2192qg \u03c6p\u2190qh\n\ng,h \u03c6p\u2192qg \u03c6p\u2190qh\n\np,q\n\n,\n\ng,h \u03c6p\u2192qg \u03c6p\u2190qh\n\n(cid:1)\n\n.\n\nP\n\n\u02c6\u03c1 =\n\np,q\n\nfor every pair (g, h) \u2208 [1, K]2. Finally, the approximate MLE of the sparsity parameter \u03c1 is\n\nwith \u02c6d =P\n\nAlternatively, we can \ufb01x \u03c1 prior to the analysis; the density of the interaction matrix is estimated\np,q Y (p, q)/N 2, and the sparsity parameter is set to \u02dc\u03c1 = (1 \u2212 \u02c6d). This latter estimator\n\n(6)\n\n(7)\n\n4\n\n\fattributes all the information in the non-interactions to the point mass, i.e., to latent sources other\nthan the block model B or the mixed membership vectors ~\u03c01:N . It can be used, however, as a quick\nrecipe to reduce the computational burden during exploratory analyses.\nIn our setting, model selection\nSeveral model selection strategies exist for hierarchical models.\ntranslates into the choice of the number of blocks, K. Below, we chose K with held-out likelihood\nin a cross-validation experiment, on large networks, and with approximate BIC, on small networks.\n\n2.3 Summarizing and De-Noising Pairwise Measurements\n\nIt is useful to consider two data analysis perspectives the MMB can offer: (i) it summarizes the\ndata, Y , in terms of the global blockmodel, B, and the node-speci\ufb01c mixed memberships, \u03a0s,\n(ii) it de-noises the data, Y , in terms of the global blockmodel, B, and interaction-speci\ufb01c single\nmemberships, Zs.\nIn both cases the model depends on a small set of unknown constants to be\nestimated: \u03b1, and B. The likelihood is the same in both cases, although, the reasons for including the\nset of latent variables Zs differ. When summarizing data, we could integrate out the Zs analytically;\nthis leads to numerical optimization of a smaller set of variational parameters, \u0393s. We choose to\nkeep the Zs to simplify inference. When de-noising, the Zs are instrumental in estimating posterior\nexpectations of each interactions individually\u2014a network analog to the Kalman Filter. The posterior\nexpectations of an interaction is computed as ~\u03c0p\n\n0 B ~\u03c6p\u2190q, in the two cases.\n\n0 B ~\u03c0q, and ~\u03c6p\u2192q\n\n3 Empirical Results\n\nWe evaluated the MMB on simulated data and on three collections of pairwise measurements. Re-\nsults on simulated data sampled accordingly to the model show that variational EM accurately recov-\ners the mixed membership map, ~\u03c01:N , and the blockmodel, B. Cross-validation suggests an accurate\nestimate for K. Nested variational scheduling of parameter updates makes inference parallelizable\nand a typically reaches a better solution than the na\u00a8\u0131ve scheduling.\nFirst we consider, whom-do-like relations among 18 novices in a New England monastery. The\nunsupervised analysis demonstrates the type of patterns that MMB recovers from data, and allows us\nto contrast the summaries of the original measurements achieved through prediction and de-noising.\nThe data was collected by Sampson during his stay at the monastery, while novices were preparing\nto join the monastic order [10]. Sampson\u2019s original analysis is rooted in direct anthropological\nobservations. He made a strong case for the existence of tight factions among the novices: the loyal\nopposition (whose members joined the monastery \ufb01rst), the young turks (who joined later on), the\noutcasts (who were not accepted in the two main factions), and the waverers (who did not take sides).\nThe events that took place during Sampson\u2019s stay at the monastery supported his observations\u2014\nmembers of the young turks resigned or were expelled over religious differences (John and Gregory).\nScholars in the social sciences typically regard the faction labels assigned by Sampson to the novices\n(and his conclusions, more in general) as ground truth to the extent of assessing the quality of results\nof quantitative analyses; we shall do the same here. Using the nested variational EM algorithm\nabove, we \ufb01t an array of mixed membership blockmodels with different values of K, and collected\nmodel estimates {\u02c6\u03b1, \u02c6B} and posterior mixed membership vectors ~\u03c01:18 for the novices. We used an\napproximation of BIC to choose the value of K supported by the data. This criterion selects \u02c6K = 3,\nthe same number of proper groups that Sampson identi\ufb01ed based on anthropological observations\u2014\nthe waverers are interstitial members, rather than a group. Figure 2 shows the patterns that the mixed\nmembership blockmodel with \u02c6K = 3 recovers from data. In particular, the top-left panel shows a\ngraphical representation of the blockmodel \u02c6B. The block that we can identify a-posteriori with the\nloyal opposition is portrayed as central to the monastery, while the block identi\ufb01ed with the outcasts\nshows the lowest internal coherence, in accordance with Sampson\u2019s observations. The top-right\npanel illustrates the posterior means of the mixed membership scores, E[~\u03c0|Y ], for the 18 monks in\nthe monastery. The model (softly) partitions the monks according to Sampson\u2019s classi\ufb01cation, with\nYoung Turks, Loyal Opposition, and Outcasts dominating each corner respectively. Notably, we can\nquantify the central role played by John Bosco and Gregory, who exhibit relations in all three groups,\nas well as the uncertain af\ufb01liations of Ramuald and Victor; Amand\u2019s uncertain af\ufb01liation, however,\nis not captured. The bottom panels contrast the different resolution of the original adjacency matrix\nof whom-do-like sociometric relations (left panel) obtained with the two analyses MMB enables.\n\n5\n\n\fFigure 2: Top- Left: Estimated blockmodel, \u02c6B. Top- Right: Posterior mixed membership vectors,\n~\u03c01:18, projected in the simplex. The estimates correspond to a model with \u02c6B top- left, and \u02c6\u03b1 = 0.058.\nNumbered points can be mapped to monks\u2019 names using the legend on the right. The colors identify\nthe four factions de\ufb01ned by Sampson\u2019s anthropological observations. Bottom: Original adjacency\nmatrix of whom- do- like sociometric relations (left), relations predicted using approximate MLEs\nfor ~\u03c01:N and B (center), and relations de- noised using the model including Zs indicators (right).\nIf the goal of the analysis if to \ufb01nd a parsimonious summary of the data, the amount of relational\ninformation that is captured by in \u02c6\u03b1, \u02c6B, and E[~\u03c0|Y ] leads to a coarse reconstruction of the original\nsociomatrix (central panel).\nIf the goal of the analysis if to de- noising a collection of pairwise\nmeasurements, the amount of relational information that is revealed by \u02c6\u03b1, \u02c6B and E[Z\u2192, Z\u2190|Y ] leads\nto a \ufb01ner reconstruction of the original sociomatrix, Y \u2014relations in Y are re- weighted according to\nhow much they make sense to the model (right panel). Substantively, the unsupervised analysis of the\nsociometric relations with MMB offers quantitative support to several of Sampson\u2019s observations.\nSecond, we consider a friendship network among a group of 69 students in grades 7\u201312. The analysis\nhere directly compares clustering results obtained with MMB to published results obtained with\ncompeting models, in a setting where a fair amount of social segregation is expected [2, 3].\nThe data is a collection of friendship relations among 69 students in a school surveyed in the Na-\ntional Study of Adolescent Health. The original population in the school of interest consisted of 71\nstudents. Two students expressed no friendship preferences and were excluded from the analysis.\nWe used variational EM algorithm to \ufb01t an array of mixed membership blockmodels with different\nvalues of K, collected model estimates, and used an approximation to BIC to select K. This proce-\ndure identi\ufb01ed \u02c6K = 6 as the model- size that best explains the data; note that six is the number of\ngrade- groups in the student population. The blocks are clearly interpretable a- posteriori in terms of\ngrades, thus providing a mapping between grades and blocks. Conditionally on such a mapping, we\nassign students to the grade they are most associated with, according to their posterior- mean mixed\nmembership vectors, E[~\u03c0n|Y ]. To be fair in the comparison with competing models, we assign\nstudents with a unique grade\u2014despite MMB allows for mixed membership. Table 1 computes the\ncorrespondence of grades to blocks by quoting the number of students in each grade- block pair, for\nMMB versus the mixture blockmodel (MB) in [2], and the latent space cluster model (LSCM) in\n[3]. The higher the sum of counts on diagonal elements is the better is the correspondence, while the\nhigher the sum of counts off diagonal elements is the worse is the correspondence. MMB performs\nbest by allocating 63 students to their grades, versus 57 of MB, and 37 of LSCM. Correspondence\nonly partially partially captures goodness of \ufb01t, however, it is a good metric in the setting we con-\nsider, where a fair amount of clustering is present. The extra- \ufb02exibility MMB offers over MB and\nLSCM reduces bias in the prediction of the membership of students to blocks, in this problem.\nIn other words, mixed membership does not absorb noise in this example, rather it accommodates\nvariability in the friendship relation that is instrumental in producing better predictions.\n\n6\n\n\fGrade\n7\n8\n9\n10\n11\n12\n\n1\n13\n0\n0\n0\n0\n0\n\n2\n1\n9\n0\n0\n0\n0\n\n3\n0\n2\n16\n0\n1\n0\n\n4\n0\n0\n0\n10\n0\n0\n\n5\n0\n0\n0\n0\n11\n0\n\n6\n0\n1\n0\n0\n1\n4\n\n1\n13\n0\n0\n0\n0\n0\n\n2\n1\n10\n0\n0\n0\n0\n\nMMB Clusters\n\nMB Clusters\n\n3\n0\n2\n10\n0\n1\n0\n\n4\n0\n0\n0\n10\n0\n0\n\n5\n0\n0\n0\n0\n11\n0\n\n6\n0\n0\n6\n0\n1\n4\n\n1\n13\n0\n0\n0\n0\n0\n\nLSCM Clusters\n2\n1\n11\n0\n0\n0\n0\n\n3\n0\n1\n7\n0\n0\n0\n\n4\n0\n0\n6\n0\n0\n0\n\n5\n0\n0\n3\n3\n3\n0\n\n6\n0\n0\n0\n7\n10\n4\n\nTable 1: Grade levels versus (highest) expected posterior membership for the 69 students, according\nto three alternative models. MMSB is the proposed mixed membership stochastic blockmodel, MSB\nis the mixture blockmodel in [2], and LSCM is the latent space cluster model in [3].\n\nThird, we consider physical interactions among 871 proteins in yeast. The analysis allows us to eval-\nuate the utility of MMB in summarizing and de-noising complex connectivity patterns quantitatively,\nusing an independent set of functional annotations\u2014consider two models that suggest different sets\nof interactions as reliable; we prefer the model that reveals functionally relevant interactions.\nThe pairwise measurements consist of a hand-curated collection of physical protein interactions\nmade available by the Munich Institute of Protein Sequencing (MIPS). The yeast genome database\nprovides independent functional annotations for each protein, which we use for evaluating the func-\ntional content of the protein networks estimated with the MMB from the MIPS data, as detailed\nbelow. We explored a large model space, K = 2 . . . 225, and used \ufb01ve-fold cross-validation to iden-\ntify a blockmodel B that reduces the dimensionality of the physical interactions among proteins in\nthe training set, while revealing robust aspects of connectivity that can be leveraged to predict phys-\nical interactions among proteins in the test set. We determined that a fairly parsimonious model,\nK = 50, provides a good description of the observed physical interaction network. This \ufb01nding\nsupports the hypothesis that proteins derived from the MIPS data are interpretable in terms func-\ntional biological contexts. Alternatively, the blocks might encode signal at a \ufb01ner resolution, such\nas that of protein complexes. If that was the case, however, we would expect the optimal number of\nblocks to be signi\ufb01cantly higher; 871/5 \u2248 175, given an average size of \ufb01ve proteins in a protein\ncomplex. We then evaluated the functional content of the posterior induced by MMB. The goal is\nto assess to what extent MMB reveals substantive information about the functionality of proteins\nthat can be used to inform subsequent analyses. To do this, \ufb01rst, we \ufb01t a model on the whole data\nset to estimate the blockmodel, B(50\u00d750), and the mixed membership vectors between proteins and\nblocks, ~\u03c01:871, and second, we either impute physical interactions by thresholding the posterior\nexpectations computed using blockmodel and node-speci\ufb01c memberships (summarization task), or\nwe de-noise the observed interactions using blockmodel and pair-speci\ufb01c memberships (de-noising\ntask). Posterior expectations of each interaction are in [0, 1]. Thresholding such expectations at q,\nfor instance, leads to a collection of binary physical interactions that are at reliable with probabil-\nity p \u2265 q. We used an independent set of functional annotations from the yeast database (SGD\nat www.yeastgenome.org) to decide which interactions are functionally meaningful; namely those\nbetween pairs of proteins that share at least one functional annotation. In this sense, between two\nmodels that suggest different sets of interactions as reliable, our evaluation assigns a higher score\n\nFigure 3: Functional content of the MIPS collection of protein interactions (yellow diamond) on a\nprecision-recall plot, compared against other published collections of interactions and microarray\ndata, and to the posterior estimates of the MMB models\u2014computed as described in the text.\n\n7\n\nMMB (K=50; MIPS de-noised with Zs & B)MMB (K=50; MIPS summarized with \u03a0s & B)Recall (unnormalized)Precision\fto the model that reveals functionally relevant interactions according to SGD. Figure 3 shows the\nfunctional content of the original MIPS collection of physical interactions (point no.2), and of the\ncollections of interactions computed using (B, \u03a0s), the light blue (\u2212\u00d7) line, and using (B, Zs),\nthe dark blue (\u2212+) line, thresholded at ten different levels\u2014precision-recall curves. The posterior\nmeans of \u03a0s provide a parsimonious representation for the MIPS collection, and lead to precise\nprotein interaction estimates, in moderate amount (\u2212\u00d7 line). The posterior means of Zs provide a\nricher representation for the data, and describe most of the functional content of the MIPS collection\nwith high precision (\u2212+ line). Importantly, the estimated networks corresponding to lower levels\nof recall for both model variants (i.e., \u00d7 and +) feature a more precise functional content than the\noriginal network. This means that the proposed latent block structure is helpful in effectively de-\nnoising the collection of interactions\u2014by ranking them properly. On closer inspection, dense blocks\nof predicted interactions contain known functional predictions that were not in the MIPS collection,\nthus effectively improving the quality of the protein binding data that instantiate cellular activity\nof speci\ufb01c biological contexts, such as biopolymer catabolism and homeostasis. In conclusion, our\nresults suggest that MMB successfully reduces the dimensionality of the data, while discovering in-\nformation about the multiple functionality of proteins that can be used to inform follow-up analyses.\n\nRemarks. A. In the relational setting, cross-validation is feasible if the blockmodel estimated\non training data can be expected to hold on test data; for this to happen the network must be of\nreasonable size, so that we can expect members of each block to be in both training and test sets.\nIn this setting, scheduling of variational updates is important; nested variational scheduling leads to\nef\ufb01cient and parallelizable inference. B. MMB includes two sources of variability, B, \u03a0s, that are\napparently in competition for explaining the data, possibly raising an identi\ufb01ability issue. This is\nnot the case, however, as the blockmodel B captures global/asymmetric relations, while the mixed\nmembership vectors \u03a0s capture local/symmetric relations. This difference practically eliminates the\nissue, unless there is no signal in the data to begin with. C. MMB generalizes to two important cases.\nFirst, multiple data collections Y1:M on the same objects can be generated by the same latent vectors.\nThis might be useful, for instance, for analyzing multivariate sociometric relations simultaneously.\nSecond, in the MMSB the data generating distribution is a Bernoulli, but B can be a matrix of\nparameterizes for any kind of distribution. For instance, technologies for measuring interactions\nbetween pairs of proteins, such as mass spectrometry and tandem af\ufb01nity puri\ufb01cation, which return\na probabilistic assessment about the presence of interactions, thus setting the range of Y \u2208 [0, 1].\n\nReferences\n[1] D. M. Blei, A. Ng, and M. I. Jordan. Latent Dirichlet allocation. Journal of Machine Learning Research,\n\n3:993\u20131022, 2003.\n\n[2] P. Doreian, V. Batagelj, and A. Ferligoj. Discussion of \u201cModel-based clustering for social networks\u201d.\n\nJournal of the Royal Statistical Society, Series A, 170, 2007.\n\n[3] M. S. Handcock, A. E. Raftery, and J. M. Tantrum. Model-based clustering for social networks. Journal\n\nof the Royal Statistical Society, Series A, 170:1\u201322, 2007.\n\n[4] P. D. Hoff, A. E. Raftery, and M. S. Handcock. Latent space approaches to social network analysis.\n\nJournal of the American Statistical Association, 97:1090\u20131098, 2002.\n\n[5] M. Jordan, Z. Ghahramani, T. Jaakkola, and L. Saul. Introduction to variational methods for graphical\n\nmodels. Machine Learning, 37:183\u2013233, 1999.\n\n[6] C. Kemp, J. B. Tenenbaum, T. L. Grif\ufb01ths, T. Yamada, and N. Ueda. Learning systems of concepts with\n\nan in\ufb01nite relational model. In Proc. of the 21st National Conference on Arti\ufb01cial Intelligence, 2006.\n\n[7] F.-F. Li and P. Perona. A Bayesian hierarchical model for learning natural scene categories. IEEE Com-\n\nputer Vision and Pattern Recognition, 2005.\n\n[8] K. Nowicki and T. A. B. Snijders. Estimation and prediction for stochastic blockstructures. Journal of\n\nthe American Statistical Association, 96:1077\u20131087, 2001.\n\n[9] J. K. Pritchard, M. Stephens, N. A. Rosenberg, and P. Donnelly. Association mapping in structured\n\npopulations. American Journal of Human Genetics, 67:170\u2013181, 2000.\n\n[10] F. S. Sampson. A Novitiate in a period of change: An experimental and case study of social relationships.\n\nPhD thesis, Cornell University, 1968.\n\n[11] S. Wasserman, G. Robins, and D. Steinley. A brief review of some recent research. In: Statistical Network\nAnalysis: Models, Issues and New Directions, Lecture Notes in Computer Science. Springer-Verlag, 2007.\n\n8\n\n\f", "award": [], "sourceid": 90, "authors": [{"given_name": "Edo", "family_name": "Airoldi", "institution": null}, {"given_name": "David", "family_name": "Blei", "institution": null}, {"given_name": "Stephen", "family_name": "Fienberg", "institution": null}, {"given_name": "Eric", "family_name": "Xing", "institution": null}]}