{"title": "Inference and communication in the game of Password", "book": "Advances in Neural Information Processing Systems", "page_first": 2514, "page_last": 2522, "abstract": "Communication between a speaker and hearer will be most efficient when both parties make accurate inferences about the other. We study inference and communication in a television game called Password, where speakers must convey secret words to hearers by providing one-word clues. Our working hypothesis is that human communication is relatively efficient, and we use game show data to examine three predictions. First, we predict that speakers and hearers are both considerate, and that both take the other\u2019s perspective into account. Second, we predict that speakers and hearers are calibrated, and that both make accurate assumptions about the strategy used by the other. Finally, we predict that speakers and hearers are collaborative, and that they tend to share the cognitive burden of communication equally. We find evidence in support of all three predictions, and demonstrate in addition that efficient communication tends to break down when speakers and hearers are placed under time pressure.", "full_text": "Inference and communication in the game of\n\nPassword\n\nYang Xu\u2217 and Charles Kemp\u2020\nMachine Learning Department\u2217\nSchool of Computer Science\u2217\nDepartment of Psychology\u2020\nCarnegie Mellon University\n\n{yx1@cs.cmu.edu, ckemp@cmu.edu}\n\nAbstract\n\nCommunication between a speaker and hearer will be most ef\ufb01cient when both\nparties make accurate inferences about the other. We study inference and com-\nmunication in a television game called Password, where speakers must convey\nsecret words to hearers by providing one-word clues. Our working hypothesis is\nthat human communication is relatively ef\ufb01cient, and we use game show data to\nexamine three predictions. First, we predict that speakers and hearers are both\nconsiderate, and that both take the other\u2019s perspective into account. Second, we\npredict that speakers and hearers are calibrated, and that both make accurate as-\nsumptions about the strategy used by the other. Finally, we predict that speakers\nand hearers are collaborative, and that they tend to share the cognitive burden of\ncommunication equally. We \ufb01nd evidence in support of all three predictions, and\ndemonstrate in addition that ef\ufb01cient communication tends to break down when\nspeakers and hearers are placed under time pressure.\n\n1 Introduction\n\nCommunication and inference are intimately linked. Suppose, for example, that Joan states that\nsome of her pets are dogs. Under normal circumstances, a hearer will infer that not all of Joan\u2019s\npets are dogs on the grounds that Joan would have expressed herself differently if all of her pets\nwere dogs [1]. Inferences like these have been widely studied by linguists and psychologists [2,\n3, 4, 5] and are often encountered in everyday settings. One compelling explanation is presented\nby Levinson [4], who points out that speaking (i.e. phonetic articulation) is substantially slower\nthan thinking (i.e. inference). As a result, communication will be maximally ef\ufb01cient if a speaker\u2019s\nutterance leaves inferential gaps that will be bridged by the hearer. Inference, however, is not only\nthe responsibility of the hearer. For communication to be maximally ef\ufb01cient, a speaker must take\nthe hearer\u2019s perspective into account (\u201cif I say X, will she infer Y?\u201d). The hearer should therefore\nallow for inferences on the part of the speaker (\u201cdid she think that saying X would lead me to infer\nY?\u201d) Considerations of this sort rapidly lead to a game-theoretic regress, and achieving ef\ufb01cient\ncommunication under these circumstances begins to look like a very challenging problem.\n\nHere we study a simple communication game that allows us to explore inferences made by speakers\nand hearers. Inference becomes especially important in settings where speakers are prevented from\ndirectly expressing the concepts they have in mind, and where utterances are constrained to be short.\nThe television show Password is organized around a game that satis\ufb01es both constraints. In this\ngame, a speaker is supplied with a single, secret word (the password) and must communicate this\nword to a hearer by choosing a single one-word clue. For example, if the password is \u201cmend\u201d, then\nthe speaker might choose \u201csew\u201d as the clue, and the hearer might guess \u201cstitch\u201d in response. Figure 1\nshows several examples drawn from the show\u2014note that communication is successful in the \ufb01rst\n\n1\n\n\fpassword:divide; clue:multiply\n\npassword:mend; clue:sew\n0\n\npassword:shovel; clue:snow\n\n0\n\nseparate\n\nsplit\n\nsubtract\n\nnumbers\n\nconquer\n\npart\n\n\u22122\n\nstitch\n\nseam\n\nheal\n\nfix\n\nb\n\nS\n\n\u22124\n\n\u22126\n\n\u22128\n\u22128\n\n0\n\nrestore\n\nbreak\n\nbend\n\nrip\n\n\u22124\nS\nf\n\ndig\n\ndigger\n\nspade\n\nscoop\nditch\n\ntool\n\npick\n\nspoon\n\ndirt\n\nwork\n\n\u22122\n\nb\n\nS\n\n\u22124\n\n\u22126\n\n\u22128\n\u22128\n\n\u22124\nS\nf\n\n0\n\nb\n\n)\n\nS\n\n(\n \n\nh\n\nt\n\ng\nn\ne\nr\nt\ns\n \n\nd\nr\na\nw\nk\nc\na\nb\n\n \n\ng\no\n\nl\n\n\u22122\n\n\u22124\n\n\u22126\n\n\u22128\n\u22128\n\nb\n\n)\n\nH\n\n(\n \n\nh\n\nt\n\ng\nn\ne\nr\nt\ns\n \n\nd\nr\na\nw\nk\nc\na\nb\n\n \n\ng\no\n\nl\n\n\u22124\n\n\u22126\n\n\u22128\n\u22128\n\nquotient\n\u22122\n\n\u22126\n\n\u22124\n\nlog forward strength (S\n)\nf\n\nclue:multiply; guess:divide\n0\n\npwd:divide\n\n\u22122\n\ndivision\n\ntimes\n\nadd\n\nfactor\nsubtract\n\nreproduce\n\nb\n\nH\n\n\u22124\n\n\u22126\n\n\u22122\n\n0\n\n\u22126\n\n\u22122\n\n0\n\n0\n\n\u22122\n\nclue:sew; guess:stitch\n\nclue:snow; guess:flake\n\nneedle\n\nthread\n\n0\n\n\u22122\n\npwd:\nshovel\n\nhail\n\npwd:mend\n\nb\n\nH\n\n\u22124\n\nski\n\nwhite\n\ncold\n\nyarn\npin\n\nfabric\n\nrabbit\n\nmath\n\n\u22126\n\n\u22124\n\n\u22122\n\nlog forward strength (H\n)\nf\n\n\u22126\n\n\u22128\n\u22128\n\n0\n\npants\n\nclothes\n\n\u22126\n\n\u22124\nH\nf\n\n\u22122\n\n0\n\n\u22126\n\n\u22128\n\u22128\n\nball\n\n\u22126\n\n\u22124\nH\nf\n\nfall\n\n\u22122\n\n0\n\nFigure 1: Three rounds from the television game show Password. Given each password, the top row\nplots the forward (Sf : password \u2192 clue) and backward (Sb: password \u2190 clue) strengths for several\npotential clues. The clue chosen by the speaker is circled. Given this clue, the bottom row plots\nthe forward (Hf : clue \u2192 guess) and backward (Hb: clue \u2190 guess) strengths for several potential\nguesses. The guess chosen by the hearer is circled and the password is indicated by an arrow. The\n\ufb01rst two columns represent two normal rounds, and the \ufb01nal column is a lightning round where\nspeakers and hearers are placed under time pressure. The gray dots in each plot show words that are\nassociated with the password (top row) or clue (bottom row) in the University of Southern Florida\nword association database. Labels for these words are included where space permits.\n\nexample but not in the remaining two. The clues and guesses generated by speakers and hearers\nare obviously much simpler than most real-world linguistic utterances, but studying a setting this\nsimple allows us to develop and evaluate formal models of communication. Our analyses therefore\ncontribute to a growing body of work that uses formal methods to explore the ef\ufb01ciency of human\ncommunication [6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16].\n\nAt \ufb01rst sight the optimal strategies for speaker and hearer may seem obvious: the speaker should\ngenerate the clue that is associated most strongly with the password, and the hearer should guess\nthe word that is associated most strongly with the clue. Note, however, that word associations are\nasymmetric. Given a pair of words such as \u201cshovel\u201d and \u201csnow\u201d, the forward association (shovel \u2192\nsnow) may be strong but the backward association (shovel \u2190 snow) may be weak. The third example\nin Figure 1 shows a case where communication fails because the speaker chooses a clue with a\nstrong forward association but a weak backward association. Although the data include examples\nlike the case just described, we hypothesize that speakers and hearers are both considerate: in other\nwords, that both parties attempt to take the other\u2019s perspective into account. We test this hypothesis\nby exploring whether speakers and hearers tend to take backward associations into account when\ngenerating their clues and guesses.\n\nOur second hypothesis is that speaker and hearer are calibrated: in other words, that both make\naccurate assumptions about the strategy used by the other. Taking the other person\u2019s perspective into\naccount is a good start, but is no guarantee of calibration. Suppose, for example, that the speaker\nattempts to make the hearer\u2019s task as easy as possible, and considers only backward associations\nwhen choosing his clue. This strategy will work best if the hearer considers only forward associates\nof the clue, but suppose that the hearer considers only backward associations, on the theory that the\nspeaker probably generated his clue by choosing a forward associate. In this case, both parties are\nconsiderate but not calibrated, and communication is unlikely to prove successful.\n\n2\n\n\fOur third hypothesis is that speakers and hearers are collaborative: in other words, that they settle on\nstrategies that tend to share the cognitive burden of communication. In operationalizing this hypoth-\nesis we assume that forward associates are easier for people to generate than backward associates.\nA pair of strategies can be calibrated but not cooperative: for example, the speaker and hearer will\nbe calibrated if both agree that the speaker will consider only forward associates, and the hearer will\nconsider only backward associates. This policy, however, is likely to demand more effort from the\nhearer than the speaker, and we propose that speakers and hearers will satisfy the principle of least\ncollaborative effort [17, 18] by choosing a calibrated pair of strategies where each person weights\nforward and backward associates equally.\n\nTo evaluate our hypotheses we use word association data to analyze the choices made by game\nshow contestants. We \ufb01rst present evidence that speakers and hearers are considerate and take both\nforward and backward associations into account. We then develop simple models of the speaker and\nhearer, and use these models to explore the extent to which speakers and hearers weight forward\nand backward associations. Our results suggest that speakers and hearers are both calibrated and\ncollaborative under normal conditions, but that calibration and collaboration tend to break down\nunder time pressure.\n\n2 Game show and word association data\n\nWe collected data from the Password game show hosted by Allen Ludden on CBS. Previous re-\nsearchers have used game show data to explore several aspects of human decision-making [19], but\nto our knowledge the game of Password has not been previously studied. In each game round, a\nsingle English word (the password) is shown to speakers on two competing teams. With each team\ntaking turns, the speaker gives a one-word clue to the hearer and the hearer makes a one-word guess\nin return. The team that performs best proceeds to the lightning rounds where the same game is\nplayed under time pressure. Our data set includes passwords, speaker-generated clues and hearer-\ngenerated guesses for 100 normal and 100 lightning rounds sampled from the show episodes during\n1962\u20131967. Each round includes a single password and potentially multiple clues and guesses from\nboth teams. For all our our analyses, we use only the \ufb01rst clue\u2013guess pair in each round.\n\nThe responses of speakers and hearers are likely to depend heavily on word associations, and we\ncan therefore use word association data to model both speakers and hearers. We used the word\nassociation database from the University of South Florida (USF) for all of our analyses [20]. These\ndata were collected using a free association task, where participants were given a cue word and asked\nto generate a single associate of the cue. More than 6000 participants contributed to the database,\nand each generated associates for 100\u2013120 English words. To allow for weak associates that were\nnot generated by these participants, we added a count of 1 to the observed frequency for each cue-\ntarget pair in the database. The forward strength (wi \u2192 wj) is de\ufb01ned as the proportion of wi trials\nwhere wj was generated as an associate. The backward strength (wi \u2190 wj) is proportional to the\nforward strength (wj \u2192 wi) but is normalized with respect to all forward strengths to wi:\n\n(wi \u2190 wj ) =\n\n(wj \u2192 wi)\n\nPk(wk \u2192 wi)\n\n.\n\n(1)\n\nNote that this normalization ensures that both forward and backward strengths can be treated as\nprobabilities. The correlation between forward strengths and backward strengths is positive but low\n(r = 0.32), suggesting that our game show analyses may be able to differentiate the in\ufb02uence of\nforward and backward associations.\n\nThe USF database includes associates for a set of 5016 words, and we used this set as the lexicon for\nall of our analyses. Some of the rounds in our game show data include passwords, clues or guesses\nthat do not appear in this lexicon, and we removed these rounds, leaving 68 password-clue and 68\nclue-guess pairs in the normal rounds and 86 password-clue pairs and 80 clue-guess pairs in the\nlightning rounds. The USF database also includes the frequency of each word in a standard corpus\nof written English [21], and we use these frequencies in our \ufb01rst analysis.\n\n3\n\n\fa) i)\n\nk\nn\na\nr\n \nd\ne\nz\n\ni\nl\n\na\nm\nr\no\nn\n \ng\no\n\nl\n\n\u22126\n\n\u22128\n\n\u221210\n\n\u221212\n\n \n\nSN\n\n \n\nmean\n\n\u22126\n\n\u22128\n\n\u221210\n\n\u221212\n\nHN\n\nb) i)\n\nSL\n\n\u22126\n\n\u22128\n\n\u221210\n\n\u221212\n\n\u22126\n\n\u22128\n\n\u221210\n\n\u221212\n\nS\nb\n\nSN\n\nS\nf\n\nii)\nr = 0.32\n\n + S\nS\nf\nb\n\nH\nf\n\nH\nb\n\nHN\n\n + H\nH\nf\nb\n\nS\nf\n\nii)\n\nS\nb\n\nSL\n\n + S\nS\nf\nb\n\nH\nf\n\nHL\n\nH\nb\n\nHL\n\n + H\nH\nf\nb\n\nk\nn\na\nr\n \nd\nr\na\nw\nk\nc\na\nb\n \ng\no\n\nl\n\n5\n\n4\n\n3\n\n2\n\n1\n\n0\n\nt\n\nn\nu\no\nc\n \n\nd\ne\nz\n\ni\nl\n\na\nm\nr\no\nn\n\n0.5\n\n0.4\n\n0.3\n\n0.2\n\n0.1\n\n0\n\n0\n\niii)\n\n2\n\n1\n4\nlog forward rank\n\n3\n\nSN\n\nBb EbWb Cf\n\nBf Ef Wf Cb\n\n5\n\n5\n\n4\n\n3\n\n2\n\n1\n\n0\n\n0.5\n\n0.4\n\n0.3\n\n0.2\n\n0.1\n\n0\n\nr = 0.70\n\n0\n\n1\n\n2\n\n3\n\n4\n\n5\n\n5\n\n4\n\n3\n\n2\n\n1\n\n0\n\nr = 0.60\n\n0\n\n1\n\n2\n\n3\n\n4\n\n5\n\nHN\n\niii)\n\n0.5\n\nSL\n\n0.4\n\n0.3\n\n0.2\n\n0.1\n\n0\n\nBb EbWb Cf\n\nBf Ef Wf Cb\n\nBb EbWb Cf\n\nBf Ef Wf Cb\n\n5\n\n4\n\n3\n\n2\n\n1\n\n0\n\n0.5\n\n0.4\n\n0.3\n\n0.2\n\n0.1\n\n0\n\nr = 0.45\n\n0\n\n1\n\n2\n\n3\n\n4\n\n5\n\nHL\n\nBb EbWb Cf\n\nBf Ef Wf Cb\n\nFigure 2: (a) Analyses of the speaker and hearer data (SN and HN ) from the normal rounds. (i)\nRanks of the human responses normalized with respect to all other words in the lexicon. Ranks are\nshown along three dimensions: forward strength (f ), backward strength (b) and combined forward\nand backward strengths. The dark square shows the mean rank, and the horizontal lines within the\nbox show the median and interquartile range. The plus symbols are outliers. (ii) Ranks of the human\nresponses along the forward and backward dimensions. (iii) \u201cMatched rank\u201d analysis exploring\nwhether human responses tend to be better along one of the dimensions than alternatives that are\nmatched along the other dimension. The four bars on the left in each subplot show normalized\ncounts based on comparisons with matches along the f dimension, and the four bars on the right are\nbased on matches along the b dimension. For example, group Bb includes human responses that are\nbetter along the b dimension compared to matches along the f dimension, and groups Eb and W b\ninclude cases where human responses are equal to or worse than the f -matches. Group Cf includes\ncases where the human response is top ranked along the f dimension. Groups Bf , Ef , W f and Cb\nare de\ufb01ned similarly. (b) Analyses of the lightning rounds.\n\n3 Speakers and hearers are considerate\n\nA speaker should \ufb01nd it easy to generate clues that are strong forward associates of a password, and a\nhearer should likewise \ufb01nd it easy to generate guesses that are strong forward associates of a clue. A\nconsiderate speaker, however, may attempt to generate strong backward associates, which will make\nit easier for the hearer to successfully guess the password. Similarly, a hearer who considers the task\nfaced by the speaker should also take backward associates into account. This section describes some\ninitial analyses that explore whether clues and guesses are shaped by backward associations.\n\nFigure 2a.i compares forward and backward strengths as predictors of the responses chosen by\nspeakers and hearers. A dimension is a successful predictor if the words chosen by contestants tend\nto have low ranks along this dimension with respect to the 5016 words in the lexicon (rank 1 is the top\nrank). We handle ties using fractional ranking, which means that it is sensible to compare mean ranks\nalong each dimension. In Figure 2a.i, Sf and Sb represent forward (password \u2192 clue) and backward\n(password \u2190 clue) strengths for the speaker, and Hf and Hb represent forward (clue \u2192 guess) and\n\n4\n\n\fbackward (clue \u2190 guess) strengths for the hearer. In addition to forward and backward strengths,\nwe also considered word frequency as a predictor. Across both normal (SN and HN ) and lightning\n(SL and HL) rounds, the ranks along the forward and backward dimensions are substantially better\nthan ranks along the frequency dimension (p < 0.01 in pairwise t-tests), and we therefore focus on\nforward and backward strengths for the rest of our analyses.\nFor data set SN the mean ranks suggest that forward and backward strengths appear to predict\nchoices about equally well. The third dimension Sf + Sb is created by combining dimensions Sf\nand Sb. Word w1 dominates w2 if it is superior along one dimension and no worse along the other,\nand the rank for each word along the combined dimension is based on the number of words that\ndominate it. For data set SN , the mean rank based on the Sf + Sb dimension is lower than that for\nSf alone, suggesting that backward strengths make a predictive contribution that goes beyond the\ninformation present in the forward associations. Note, however, that the difference between mean\nranks for Sf and Sf + Sb is not statistically signi\ufb01cant.\nFor data set HN , Figure 2a.i provides little evidence that backward strengths make a contribution\nthat goes beyond the forward strengths. Figure 2a.ii plots the rank of each guess along the dimen-\nsions of forward and backward strength. The correlation between the dimensions is relatively high,\nsuggesting that both dimensions tend to capture the information present in the other. As a result, the\nhearer data set HN may offer little opportunity to explore whether backward and forward associa-\ntions both contribute to people\u2019s responses.\n\nFigure 2a.iii shows the results of an analysis that explores more directly whether each dimension\nmakes a contribution that goes beyond the other. We compared each \u201cactual word\u201d (i.e. each clue\nor guess chosen by a contestant) to \u201cmatched words\u201d that are matched in rank along one of the\ndimensions. For example, if the backward dimension matters, then the actual words should tend to\nbe better along the b dimension than words that are matched along the f dimension. The \ufb01rst group\nof bars in Figure 2a.iii shows the proportion of actual words that are better (Bb), equivalent (Eb) or\nworse (W b) along the backward dimension than matches along the forward dimension. The Bb bar\nis higher than the others, suggesting that the backward dimension does indeed make a contribution\nthat goes beyond the forward dimension. Note that a match is de\ufb01ned as a word that is ranked\nthe same as the actual word, or in cases where there are no ties, a word that is ranked one step\nbetter. The fourth bar (Cf , for champion along the forward dimension) includes all cases where a\nword is ranked best along the forward dimension, which means that no match can be found. Our\npolicy for identifying matches is conservative\u2014all other things being equal, actual words should be\nequivalent (Eb) or worse (W b) than the matched words, which means that the large Bb bar provides\nstrong evidence that the backward dimension is important. A binomial test con\ufb01rms that the Bb\nbar is signi\ufb01cantly greater than the W b bar (p < 0.05). The Bf bar for the speaker data is also\nhigh, suggesting that the forward dimension makes a contribution that goes beyond the backward\ndimension. In other words, Figure 2a.iii suggests that both dimensions in\ufb02uence the responses of\nthe speaker.\nThe results for the hearer data HN provide additional support for the idea that neither dimension\npredicts hearer guesses better than the other. Note, for example, that the second group of four bars\nin Figure 2a.iii suggests that the forward dimension is not predictive once the backward dimension\nis taken into account (Bf is smaller than W f ). This result is consistent with our previous \ufb01nding\nthat forward and backward strengths are highly correlated in the case of the hearer, and that neither\ndimension makes a contribution after controlling for the other.\n\nOur analyses so far suggest that forward and backward strengths both make independent contri-\nbutions to the choices made by speakers, but that the hearer data do not allow us to discriminate\nbetween these dimensions. Figure 2b shows similar analyses for the lightning rounds. The most no-\ntable change is that backward strengths appear to play a much smaller role when speakers are placed\nunder time pressure. For example, Figure 2b.i suggests that backward strengths are now worse than\nforward strengths at predicting the clues chosen by speakers. Relative to the results for the normal\nrounds SN , the Bb counts for SL in Figure 2b.iii show a substantial drop (53% decrease) and the\nBf counts show an increase of similar scale. \u03c72 goodness-of-\ufb01t tests show that the distributions\nof counts for both {Bb, Eb, W b, Cf } and {Bf, Ef, W f, Cb} in the lightning rounds signi\ufb01cantly\ndeviate from those in the normal rounds (p < 0.01). This result provides further evidence that\nspeakers tend to rely more heavily on forward associations than backward associations when placed\nunder time pressure.\n\n5\n\n\fSpeaker distribution pS(c|w)\n\nHearer distribution pH (w|c)\n\nS0\nS1\nS2\n\n(w \u2192 c)\n(w \u2190 c)\n\u03b1(2)\nS (w \u2192 c) + \u03b2(2)\n\nS (w \u2190 c)\n\n...\n\nH0\nH1\nH2\n\n(c \u2192 w)\n(c \u2190 w)\n\u03b1(2)\nH (c \u2192 w) + \u03b2(2)\n\nH (c \u2190 w)\n\n...\n\nSn \u03b1(n)\n\nS (w \u2192 c) + \u03b2(n)\n\nS (w \u2190 c)\n\nHn \u03b1(n)\n\nH (c \u2192 w) + \u03b2(n)\n\nH (c \u2190 w)\n\nTable 1: Strategies for speaker and hearer. In each case we assume that the speaker and hearer\nsample words from distributions pS(c|w) and pH (w|c) based on the expressions shown. At level 0,\nboth speaker and hearer rely entirely on forward associates, and at level 1, both parties rely entirely\non backward associates. For each party, the strategy at level k is the best choice assuming that the\nother person uses a strategy at a level lower than k.\n\nOur previous analyses found little evidence that forward and backward strengths make separate\ncontributions in the case of the hearer, but the lightning data HL suggest that these dimensions\nmay indeed make separate contributions. Figure 2b.iii suggests that time pressure affects these\ndimensions differently: note that Bb counts decrease by 19% and Bf counts increase by 64%. \u03c72\ntests con\ufb01rm that the distributions of {Bb, Eb, W b, Cf} and {Bf, Ef, W f, Cb} in the lightning\nrounds signi\ufb01cantly deviate from those in the normal rounds (p < 0.01), suggesting that the hearer\n(like the speaker) tends to rely on forward strengths rather than backward strengths in the lightning\nrounds.\n\nTaken together, the full set of results in Figure 2 suggests that the responses of speakers and hearers\nare both shaped by backward associates\u2014in other words, that both parties are considerate of the\nother person\u2019s situation. The evidence in the case of the speaker is relatively strong and all of the\nanalyses we considered suggest that backward associations play a role. The evidence is weaker in\nthe case of the hearer, and only the comparison between normal and lightning rounds suggests that\nbackward associations play some role.\n\n4 Ef\ufb01cient communication: calibration and collaboration\n\nOur analyses so far provide some initial evidence that speakers and hearers are both in\ufb02uenced by\nforward and backward associations. Given this result, we now consider a model that explores how\nforward and backward associations are combined in generating a response.\n\n4.1 Speaker and hearer models\n\nSince both kinds of associations appear to play a role, we explore a simple speaker model which\nassumes that the clue c chosen for the password w is sampled from a mixture distribution\n\npS(c|w) = \u03b1S(w \u2192 c) + \u03b2S(w \u2190 c)\n\n(2)\n\nwhere (w \u2192 c) indicates the forward strength from w to c, (w \u2190 c) indicates the backward strength\nfrom c to w, and \u03b1S and \u03b2S are mixture weights that sum to 1. The corresponding hearer model\nassumes that guess w given clue c is sampled from the mixture distribution\n\npH (w|c) = \u03b1H (c \u2192 w) + \u03b2H (c \u2190 w).\n\n(3)\n\nSeveral possible mixture distributions for speaker and hearer are shown in Table 1. For example,\nthe level 0 distributions assume that speaker and hearer both rely entirely on forward associates, and\nthe level 1 distributions assume that both rely entirely on backward associates. By \ufb01tting mixture\nweights to the game show data we can explore the extent to which speaker and hearer rely on forward\nand backward associations.\n\nThe mixture models in Equations 2 and 3 can be derived by assuming that the hearer relies on\nBayesian inference. Using Bayes\u2019 rule, the hearer distribution pH (w|c) can be expressed as\n\npH(w|c) \u221d pS(c|w)p(w).\n\n(4)\n\n6\n\n\fTo simplify our analysis we make three assumptions. First, we assume that the prior p(w) in Equa-\ntion 4 is uniform. Second, we assume that contestants are near-optimal in many respects but that\nthey sample rather than maximize. In other words, we assume that the hearer samples a guess w\nfrom the distribution pH (w|c) in Equation 4, and that the speaker samples a clue from a distribution\npS(c|w) \u221d pH (w|c). Finally, we assume that the normalizing constant in Equation 1 is 1 for all\nwords wi. This assumption seems reasonable since for our smoothed data set the mean value of the\nnormalizing constant is 1 and the standard deviation is 0.04. Our \ufb01nal assumption simpli\ufb01es matters\nconsiderably since it implies that (wi \u2192 wj ) = (wj \u2190 wi) for all pairs wi and wj.\nGiven these assumptions it is straightforward to show that the level 0 strategies in Table 1 are the\nbest responses to the level 1 strategies, and vice versa. For example, if the speaker uses strategy S0\nand samples a clue c from the distribution pS(c|w) = w \u2192 c, then Equation 4 suggests that the\nhearer should sample a guess w from the distribution pH (c|w) \u221d (w \u2192 c) = (c \u2190 w). Similarly,\nif the speaker uses the strategy S1 and samples a clue c from the distribution pS(c|w) = (w \u2190 c),\nthen Equation 4 suggests that the hearer should sample a guess w from the distribution pH (c|w) \u221d\n(w \u2190 c) = (c \u2192 w).\nSuppose now that the hearer is uncertain about the strategy used by the speaker. A level 2 hearer\nassumes that the speaker could use strategy S0 or strategy S1 and assigns prior probabilities of \u03b2(2)\nand \u03b1(2)\nH to these speaker strategies. Since H1 is the appropriate response to S0 and H0 is the\nappropriate response to S1, the level 2 hearer should sample from the distribution\n\nH\n\npH (w|c) = p(S1)pH (w|c, S1) + p(S0)pH (w|c, S0)\n\n= \u03b1(2)\n\nH (c \u2192 w) + \u03b2(2)\n\n(5)\nMore generally, suppose that a level n hearer assumes that the speaker uses a strategy from the\nset {S0, S1, . . . , Sn\u22121}. Since the appropriate response to any one of these strategies is a mixture\nsimilar to Equation 5, it follows that strategy Hn is also a mixture of the distributions (w \u2192 c) and\n(w \u2190 c). A similar result holds for the speaker, and strategy Sn in Table 1 also takes the form of\na mixture distribution. Our Bayesian analysis therefore suggests that ef\ufb01cient speakers and hearers\ncan be characterized by the mixture models in Equations 2 and 3.\n\nH (c \u2190 w).\n\nSome pairs of mixture models are calibrated in the sense that the hearer model is the best choice\ngiven the speaker model and vice versa. Equation 4 implies that calibration is achieved when the\nforward weight for the speaker matches the backward weight for the hearer (\u03b1S = \u03b2H) and the\nbackward weight for the speaker matches the forward weight for the hearer (\u03b2S = \u03b1H). If game\nshow contestants achieve ef\ufb01cient communication, then mixture weights \ufb01t to their responses should\ncome close to satisfying this calibration condition.\n\nThere are many sets of weights that satisfy the calibration condition. For example, calibration is\nachieved if the speaker uses strategy S0 and the hearer uses strategy H1. If generating backward\nassociates is more dif\ufb01cult than thinking about forward associates, this solution seems unbalanced\nsince the hearer alone is required to think about backward associates. Consistent with the principle\nof least collaborative effort, we make a second prediction that speaker and hearer will collaborate\nand share the communicative burden equally. More precisely, we predict that both parties will assign\nthe same weight to backward associates and that \u03b2S will equal \u03b2H. Combining our two predictions,\nwe expect that the weights which best characterize human responses will have \u03b1S = \u03b2S = \u03b1H =\n\u03b2H = 0.5.\n\n4.2 Fitting forward and backward mixture weights to the data\n\nTo evaluate our predictions we assumed that the speaker and hearer are characterized by Equations 2\nand 3 and identi\ufb01ed the mixture weights that best \ufb01t the game show data. Assuming that the M game\nrounds are independent, the log likelihood for the speaker data is\n\nL = log\n\nM\n\nY\n\nm=1\n\nP (cm|wm) =\n\nM\n\nX\n\n[\u03b1S log(wm \u2192 cm) + \u03b2S log(wm \u2190 cm)]\n\n(6)\n\nm=1\n\nand a similar expression is used for the hearer data. We \ufb01t the weights \u03b1S and \u03b2S by maximizing the\nlog likelihood in Equation 6. Since this likelihood term is convex and there is a single free parameter\n(\u03b1S +\u03b2S = 1), the global optimum can be found by a simple line search over the range 0 < \u03b1S < 1.\n\n7\n\n\fa)\n\n1\n\ns\nt\n\ni\n\n \n\nh\ng\ne\nw\ne\nr\nu\nt\nx\nm\n\ni\n\n0.8\n\n0.6\n\n0.4\n\n0.2\n\n0\n\n \n\nb)\n\n \n\n\u03b1\n\u03b2\n\n)\n\u03b1 \u03b2\n(\ng\no\nl\n\n4\n\n3\n\n2\n\n1\n\n0\n\n\u22121\n\n\u22122\n\nc)\n\n)\nc\ne\ns\n(\n \ne\nm\n\ni\nt\n \n\ne\ns\nn\no\np\ns\ne\nr\n\n10\n\n5\n\n0\n\n \n\n \n\nnormal\nlightning\n\nS\n\nH\n\nS\nN\n\nH\nN\n\nS\nL\n\nH\nL\n\nS\nN\n\nH\nN\n\nS\nL\n\nH\nL\n\nFigure 3: (a) Fitted mixture weights for the speaker (S) and hearer (H) models based on boot-\nstrapped normal (N) and lightning (L) rounds. \u03b1 and \u03b2 are weights on the forward and backward\nstrengths. (b) Log-ratios of \u03b1 and \u03b2 weights estimated from bootstrapped normal and lightning\nrounds. (c) Average response times for speakers choosing clues and hearers choosing guesses in\nnormal and lightning rounds. Averages are computed over 30 rounds randomly sampled from the\ngame show.\n\nWe ran separate analyses for normal and lightning rounds, and ran similar analyses for the hearer\ndata. 1000 estimates of each mixture weight were computed by bootstrapping game show rounds\nwhile keeping tallies of normal and lightning rounds constant.\n\nConsistent with our predictions, the results in Figure 3a suggest that all four mixure weights for\nthe normal rounds are relatively close to 0.5. Both speaker and hearer appear to weight forward\nassociates slightly more heavily than backward associates, but 0.5 is within one standard deviation\nof the bootstrapped estimates in all four cases. The lightning rounds produce a different pattern\nof results and suggest that the speaker now relies much more heavily on forward than backward\nassociates. Figure 3b shows log ratios of the mixture weights, and indicates that these ratios lie close\nto 0 (i.e. \u03b1 = \u03b2) in all cases except for the speaker in the lightning rounds. Further con\ufb01dence tests\nshow that the percentage of bootstrapped ratios exceeding 0 is 100% for the speaker in the lightning\nrounds, but 85% or lower in the three remaining cases. Consistent with our previous analyses, this\nresult suggests that coordinating with the hearer requires some effort on the part of the speaker,\nand that this coordination is likely to break down under time pressure. The \ufb01tted mixture weights,\nhowever, do not con\ufb01rm the prediction that time pressure makes it dif\ufb01cult for the hearer to consider\nbackward associations. Figure 3c helps to explain why mixture weights for the speaker but not the\nhearer may differ across normal and lightning rounds. The difference in response times between\nnormal and lightning rounds is substantially greater for the speaker than the hearer, suggesting that\nany differences between normal and lightning rounds are more likely to emerge for the speaker than\nthe hearer.\n5 Conclusion\n\nWe studied how speakers and hearers communicate in a very simple context. Our results suggest\nthat both parties take the other person\u2019s perspective into account, that both parties make accurate\nassumptions about the strategy used by the other, and that the burden of communication is equally\ndivided between the two. All of these conclusions support the idea that human communication\nis relatively ef\ufb01cient. Our results, however, suggest that ef\ufb01cient communication is not trivial to\nachieve, and tends to break down when speakers are placed under time pressure.\n\nAlthough we worked with simple models of the speaker and hearer, note that neither model is in-\ntended to capture psychological processing. Future studies can explore how our models might be\nimplemented by psychologically plausible mechanisms. For example, one possibility is that speak-\ners sample a small set of words with high forward strengths, then choose the word in this sample\nwith greatest backward strength. Different processing models might be considered, but we believe\nthat any successful model of speaker or hearer will need to include some role for inferences about\nthe other person.\nAcknowledgments This work was supported in part by the Richard King Mellon Foundation (YX)\nand by NSF grant CDI-0835797 (CK).\n\n8\n\n\fReferences\n\n[1] L. Horn. Toward a new taxonomy for pragmatic inference: Q-based and R-based implicature.\nIn Meaning, Form, and Use in Context: Linguistic Applications. Georgetown University Press,\n1984.\n\n[2] P. Grice. Studies in the Way of Words. Harvard University Press, Cambridge, 1989.\n[3] D. Sperber. Relevance: Communication and Cognition. Blackwell, Oxford, 1986.\n[4] S. Levinson. Presumptive Meanings: The Theory of Generalized Implicature. MIT Press,\n\nCambridge, 2000.\n\n[5] D. Jurafsky. Pragmatics and computational linguistics. In L. R. Horn and G. Ward, editors,\n\nHandbook of Pragmatics, pages 578\u2013604. Blackwell, Oxford, 2005.\n\n[6] G. K. Zipf, editor. Human behaviour and the principle of least effort: An introduction to human\n\necology. Addison-Wesley Press, Cambridge, 1949.\n\n[7] R. Levy and T. F. Jaeger. Speakers optimize information density through syntactic reduction.\n\nIn Advances in Neural Information Processing Systems, 2007.\n\n[8] T. F. Jaeger. Redundancy and reduction: Speakers manage syntactic information density. Cog-\n\nnitive Psychology, 61(1):23\u201362, 2010.\n\n[9] M. Aylett and A. Turk. The smooth signal redundancy hypothesis: A functional explanation for\nrelationships between redundancy, prosodic prominence, and duration in spontaneous speech.\nLanguage and Speech, 47(1):31\u201356, 2004.\n\n[10] S. T. Piantadosi, H. J. Tily, and E. Gibson. The communicative lexicon hypothesis. In The 31st\n\nannual meeting of the Cognitive Science Society, 2009.\n\n[11] R. Baddeley and D. Attewell. The relationship between language and the environment: in-\nformation theory shows why we have only three lightness terms. Psychological Science,\n20(9):1100\u20131107, 2009.\n\n[12] J. Hawkins. Ef\ufb01ciency and complexity in grammars. Oxford University Press, Oxford, 2004.\n[13] N. Chomsky. Language and mind: current thoughts on ancient problems. In L. Jenkins, editor,\n\nVariations and universals in biolinguistics, pages 379\u2013405. Elsevier, Amsterdam, 2004.\n\n[14] R. van Rooy. Conversational implicatures and communication theory. In J. van Kuppevelt and\n\nR. Smith, editors, Current and New Directions in Discourse and Dialogue. Kluwer, 2003.\n\n[15] C. R. M. McKenzie and J. D. Nelson. What a speaker\u2019s choice of frame reveals: reference\n\npoints, frame selection, and framing effects. Psychonomic Bullentin and Review, 10, 2003.\n\n[16] S. Sher and C. R. M. McKenzie. Information leakage from logically equivalent frames. Cog-\n\nnition, 101:467\u2013494, 2006.\n\n[17] H. H. Clark and D. Wilkes-Gibbs. Referring as a collaborative process. Cognition, 22:1\u201339,\n\n1986.\n\n[18] H. H. Clark. Using language. Cambridge University Press, Cambridge, 1996.\n[19] J. B. Berk, E. Hughson, and K. Vandezande. The price is right, but are the bids? An investiga-\n\ntion of rational decision theory. The American Economic Review, 86(4):654\u2013970, 1996.\n\n[20] D. L. Nelson, C. L. McEvoy, and T. A. Schreiber. The University of South Florida word\n\nassociation, rhyme, and word fragment norms. http://www.usf.edu/FreeAssociation/, 1998.\n\n[21] H. Kucera and W. N. Francis. Computational Analysis of Present-day American Engish. Brown\n\nUniversity Press, Providence, 1967.\n\n9\n\n\f", "award": [], "sourceid": 1014, "authors": [{"given_name": "Yang", "family_name": "Xu", "institution": null}, {"given_name": "Charles", "family_name": "Kemp", "institution": null}]}