{"title": "An Information Maximization Model of Eye Movements", "book": "Advances in Neural Information Processing Systems", "page_first": 1121, "page_last": 1128, "abstract": null, "full_text": "An Information Maximization Model of \n\nEye Movements \n\n \n\n \n\nLaura Walker Renninger, James Coughlan, Preeti Verghese \n\nSmith-Kettlewell Eye Research Institute \n\n{laura, coughlan, preeti}@ski.org \n\n \n\nJitendra Malik \n\nUniversity of California, Berkeley \n\nmalik@eecs.berkeley.edu \n\nAbstract \n\nWe propose a sequential information maximization model as a \ngeneral strategy for programming eye movements. The model \nreconstructs high-resolution visual information from a sequence of \nfixations, taking into account the fall-off in resolution from the \nfovea to the periphery. From this framework we get a simple rule \nfor predicting fixation sequences: after each fixation, fixate next at \nthe location that minimizes uncertainty (maximizes information) \nabout the stimulus. By comparing our model performance to human \neye movement data and to predictions from a saliency and random \nmodel, we demonstrate that our model is best at predicting fixation \nlocations. Modeling additional biological constraints will improve \nthe prediction of fixation sequences. Our results suggest that \ninformation maximization is a useful principle for programming \neye movements. \n\n1 Introduction \n\nSince the earliest recordings [1, 2], vision researchers have sought to understand the \nnon-random yet idiosyncratic behavior of volitional eye movements. To do so, we \nmust not only unravel the bottom-up visual processing involved in selecting a \nfixation location, but we must also disentangle the effects of top-down cognitive \nfactors such as task and prior knowledge. Our ability to predict volitional eye \nmovements provides a clear measure of our understanding of biological vision. \nOne approach to predicting fixation locations is to propose that the eyes move to \npoints that are \u201csalient\u201d. Salient regions can be found by looking for center-\nsurround contrast in visual channels such as color, contrast and orientation, among \nothers [3, 4]. Saliency has been shown to correlate with human fixation locations \nwhen observers \u201clook around\u201d an image [5, 6] but it is not clear if saliency alone \ncan explain why some locations are chosen over others and in what order. Task as \nwell as scene or object knowledge will play a role in constraining the fixation \n\n\f \n\nlocations chosen [7]. Observations such as this led to the scanpath theory, which \nproposed that eye movement sequences are tightly linked to both the encoding and \nretrieval of specific object memories [8]. \n\n1.1 Our Approach \n\nWe propose that during natural, active vision, we center our fixation on the most \ninformative points in an image in order to reduce our overall uncertainty about what \nwe are looking at. This approach is intuitive and may be biologically plausible, as \noutlined by Lee & Yu [9]. The most informative point will depend on both the \nobserver\u2019s current knowledge of the stimulus and the task. The quality of the \ninformation gathered with each fixation will depend greatly on human visual \nresolution limits. This is the reason we must move our eyes in the first place, yet it \nis often ignored. A sequence of eye movements may then be understood within a \nframework of sequential information maximization. \n\n2 Human eye movements \n\nWe investigated how observers examine a novel shape when they must rely heavily \non bottom-up stimulus information. Because eye movements will be affected by the \ntask of the observer, we constructed a learn-discriminate paradigm. Observers are \nasked to carefully study a shape and then discriminate it from a highly similar one. \n\n2.1 Stimuli and Design \n\nWe use novel silhouettes to reduce the influence of object familiarity on the pattern \nof eye movements and to facilitate our computations of information in the model. \nEach silhouette subtends 12.5\u00ba to ensure that its entire shape cannot be characterized \nwith a single fixation. \nDuring the learning phase, subjects first fixated a marker and then pressed a button \nto cue the appearance of the shape which appeared 10\u00ba to the left or right of fixation. \nSubjects maintained fixation for 300ms, allowing for a peripheral preview of the \nobject. When the fixation marker disappeared, subjects were allowed to study the \nobject for 1.2 seconds while their eye movements were recorded. During the \ndiscrimination phase, subjects were asked to select the shape they had just studied \nfrom a highly similar shape pair (Figure 1). Performance was near 75% correct, \nindicating that the task was challenging yet feasible. Subjects saw 140 shapes and \ngiven auditory feedback. \n\nrelease fixation,\nrelease fixation,\nview object freely\nview object freely\n\n(1200ms)\n(1200ms)\n\nmaintain fixation\nmaintain fixation\n\n(300ms)\n(300ms)\n\nfixate, initiate trial\nfixate, initiate trial\n\nWhich shape is a match?\nWhich shape is a match?\n\n \n\nFigure 1. Temporal layout of a trial during the learning phase (left). Discrimination of \nlearned shape from a highly similar one (right). \n\n\f \n\n2.2 Apparatus \n\nRight eye position was measured with an SRI Dual Purkinje Image eye tracker while \nsubjects viewed the stimulus binocularly. Head position was fixed with a bitebar. A \n25 dot grid that covered the extent of the presentation field was used for calibration. \nThe points were measured one at a time with each dot being displayed for 500ms. \nThe stimuli were presented using the Psychtoolbox software [10]. \n\n3 Model \n\nWe wish to create a model that builds a representation of a shape silhouette given \nimperfect visual information, and which updates its representation as new visual \ninformation is acquired. The model will be defined statistically so as to explicitly \nencode uncertainty about the current knowledge of the shape silhouette. We will use \nthis model to generate a simple rule for predicting fixation sequences: after each \nfixation, fixate next at the location that will decrease the model\u2019s uncertainty as \nmuch as possible. Similar approaches have been described in an ideal observer \nmodel for reading [11], an information maximization algorithm for tracking \ncontours in cluttered images [12] and predicting fixation locations during object \nlearning [13]. \n\n3.1 Representing information \n\nThe information in silhouettes clearly resides at its contour, which we represent \nwith a collection of points and associated tangent orientations. These points and \ntheir associated orientations are called edgelets, denoted e1, e2, ... eN, where N is the \ntotal number of edgelets along the boundary. Each edgelet ei is defined as a triple \nei=(xi, yi, zi) where (xi, yi) is the 2D location of the edgelet and zi is the orientation \nof the tangent to the boundary contour at that point. zi can assume any of Q possible \nvalues 1, 2, \u2026, Q, representing a discretization of Q possible orientations ranging \nfrom 0 to\u03c0, and we have chosen Q=8 in our experiments. The goal of the model is \nto infer the most likely orientation values given the visual information provided by \none or more fixations. \n\n3.2 Updating knowledge \n\nThe visual information is based on indirect measurements of the true edgelet values \ne1, e2, ... eN. Although our model assumes complete knowledge of the number N and \nlocations (xi, yi) of the edgelets, it does not have direct access to the orientations zi.1 \nOrientation information is instead derived from measurements that summarize the \nlocal frequency of occurrence of edgelet orientations, averaged locally over a coarse \nscale (corresponding to the spatial scale at which resolution is limited by the human \nvisual system). These coarse measurements provide indirect information about \nindividual edgelet orientations, which may not uniquely determine the orientations. \nWe will use a simple statistical model to estimate the distribution of individual \norientation values conditioned on this information. \nOur measurements are defined to model the resolution limitations of the human \nvisual system, with highest resolution at the fovea and lower resolution in the \n \n1 Although the visual system does not have precise knowledge of location coordinates, the \nmodel is greatly simplified by assuming this knowledge. It is reasonable to expect that \nlocation uncertainty will be highly correlated with orientation uncertainty, so that the \ninclusion of location should not greatly affect the model's decisions of where to fixate next. \n\n\f)\n\n \n\n(\n\n(\n\n)\n\n,\n\nx\n\n)\n\n(\n\nf\n\n+\n\n=\n\nr\nf\n\n=)\n\nr \u2212\nx\n\nFPH\n\nx =r\n\nf =r\n\nyx\n,(\n\n2EE\n\nperiphery. Distance to the fovea is measured as eccentricity E, the visual angle \nis the location of a point in an image \nbetween any point and the fovea. If \nand \n is the fixation (i.e. foveal) location in the image then the \nf\ny\n, measured in units of visual degrees. The effective \neccentricity is \nE\nresolution of orientation discrimination falls with increasing eccentricity as \n where r(E) is an effective radius over which the visual system \n(Er\nspatially pools information and F\nOur model represents pooled information as a histogram of edge orientations within \nthe effective radius. For each edgelet ei we define the histogram of all edgelet \norientations ej within radius ri = r(E) of ei , where E is the eccentricity of \n)\n \n. To define the histogram more \nrelative to the current fixation \nprecisely we will introduce the neighborhood set Ni of all indices j corresponding to \n}i\nedgelets within radius ri of ei : \n, with number of \nr\nneighborhood edgelets |Ni|. The (normalized) histogram centered at edgelet ei is then \ndefined as \n\nPH =0.1 and E2=0.8 [14]. \n\nr\nxtsj\ni\n\nfr , i.e. \n\n{\nall\n\nyx\n,\ni\ni\n\nx =r\ni\n\nr \u2212\nx\ni\n\n..\n\nN\n\n\u2264\n\n\u2212\n\n=\n\nr\nx\n\nE\n\nr\nf\n\n=\n\ni\n\nj\n\nh\niz\n\n=\n\n1\nN\n\ni\n\n\u03b4 , \n\nzz\n,\n\nj\n\n\u2211\n\nNj\n\u2208\n\ni\n\nwhich is the proportion of edgelet orientations that assume value z in the \n(eccentricity-dependent) neighborhood of edgelet ei.2 \n \n\nFigure 2. Relation between eccentricity E and radius r(E) of the neighborhood (disk) \nwhich defines the local orientation histogram (hiz). Left and right panels show two \nfixations for the same object. \n\n \n\n \nUp to this point we have restricted ourselves to the case of a single fixation. To \ndesignate a sequence of multiple fixations we will index them by k=1, 2, \u2026, K (for \nK total fixations). The kth fixation location is denoted by\n. The \nquantities ri , Ni and hiz depend on fixation location and so to make this dependence \nexplicit we will augment them with superscripts as r\ni\n\n and h\n(k\niz\n\niN\n(k\n\n. )\n\n, \n\nr\nf\n\n=\n\nk\nx\n\nk\ny\n\n(k\n\n(\n\n)\n\n,\n\nf\n\nf\n\n,\n\nk\n\n)\n\n)\n\n)\n\n(\n\n \n2 \nyx,\u03b4 is the Kronecker delta function, defined to equal 1 if \n\nx = and 0 if \n\ny\n\ny\nx \u2260 . \n\n\f \n\n, z\n\nNow we describe the statistical model of edgelet orientations given information \nobtained from multiple fixations. Ideally we would like to model the exact \non \nconditioned \ndistribution \nof \norientations \nthe \ndata: \n | {h\nP(z\n, ... z\n{h}\n{h\n,\n,\nK\n)1(\n)2(\n(\nrepresents all histogram \n, where {h\n}k\n)(\nK\niz\niz\niz\niz\n)(k\ncomponents z at every edgelet e\n. This exact distribution is \ni for fixation \nintractable, so we will use a simple approximation. We assume the distribution \nfactors over individual edgelets: \n\nhistogram \n\n}\n)\n\n}\n,\n\nfr\n\nN\n\n2\n\n)\n\ni\n\nP(z\n\ni\n\n, z\n\n2\n\n, ... z\n\nN\n\n | {h\n)1(\niz\n\n{h}\n,\n)2(\niz\n\n}\n,\n\n)\n\n,\n\n{h\nK\n(\niz\n\n})\n\n=\n\nK\n\n \n\n(zg\ni\n\ni\n\n)\n\nN\n\n\u220f\n\ni\n\n1\n=\n\nwhere gi(zi) is the marginal distribution of orientation zi. Determining these \nmarginal distributions is still difficult even with the factorization assumption, so we \n, where Zi is a suitable \nwill make an additional approximation: \n\nK\n\n)\n\n(zg\ni\n\ni\n\n)\n\n=\n\n1\nZ\n\ni\n\nh\nk\n(\niz\n\n\u220f\n\nk\n\n1\n=\n\nnormalization factor. This approximation corresponds to treating \nas a likelihood \nfunction over z, with independent likelihoods for each fixation k. While the \napproximation has some undesirable properties (such as making the marginal \ni(zi) more peaked if the same fixation is made repeatedly), it provides a \ndistribution g\nsimple mechanism for combining histogram evidence from multiple, distinct \nfixations. \n\n)\n\nizh\n(k\n\n3.3 Selecting the next fixation \n\n( +Kfr\nGiven the past K fixations, the next fixation \nentropy of the edgelet orientations. In other words, \n\n)1\n\n)1\n\nis chosen to minimize \n\nis chosen to minimize the model \n( +Kfr\n, where the entropy of a \n{h\n(K+\niz\n. In practice, we minimize the \n\n})\n]\n\n)1\n\ni\n\n(\n\n2\n\nN\n\n)1\n\n)\n\nK +\n\n=\n\n, z\n\nP(z\n\n, ... z\n\nentropy\n[\n\nr\n{h}\nfH\n | {h\n,\n(\n)2(\n)1(\niz\niz\ndistribution P(x) is defined as \u2211\u2212\nxP\nlog)(\nentropy by evaluating it across a set of candidate locations \nwhich forms a \nregularly sampled grid across the image.3 We note that this selection rule makes \ndecisions that depend, in general, on the full history of previous K fixations. \n\n}\n,\n,\nK\nxP\n)(\n\n( +Kfr\n\n)1\n\nx\n\n4 Results \n\nFigure 3 shows an example of one observer\u2019s eye movements superimposed over the \nshape (top row), the prediction from a saliency model (middle row) [3] and the \nprediction from the information maximization model (bottom row). The information \nmaximization model updates its prediction after each fixation. \nAn ideal sequence of fixations can be generated by both models. The saliency model \nselects fixations in order of decreasing salience. The information maximization \nmodel selects the maximally informative point after incorporating information from \nthe previous fixations. To provide an additional benchmark, we also implemented a \n\n \n3 This rule evaluates the entropy resulting from every possible next fixation before making a \ndecision. Although this rule is suitable for our modeling purposes, it would be inefficient to \nimplement in a biological or machine vision system. A practical decision rule would use \ncurrent knowledge to estimate the expected (rather than actual) entropy. \n\n\f \n\n \n\nFigure 3. Example eye movement pattern, superimposed over the stimulus (top row), \nsaliency map (middle row) and information maximization map (bottom row). \n\nmodel that selects fixations at random. One way to quantify the performance is to \nmap a subject\u2019s fixations onto the closest model predicted fixation locations, \nignoring the sequence in which they were made. In this analysis, both the saliency \nand information maximization models are significantly better than random at \npredicting candidate locations (p < 0.05; t-test) for three observers (Figure 4, left). \nThe information maximization model performs slightly but significantly better than \nthe saliency model for two observers (lm, kr). If we match fixation locations while \nretaining the sequence, errors become quite large, indicating that the models cannot \naccount for the observed behavior (Figure 4, right). \n\nLocation Error\nLocation Error\n\nSequence Error\nSequence Error\n\n)\n)\ng\ng\ne\ne\nd\nd\n(\n(\n \n \ne\ne\ng\ng\nn\nn\nA\nA\n\nl\nl\n\n \n \nl\nl\n\na\na\nu\nu\ns\ns\nV\nV\n\ni\ni\n\nR S I\nR S I\n\nR S I\nR S I\n\nR S I\nR S I\n\nR S I\nR S I\n\nR S I\nR S I\n\nR S I\nR S I\n\n \n\nFigure 4. Prediction error of three models: random (R), saliency (S) and information \nmaximization (I) for three observers (pv, lm, kr). The left panel shows the error in \npredicting fixation locations, ignoring sequence. The right panel shows the error when \nsequence is retained before mapping. Error bars are 95% confidence intervals. \n\nThe information maximization model incorporates resolution limitations, but there \nare further biological constraints that must be considered if we are to build a model \nthat can fully explain human eye movement patterns. First, saccade amplitudes are \ntypically around 2-4\u00ba and rarely exceed 15\u00ba [15]. When we move our eyes, the \nimage of the visual world is smeared across the retina and our perception of it is \nactively suppressed [16]. Shorter saccade lengths may be a mechanism to reduce \nthis cost. This biological constraint would cause a fixation to fall short of the \nprediction if it is distant from the current fixation (Figure 5). \n\n\f \n\nFigure 5. Cost of moving the eyes. Successive fixations may fall short of the maximally \nsalient or informative point if it is very distant from the current fixation. \n\nSecond, the biological system may increase its sampling efficiency by planning a \nseries of saccades concurrently [17, 18]. Several fixations may therefore be made \nbefore sampled information begins to influence target selection. The information \nmaximization model currently updates after each fixation. This would create a \ndiscrepancy in the prediction of the eye movement sequence (Figure 6). \n\n \n\nFigure 6. Three fixations are made to a location that is initially highly informative \naccording to the information maximization model. By the fourth fixation, the subject \nfinally moves to the next most informative point. \n\n \n\n5 Discussion \n\nOur model and the saliency model are using the same image information to \ndetermine fixation locations, thus it is not surprising that they are roughly similar in \ntheir performance of predicting human fixation locations. The main difference is \nhow we decide to \u201cshift attention\u201d or program the sequence of eye movements to \nthese locations. The saliency model uses a winner-take-all and inhibition-of-return \nmechanism to shift among the salient regions. We take a completely different \napproach by saying that observers adopt a strategy of sequential information \nmaximization. In effect, the history of where we have been matters because our \nmodel is continually collecting information from the stimulus. We have an implicit \n\u201cinhibition-of-return\u201d because there is little to be gained by revisiting a point. \nSecond, we attempt to take biological resolution limits into account when \ndetermining the quality of information gained with each fixation. By including \nadditional biological constraints such as the cost of making large saccades and the \n\n\f \n\nnatural time course of information update, we may be able to improve our prediction \nof eye movement sequences. \nWe have shown that the programming of eye movements can be understood within a \nframework of sequential information maximization. This framework is portable to \nany image or task. A remaining challenge is to understand how different tasks \nconstrain the representation of information and to what degree observers are able to \nutilize the information. \n\nAcknowledgments \nSmith-Kettlewell Eye Research Institute, NIH Ruth L. Kirschstein NRSA, ONR #N00014-\n01-1-0890, NSF #IIS0415310, NIDRR #H133G030080, NASA #NAG 9-1461. \n\nReferences \n[1] Buswell (1935). How people look at pictures. Chicago: The University of Chicago Press. \n[2] Yarbus (1967). Eye movements and vision. New York: Plenum Press. \n[3] Itti & Koch (2000). A saliency-based search mechanism for overt and covert shifts of \nvisual attention. Vision Research, 40, 1489-1506. \n[4] Kadir & Brady (2001). Scale, saliency and image description. International Journal of \nComputer Vision, 45(2), 83-105. \n[5] Parkhurst, Law, and Niebur (2002). Modeling the role of salience in the allocation of \novert visual attention. Vision Research, 42(1), 107-123. \n[6] Nothdurft (2002). Attention shifts to salient targets. Vision Research, 42, 1287-1306. \n[7] Oliva, Torralba, Castelhano & Henderson (2003). Top-down control of visual attention in \nobject detection. Proceedings of the IEEE International Conference on Image Processing, \nBarcelona, Spain. \n[8] Noton & Stark (1971). Scanpaths in eye movements during pattern perception. Science, \n171, 308-311. \n[9] Lee & Yu (2000). An information-theoretic framework for understanding saccadic \nbehaviors. Advanced in Neural Processing Systems, 12, 834-840. \n[10] Brainard (1997). The psychophysics toolbox. Spatial Vision, 10 (4), 433-436. \n[11] Legge, Hooven, Klitz, Mansfield & Tjan (2002). Mr.Chips 2002: new insights from an \nideal-observer model of reading. Vision Research, 42, 2219-2234. \n[12] Geman & Jedynak (1996). An active testing model for tracking roads in satellite images. IEEE \nTrans. Pattern Analysis and Machine Intel, 18(1), 1-14. \n[13] Renninger & Malik (2004). Sequential information maximization can explain eye \nmovements in an object learning task. Journal of Vision, 4(8), 744a. \n[14] Levi, Klein & Aitesbaomo (1985). Vernier acuity, crowding and cortical magnification. \nVision Research, 25(7), 963-977. \n[15] Bahill, Adler & Stark (1975). Most naturally occurring human saccades have \nmagnitudes of 15 degrees or less. Investigative Ophthalmology, 14, 468-469. \n[16] Burr, Morrone & Ross (1994). Selective suppression of the magnocellular visual \npathway during saccadic eye movements. Nature, 371, 511-513. \n[17] Caspi, Beutter & Eckstein (2004). The time course of visual information accrual guiding \neye movement decisions. Proceedings of the Nat\u2019l Academy of Science, 101(35), 13086-90. \n[18] McPeek, Skavenski & Nakayama (2000). Concurrent processing of saccades in visual \nsearch. Vision Research, 40, 2499-2516. \n\n\f", "award": [], "sourceid": 2660, "authors": [{"given_name": "Laura", "family_name": "Renninger", "institution": null}, {"given_name": "James", "family_name": "Coughlan", "institution": null}, {"given_name": "Preeti", "family_name": "Verghese", "institution": null}, {"given_name": "Jitendra", "family_name": "Malik", "institution": null}]}