{"title": "A Productive, Systematic Framework for the Representation of Visual Structure", "book": "Advances in Neural Information Processing Systems", "page_first": 10, "page_last": 16, "abstract": null, "full_text": "A productive, systematic framework for \nthe representation of visual structure \n\nShimon Edelman \n\nNathan Intrator \n\n232 Uris Hall, Dept. of Psychology \n\nInstitute for Brain and Neural Systems \n\nCornell University \n\nIthaca, NY 14853-7601 \n\nse37@cornell.edu \n\nBox 1843, Brown University \n\nProvidence, RI 02912 \n\nN athan_Intrator@brown. edu \n\nAbstract \n\nWe describe a unified framework for the understanding of struc(cid:173)\nture representation in primate vision. A model derived from this \nframework is shown to be effectively systematic in that it has the \nability to interpret and associate together objects that are related \nthrough a rearrangement of common \"middle-scale\" parts, repre(cid:173)\nsented as image fragments. The model addresses the same concerns \nas previous work on compositional representation through the use \nof what+where receptive fields and attentional gain modulation. It \ndoes not require prior exposure to the individual parts, and avoids \nthe need for abstract symbolic binding. \n\n1 The problem of structure representation \n\nThe focus of theoretical discussion in visual object processing has recently started to \nshift from problems of recognition and categorization to the representation of object \nstructure. Although view- or appearance-based solutions for these problems proved \neffective on a variety of object classes [1], the \"holistic\" nature of this approach \n- the lack of explicit representation of relational structure - limits its appeal as a \ngeneral framework for visual representation [2]. \n\nThe main challenges in the processing of structure are productivity and system(cid:173)\naticity, two traits commonly attributed to human cognition. A visual system is \nproductive if it is open-ended, that is, if it can deal effectively with a potentially \ninfinite set of objects. A visual representation is systematic if a well-defined change \nin the spatial configuration of the object (e.g., swapping top and bottom parts) \ncauses a principled change in the representation (e.g., the interchange of the rep(cid:173)\nresentations of top and bottom parts [3, 2]). A solution commonly offered to the \ntwin problems of productivity and systematicity is compositional representation, in \nwhich symbols standing for generic parts drawn from a small repertoire are bound \ntogether by categorical symbolically coded relations [4]. \n\n\f2 The Chorus of Fragments \n\nIn visual representation, the need for symbolic binding may be alleviated by us(cid:173)\ning location in the visual field in lieu of the abstract frame that encodes object \nstructure. Intuitively, the constituents of the object are then bound to each other \nby virtue of residing in their proper places in the visual field; this can be thought \nof as a pegboard, whose spatial structure supports the arrangement of parts sus(cid:173)\npended from its pegs. This scheme exhibits shallow compositionality, which can \nbe enhanced by allowing the \"pegboard\" mechanism to operate at different spatial \nscales, yielding effective systematicity across levels of resolution. Coarse coding the \nconstituents (e.g., representing each object fragment in terms of its similarities to \nsome basis shapes) will render the scheme productive. We call this approach to the \nrepresentation of structure the Chorus of Fragments (CoF; [5]). \n\n2.1 Neurobiological building blocks \n\nWhat+ Where cells. The representation of spatially anchored object fragments pos(cid:173)\ntulated by the CoF model can be supported by what+where neurons, each tuned \nboth to a certain shape class and to a certain range of locations in the visual field. \nSuch cells have been found in the monkey in areas V 4 and posterior IT [6], and in \nthe prefrontal cortex [7]. \n\nAttentional gain fields. To decouple the representation of object structure from its \nlocation in the visual field, one needs a version of the what+where mechanism in \nwhich the response of the cell depends not merely on the location of the stimulus \nwith respect to fixation (as in classical receptive fields), but also on its location \nwith respect to the focus of attention. Indeed, modulatory effects of object-centered \nattention on classical RF structure (gain fields) have been found in area V 4 [8]. \n\n2.2 \n\nImplemented model \n\nOur implementation of the CoF model involves what+where cells with attention(cid:173)\nmodulated gain fields, and is aimed at productive and systematic treatment of \ncomposite shapes in object-centered coordinates. It operates directly on gray-level \nimages, pre-processed by a model of the primary visual cortex [9], with complex(cid:173)\ncell responses modified to use the MAX operation suggested in [10]. In the model, \none what+where unit is assigned to the top and one to the bottom fragment of the \nvisual field, each extracted by an appropriately configured Gaussian gain profile \n(Figure 2, left). The units are trained (1) to discriminate among five objects, \n(2) to tolerate translation within the hemifield, and (3) to provide an estimate of \nthe reliability of its output, through an autoassociation mechanism attempting to \nreconstruct the stimulus image [11, 12]. Within each hemifield, the five outputs \nof a unit can provide a coarse coding of novel objects belonging to the familiar \ncategory, in a manner useful for translation-tolerant recognition [13]. The reliability \nestimate carries information about category, allowing outputs for objects from other \ncategories to be squelched. Most importantly, due to the spatial localization of the \nunit's receptive field, the system can distinguish between different configurations of \nthe same shapes, while noting the fragment-wise similarities. \n\nWe assume that during learning the system performs multiple fixations of the target \nobject, effectively providing the what+where units with a basis for spanning the \n\n\fo \n\n\"9\" \n\n\"1 above center\" \n\nI I \n\nL_ ----\n\nWld'~' \n\n~~~~~~.-+ _---'''6'' \n\nwhat i r ! \n\n\"1\" \n\nj \n\nwherel \n\n, \n\"1 below center\" \n\n\"something \nbelow center\" \n\nFigure 1: Left: the CoF model conceptualized as a \"computation cube\" trained to \ndistinguish among three fragments (1, 6, 9), each possibly appearing at two loca(cid:173)\ntions (above or below the center of attention). A parallel may be drawn between \nthe computation cube and a cortical hypercolumn; in the inferotemporal cortex, \ncells selective for specific shapes may be arranged in columns, with the dimension \nperpendicular to the cortical surface encoding different variants of the same shape \n[14]. It is not known whether the attention-centered location of the shape, which \naffects the responses of V 4 cells [8], is mapped in an orderly fashion onto some phys(cid:173)\nical dimension(s) of the cortex. Right: the estimation of the marginal probabilities \nof shapes, which can be used to decide whether to allocate a unit coding for their \ncomposition, can be carried out simply by summing the activities of units along the \ndifferent dimensions of the computation cube. \n\nspace of stimulus translations. It is up to the model, however, to figure out that the \nobjects may be composed of recurring fragments, and to self-organize in a manner \nthat would allow it to deal with novel configurations of those fragments. This \nproblem, which arises both at the level of fragments and of their constituent features, \ncan be addressed within the Minimum Description Length (MDL) framework. \n\nSpecifically, we propose to construct receptive fields (RFs) for composite objects so \nas to capture the deviation from independence between the probability distributions \nof the responses of RFs tuned to their fragments. This implies a savings in the \ndescription length of the composite object. Suppose, for example, that r /l is the \nresponse of a unit tuned roughly to the top half of the character 6 and r h - the \nresponse of a unit tuned to its bottom half. The construction of a more complex \nRF combining the responses of these two units will be justified when \n\nP(r/l,rh)>> P(r/l)P(rh) \n\n(1) \n\nor, more practically, when some measure of deviation from independence between \nP(rfd and P(rh) is large (the simplest such measure would be the covariance, \nnamely, the second moment of the joint distribution but we believe that higher \nmoments may also be required, as suggested by the extensive work on measuring \ndeviation from Gaussian distributions). \n\nBy this criterion, a composite RF will be constructed that recognizes the two \"parts\" \n\n\fof the character 6 when they are appropriately located: the probability on the LHS \nof eq. 1 in that case would be proportional to 1/10, while the probability of the RHS \nwould be proportional to 1/100 (assuming that all characters are equiprobable, and \nthat their fragments never appear in isolation). At the same time, a composite RF \ntuned, say, to 6 above 3 (see section 3) will not be allocated, because the probability \nof such a complex feature as measured by either the RHS or the LHS of eq. 1 is \nproportional to 1/100. We note that this feature analysis can be performed on \nthe marginal probabilities of the corresponding fragments, which are by definition \nless sensitive to image parameters such as the exact location or scale, and can be \nbased on a family of features (cf. Figure 1). A discussion of this approach and \nof its relationship to the reconstruction constraint we impose when training the \nfragment-tuned modules is beyond the scope of this paper. \n\nA parallel can be drawn between the MDL framework just outlined and the findings \nconcerning what+where cells and gain fields in the shape processing pathway in the \nmonkey cortex. Under the interpretation we propose, the features at all levels \nof the hierarchy are coarsely coded, and each feature is associated with a rough \nlocation in the visual field, so that composite features necessarily represent more \ncomplex spatial structure than their constituents, without separately implemented \nbinding, and without a combinatorial proliferation of features. The computational \nexperiments described below concentrate on these novel characteristics of our model, \nrather than on the standard MDL machinery. \n\nReconstruction error Classification \n(modulatory signal) \n(output signal) \n\nFigure 2: The CoF model, trained on five composite objects (lover 6,2 over 7, etc.). \nLeft: the model consists of two what+where units, responsible for the bottom and \nthe top fragments of the stimulus, respectively. Gain fields (boxes labeled below \ncenter and above center) steer each input fragment to the appropriate unit. The \nlearning mechanism (RIC, for Reconstruction/Classification) was implemented as a \nradial basis function network. The reconstruction error (~) modulates the classifi(cid:173)\ncation outputs. Right: training the model, viewed as a computation cube. Multiple \nfixations of the stimulus (of which three are illustrated), along with Gaussian win(cid:173)\ndows selecting stimulus fragments, allow the system to learn what+where responses. \nA cell would only be allocated to a given fragment if it recurs in the company of a \nvariety of other fragments, as warranted by the ratio between their joint probability \nand the product of the corresponding marginal probabilities (cf. eq. 1 and Figure 1, \nright; this criterion has not yet been incorporated into the CoF training scheme). \n\n\f3 Computational experiments \n\nWe conducted three experiments that examined the properties of the structured \nrepresentations emerging from the CoF model. The first experiment (reported else(cid:173)\nwhere [13)), involved animal-like shapes and aimed at demonstrating basic produc(cid:173)\ntivity and systematicity. We found that the CoF model is capable of systematically \ninterpreting composite objects to which it was not previously exposed (for example, \na half-goat and half-lion chimera is represented as such, by an ensemble of units \ntrained to discriminate between three altogether different animals). \nIn the second experiment, a version of the CoF model (Figure 2) was charged with \nlearning to reuse fragments of the members of the training set -\nfive bipartite \nobjects composed of shapes of numerals from 1 through 0 -\nin interpreting novel \ncompositions of the same fragments. The gain field mechanism built into the CoF \nmodel allowed it to respond largely systematically to the learned fragments even \nwhen these were shown in novel locations, both absolute, and relative (Figure 3, \nleft). \n\nThe third experiment addressed a basic prediction of the CoF model, stemming \nfrom its reliance on what+where mechanisms: the interaction between effects of \nshape and location in object representation. Such interaction had been found in a \npsychophysical study [15], in which the task was 4-alternative forced-choice classi(cid:173)\nfication of two-part stimuli consisting of simple geometric shapes (cube, cylinder, \nsphere, cone). The composite stimuli were defined by two variables, shape and loca(cid:173)\ntion, each of which could be same, neutral, or different in the prime and the target \n(yielding 9 conditions altogether). Response times of human subjects revealed ef(cid:173)\nfects of shape and location (what+where) , but not of shape alone; the pattern of \npriming across the nine conditions was replicated by the CoF model (correlation \nbetween model and human data r = 0.85), using the same stimuli as in the psy(cid:173)\nchophysical experiment. \n\n4 Discussion \n\nBecause CoF relies on retinotopy rather than on abstract binding, its representation \nof spatial structure is location-specific; so is the treatment of structure by the human \nvisual system, as indicated by a number of findings. For example, priming in a \nsubliminal perception task was found to be confined to a quadrant of the visual field \n[16]. The notion that the representation of an object may be tied to a particular \nlocation in the visual field where it is first observed is compatible with the concept of \nobject file, a hypothetical record created by the visual system for every encountered \nobject, which persists as long as the object is observed. Moreover, location (as it \nfigures in the CoF model) should be interpreted relative to the focus of attention, \nrather than retinotopically [17]. \n\nThe idea that global relationships (hence, large-scale structure) have precedence \nover local ones [18], which is central to our approach, has withstood extensive testing \nin the past two decades. Even with the perceptual salience of the global and local \nstructure equated, subjects are able to process the relations among elements before \nthe elements themselves are identified [19]. More generally, humans are limited \nin their ability to represent spatial structure, in that the representation of spatial \nrelations requires spatial attention. For example, visual search is difficult when \n\n\fabove ___ 1.1 ____ 0 . 7 \nbelow __ 1 ____ 1_-\n\nbelow \n\n0 . 6 \n\n6 \n\n0 \n\n8 \n\n9 \n\n1 \n\n2 \n\n3 \n\n4 \n\n5 \n\n7 \n\n0.5 \n\n0.4 \n\n0.3 \n\n!J \n0 . 2 ~ \n\n0.1 \n\nabove \n\nmean \ncorrect ~ entropy \nper unit \n2 . 00 \n\n0 . 9 rate \n0.8 \n\nmean \n\nI 1. 75 \n\nf \\ \n\\ \nf \nf \n\\ \n\\ \nf \nI!J.... \nf \nf \n, \nf \n\n0 . 75 \n\n1.50 \n\n1.25 \n\n1.00 \n\ni!I, \n\n'c \n\n0 . 50 \n\n0 . 25 \n\n1 \n\n2 \n\n3 \n\n4 \n\n5 \n\n6 \n\n7 \n\n8 \n\n9 \n\n0 \n\n0 . 1 \n\n0.05 \n\n0.01 \n\n0.005 \n\nFigure 3: Left: the response of the CoF model to a novel composite object, 6 (which \nonly appeared in the bottom position in the training set) over 3 (which was only seen \nin the top position) . The interpretations offered by the model were correct in 94 out \nof the 100 possible test cases (10 digits on top x 10 digits on the bottom) in this \nexperiment. Note: in the test scenario, each unit (above and below) must be fed \neach of the two input fragments (above and below), hence the 20 bars in the plots \nof the model's output. Right: the non-monotonic dependence of the mean entropy \nper output unit (ordinate axis on the right; dashed line) on the spread constant a \nof the radial basis functions (abscissa) indicates that entropy alone should not be \nused as a training criterion in object representation systems. \n\ntargets differ from distractors only in the spatial relation between their elements, \nas if \" ... attention is required to bind features ... \" [20]. \n\nThe CoF model offers a unified framework, rooted in the MDL principle, for the \nunderstanding of these behavioral findings and of the functional significance of \nwhat+where receptive fields and attentional gain modulation. It extends the previ(cid:173)\nous use of gain fields in the modeling of translation invariance [21] and of object(cid:173)\ncentered herni-neglect [22], and highlights a parallel between whaHwhere cells and \nprobabilistic approaches to structure representation in computational vision (e.g., \n[23]). The representational framework we described is both productive and effec(cid:173)\ntively systematic. Specifically, it has the ability, as a matter of principle, to recog(cid:173)\nnize as such objects that are related through a rearrangement of mesoscopic parts, \nwithout being taught those parts individually, and without the need for abstract \nsymbolic binding. \n\nReferences \n[1] S. Edelman. Computational theories of object recognition. Trends in Cognitive Sci(cid:173)\n\nence, 1:296- 304, 1997. \n\n[2] J. E. Hummel. Where view-based theories of human object recognition break down: \nthe role of structure in human shape perception. In E. Dietrich and A. Markman, eds., \nCognitive Dynamics: conceptual change in humans and machines, ch. 7. Erlbaum, \nHillsdale, NJ, 2000. \n\n[3] R. F. Hadley. Cognition, systematicity, and nomic necessity. Mind and Language, \n\n12:137-153, 1997. \n\n\f[4] E. Bienenstock, S. Geman, and D. Potter. Compositionality, MDL priors, and object \nrecognition. In M. C. Mozer, M. I. Jordan, and T. Petsche, editors, NIPS 9. MIT \nPress, 1997. \n\n[5] S. Edelman. Representation and recognition in vision. MIT Press, Cambridge, MA, \n\n1999. \n\n[6] E. Kobatake and K. Tanaka. Neuronal selectivities to complex object features in the \nventral visual pathway of the macaque cerebral cortex. J. Neurophysiol., 71 :856- 867, \n1994. \n\n[7] S. C. Rao, G. Rainer, and E. K. Miller. Integration of what and where in the primate \n\nprefrontal cortex. Science, 276:821- 824, 1997. \n\n[8] C. E. Connor, D. C. Preddie, J. L. Gallant, and D. C. Van Essen. Spatial attention \n\neffects in macaque area V4. J. of Neuroscience, 17:3201- 3214, 1997. \n\n[9] D. J . Heeger, E. P. Simoncelli, and J. Anthony Movshon. Computational models of \n\ncortical visual processing. Proc. Nat. Acad. Sci., 93:623- 627, 1996. \n\n[10] M. Riesenhuber and T. Poggio. Hierarchical models of object recognition in cortex. \n\nNature Neuroscience, 2:1019- 1025, 1999. \n\n[11] D. Pomerleau. Input reconstruction reliability estimation. In C. L. Giles, S. J. Hanson, \n\nand J. D. Cowan, editors, NIPS 5, pages 279- 286. Morgan Kaufmann, 1993. \n\n[12] I. Stainvas, N. Intrator, and A. Moshaiov. Improving recognition via reconstruction, \n\n2000. preprint. \n\n[13] S. Edelman and N. Intrator. (Coarse Coding of Shape Fragments) + (Retinotopy) ~ \n\nRepresentation of Structure. Spatial Vision, 13:255- 264, 2000. \n\n[14] I. Fujita, K. Tanaka, M. Ito, and K. Cheng. Columns for visual features of objects in \n\nmonkey inferotemporal cortex. Nature, 360:343- 346, 1992. \n\n[15] S. Edelman and F. N. Newell. On the representation of object structure in human \nvision: evidence from differential priming of shape and location. CSRP 500, University \nof Sussex, 1998. \n\n[16] M. Bar and I. Biederman. Subliminal visual priming. Psychological Science, 9(6):464-\n\n469, 1998. \n\n[17] A. Treisman. Perceiving and re-perceiving objects. American Psychologist, 47:862-\n\n875, 1992. \n\n[18] D. Navon. Forest before trees: The precedence of global features in visual perception. \n\nCognitive Psychology, 9:353- 383, 1977. \n\n[19] B. C. Love, J. N. Rouder, and E. J. Wisniewski. A structural account of global and \n\nlocal processing. Cognitive Psychology, 38:291- 316, 1999. \n\n[20] A. M. Treisman and N. G. Kanwisher. Perceiving visually presented objects: recogni(cid:173)\ntion, awareness, and modularity. Current Opinion in Neurobiology, 8:218- 226, 1998. \n\n[21] E. Salinas and L. F. Abbott. Invariant visual responses from attentional gain fields. \n\nJ. of Neurophysiology, 77:3267- 3272, 1997. \n\n[22] S. Deneve and A. Pouget . Neural basis of object-centered representations. In M. I. \nJordan, M. J. Kearns, and S. A. Solla, editors, NIPS 11, Cambridge, MA, 1998. MIT \nPress. \n\n[23] M. C. Burl, M. Weber, and P. Perona. A probabilistic approach to object recogni(cid:173)\ntion using local photometry and global geometry. In Proc. 4th Europ. Conf. Com(cid:173)\nput. Vision, H. Burkhardt and B. Neumann (Eds.), LNCS-Series Vol. 1406- 1407, \nSpringer- Verlag, pages 628- 641, June 1998. \n\n\f", "award": [], "sourceid": 1827, "authors": [{"given_name": "Shimon", "family_name": "Edelman", "institution": null}, {"given_name": "Nathan", "family_name": "Intrator", "institution": null}]}