{"title": "Integration of Visual and Somatosensory Information for Preshaping Hand in Grasping Movements", "book": "Advances in Neural Information Processing Systems", "page_first": 311, "page_last": 318, "abstract": null, "full_text": "Integration of Visual and Somatosensory \n\nInformation for Preshaping Hand \n\nin Grasping Movements \n\nYoji  Uno \n\nATR Human Information Processing \n\nResearch Laboratories \n\n2-2 Hikaridai, Seika-cho, Soraku-gun, \n\nKyoto 619-02, Japan \n\nNaohiro Fukumura* \nFaculty of Engineering \n\nUniversity of Tokyo \n\n7-3-1  Hongo, Bunkyo-ku, \n\nTokyo  113, Japan \n\nRyoji Suzuki \n\nFaculty of Engineering \n\nUniversity of Tokyo \n\n7-3-1  Hongo, Bunkyo-ku, \n\nTokyo 113, Japan \n\nMitsuo Kawato \n\nATR  Human Information Processing \n\nResearch Laboratories \n\n2-2 Hikaridai, Seika-cho, Soraku-gun, \n\nKyoto 619-02, Japan \n\nAbstract \n\nThe primate brain must solve two important problems in grasping move(cid:173)\nments.  The  first  problem  concerns  the  recognition  of grasped objects: \nspecifically,  how  does  the  brain integrate  visual  and motor information \non a grasped object? The second problem concerns hand shape planning: \nspecifically, how does the brain design the hand configuration suited to the \nshape of the  object and the manipulation task?  A neural network model \nthat solves these problems has been developed.  The operations of the net(cid:173)\nwork are divided into a learning phase and an optimization phase.  In the \nlearning phase, internal representations, which depend on the grasped ob(cid:173)\njects and the task,  are  acquired by integrating visual and somatosensory \ninformation.  In the optimization phase, the  most suitable hand shape for \ngrasping an object is determined by using a relaxation computation of the \nnetwork. \n\n* Present  Address: \n\nParallel  Distributed  Processing  Research  Dept.,  Sony  Corporation, \n\n6-7-35 Kitashinagawa, Shinagawa-ku, Tokyo 141, Japan \n\n311 \n\n\f312 \n\nUno,  Fukumura, Suzuki,  and Kawato \n\nINTRODUCTION \n\n1 \nIt has  previously been established that,  while reaching out to  grasp an object,  the  human \nhand preshapes according to  the  shape of the object and the planned manipulation (Jean(cid:173)\nnerod,  1984;  Arbib et al.,  1985).  The preshaping of the human hand suggests that prior to \ngrasping an object the 3-dimensional form of the object is recognized and the most suitable \nhand configuration is preset depending on the manipulation task. \n\nIt is supposed that the human recognizes objects using not only visual information but also \nsomatosensory information when the  hand grasps them.  Visual information is  made from \nthe  2-dimensional  image  in  the  visual  system  of the  brain.  Somatosensory  information \nis  closely related  to  motor  information,  because  it depends on the  prehensile hand shape \n(Le.,  finger  configuration).  We  hypothesize  that  an  internal  representation  of a  grasped \nobject is  formed in the  brain by  integrating visual and somatosensory information.  Some \nphysiological studies support our hypothesis.  For example, Taira et al.  (1990)  found that \nthe activity of hand-movement-related neurons in the posterior parietal association cortex \nwere highly selective to the shape and/or the orientation of manipulated switches. \n\nHow  can  the  neural  network integrate different kinds of information?  Merely uniting vi(cid:173)\nsual image with somatosensory information does not lead to any interesting representation. \nOur basic idea is that information compression is applied to integrating different kinds of \ninformation.  It is useful to extract the essential information by compressing the visual and \nsomatosensory information. \n\nlrie  &  Kawato  (1991)  pointed out  that  multi-layered perceptrons  have  the  ability  to  ex(cid:173)\ntract  features  from  the  input  signals  by  compressing the  information from  input  signals. \nKatayama  &  Kawato  (1990)  proposed  a learning  schema in  which  an internal  represen(cid:173)\ntation of the grasped object was acquired using information compression.  Developing the \nschema of Katayama et al., we have devised a neural network model for recognizing objects \nand planning hand shapes (e.g., Fukumura et al.  1991). This neural network consists of five \nlayers of neurons with only forward connections as  shown in Figure 1.  The input layer (1 st \nlayer)  and the  output layer (5th layer) of the  network have  the  same  structure.  There are \nfewer neurons in the 3rd layer than in the 1st and 5th layers.  The operations of the network \nare  divided into the  learning  phase,  which is  discussed in section  2 and  the optimization \nphase, which is discussed in section 3. \n\n2 \n\nINTEGRATION OF VISUAL AND SOMATOSENSORY \nINFORMATION USING NETWORK LEARNING \n\nIn the learning phase, the neural network learns the relation between the visual information \n(Le.,visual image) and the somatosensory information which, in this paper,  is regarded as \ninformation on the prehensile hand configuration (Le., finger configuration). \n\nBoth vector x representing the visual image of an object and vector y representing the pre(cid:173)\nhensile hand configuration to grasp it are fed into the 1 st layer (the input layer).  The synaptic \nweights of the network are repeatedly adjusted so that the 5th layer outputs  the same vec(cid:173)\ntors  x and y as  are  fed  into the  1st layer.  In other words,  the network comes to realize the \nidentity map between the  1st layer and the 5th layer through a learning process.  The most \nimportant point of the neural network model is that the number of neurons in the 3rd layer \nis  smaller  than  the  number  of neurons  in the  1st  layer  (which is equal  to  the  number of \nneurons in the  5th layer).  Therefore, the information from x and y is compressed between \n\n\fVisual  &  Somatosensory  Information for  Preshaping Hand in  Grasping Movements \n\n313 \n\ny \n\n(hand) \n\nFigure  1:  A neural network model for integrating visual image x and prehensile hand con(cid:173)\nfiguration y.  The internal representation z of a grasped object is  acquired in the third layer. \n\nthe  1st layer and the 3rd layer, and restored between the  3rd layer and the 5th layer.  Once \nthe network learning process is complete, visual image x and prehensile hand configuration \nyare integrated in the network.  Consequently, the internal representation z of the grasped \nobject, which should include enough information to reproduce x and y, is formed in the 3rd \nlayer. \n\nPrehensile hand configuration in grasping movements  were measured and the learning of \nthe network was simulated by a computer. In behavioral experiments, three kinds of wooden \nobjects were prepared:  five circular cylinders whose diameters were 3 cm, 4 cm, 5 cm, 6 cm \nand 7 cm;  four quadrangular prisms whose side lengths were 3 cm, 4 cm, 5 cm and 6 cm; \nand three spheres  whose diameters were  3 cm, 4 cm and 5 cm.  Data input to  the  network \nwas comprised of visual image x and prehensile hand configuration y. \n\nVisual images of objects are formed through complicated processes in the visual system of \nthe brain.  For simplicity, however, projections of objects onto a side plane and/or a bottom \nplane  were  used  instead of real  visual  images.  The  area  of each  pixel  of the ,e-rojected \nimage was fed into the network as an element of visual image x.  A DataGloveT \n(V P L) \nwas  used  to  measure  finger  configurations  in grasping movements.  We  attached  sixteen \noptical fibers,  whose outputs were roughly inversely proportional to finger joint-angles, to \nthe  DataGlove.  The  subject was  instructed to  grasp the  objects on the  table  tightly  with \nthe palm and all  the  fingers.  The  subject grasped twelve objects thirty  times each,  which \nproduced 360 prehensile patterns for use as training data for network learning. \n\nIn  the computer simulation,  six  neurons were  set in  the  3rd layer.  The baCk-propagation \nlearning method was applied in order to modify the synaptic weights in the network.  Fig(cid:173)\nure 2 shows the activity of neurons in the  3rd layer after the learning had sufficiently been \nperformed.  Some interesting features  of the  internal representations were found in Figure \n2.  The first is  that the level of neuron activity in the  3rd layer increased monotonically as \nthe  size  of the  object increased.  The second is that, except for the  magnitude,  the  neuron \nactivation patterns  for  the  same kinds of objects were almost the  same.  Furthermore, the \nactivation patterns were  similar for  circular cylinders and quadrangular prisms,  but were \nquite different for  spheres.  In other words, similar representations were acquired for  simi(cid:173)\nlarly shaped objects. We concluded that the internal representations were formed in the 3rd \n\n\f314 \n\nUno,  Fukumura,  Suzuki,  and Kawato \n\n1.00 \n\nNeuron activity \np, \nQ'  \\'  0 \n~  \\ 0--, \nYI \n\\, \n\n,../  ~ \\ \n\n0.00 \n\nDiameter \n\n--3cm \n-o--4cm \n\n/ \nd  ,(  \\ \n\n\\6 Q, b -Q-Scm \nI P,  V\\'  -{)-6cm  o. \nj\\\"-o ~ \n\n' -O '7cm \n\n, \nJ \n\n\\ \n\n' \n\nNeuron activity \n\u00a3? \n1 \nI \n10  tJ \n\n1.00 \n\nP \n\nSide Length \n-3cm \n\n0 \n\\ \n\\ \n\\ \n\\  1 '\\ \n\\ \n'0/  b/\\ \\ \no.J /\\ \n\\0 \n~\\~~ \n1~ \n\n-Q'-4cm \n-D-Scm \n-D-6cm  0.00 \n\nNeuron activity \n\n1.00 \n\nDiameter \n..... -4cm \n\n-i:rScm \n\n-/:r6cm \n\n-I.OO-l--~1 ~2 ~3 ----;=4 -5 ::\"\":;:\"'6 ~ \n\n-1.00 \n\nNeuron index of the 3rd layer \na)  Circular  cylinder \n\n123456  \nNeuron index of the 3rd layer \nb) Quadrangular prism \n\n-1.00-l--~1 -2......-13~4~5~6~ \n\nNeuron index of the 3rd layer \n\nc) Sphere \n\nFigure 2:  Internal representations of grasped objects.  Graph a) shows the  neuron activa(cid:173)\ntion patterns for five  circular cylinders whose diameters were 3 cm, 4 cm, 5 cm, 6 cm and \n7 cm.  Graph b)  shows the neuron activation patterns  for four quadrangular prisms whose \nside lengths were 3 cm, 4 cm, 5 cm and 6 cm.  Finally, Graph c)  shows the neuron activa(cid:173)\ntion patterns  for  three spheres whose diameters were  3 cm,  4 cm  and 5 cm.  The abscissa \nrepresents the index of the six neurons in the  3rd layer, while the ordinate represents their \nactivity.  These values were normalized from -1  to + 1. \n\nlayer and changed topologically according to the shapes and sizes of the grasped objects. \n\n3  DESIGN OF PREHENSILE HAND SHAPES \nThe neural  network that has completed the  learning can design hand shapes to  grasp  any \nobjects in the optimization phase.  Determining prehensile hand shape (i.e.,  finger config(cid:173)\nuration) is  an ill-posed problem,  because there are  many  ways to grasp any given object. \nIn other words, prehensile hand configuration cannot be determined uniquely  for anyone \nobject.  In order to solve this indeterminacy, a criterion, a measure of performance for  any \npossible prehensile configuration is introduced. \n\nThe criterion should normally be defined based on the dynamics of the human hand and the \nmanipulation task.  However, for simplicity, the criterion is defined based only on the static \nconfiguration of the fingers,  which is represented by vector y.  We  assumed that the central \nnervous system adopts a stable hand configuration to grasp an object, which corresponds to \nflexing  the fingers  as  much as  possible.  The output of the DataGlove sensor decreases as \nfinger flexion increases. Therefore, the criterion C l (y) is defined as follows: \n\n(1) \n\nwhere Yi  represents the ith output of the sixteen DataGlove sensors.  Minimizing the crite(cid:173)\nrion C l (y) requires as much finger flexing as possible. \n\nFinding values of Yi  (i =  1,2, ... , 16) so as  to minimize C l (y) is an optimization problem \n\n\fVisual  &  Somatosensory  Information for Preshaping Hand in Grasping Movements \n\n315 \n\nwith constraints.  In the optimization phase, the neural network can solve this optimization \nproblem using a relaxation computation as follows.  When an object is  specified, the visual \nimage x*  of the object is input to the  1 st layer as an input signal and given to the 5th layer \nas a reference signal.  We  call neurons in the  1st and the 5th layers which represent visual \nimage x image neurons, and call neurons in the 1st and the 5th layers which represent finger \nconfiguration y hand neurons.  Let us define the following energy function of the network. \n\n1 ~_2 \nE(y) = 2 ~(Xi - Xi)  + 2 ~(Yj - Yj)  + A' 2:  ~ Yj' \n\n1 ~  I  2 \n\n1 ~ * \n\nI  2 \n\ni \n\nj \n\nj \n\n(2) \n\nHere,  xi  is  the  ith element of the  image x*  which is  fed  into the  ith  image  neuron in the \n1 st  layer,  and  x~ is  the  output of the  ith image neuron in the 5th layer.  Yj  is  the activity \nof the  jth hand neuron in the  1st layer,  and yj  is  the output of the jth hand neuron in the \n5th layer.  A is  a positive regularization parameter which decreases gradually during  the \nrelaxation  computation.  The  first  term  and  the  second  term  of equation  (2)  require  that \nthe  network realizes the identity map between the  input layer and the output layer as  well \nas  in the  learning phase.  This requirement guarantees  that  a hand whose configuration is \nspecified  by  vector y  can  grasp  an object  whose  visual  image  is  x*.  The  third  term  of \nequation (2)  represents  the  criterion  C 1 (y).  In the  optimization  phase,  the  values of the \nsynaptic weights are fixed.  Instead. the  hand neuron changes its state autonomously while \nobeying the following differential equation: \n\ndYk \nc  ds  = - aYk'  k = 1,2, ... ,16. \n\naE \n\n(3) \n\nHere.  s is the relaxation time  required for  the  state change of the  hand neuron,  and c is  a \npositive time constant.  The right-hand side of equation (3) can be transformed as follows: \n\n_  aE  = ~(x; _  x~) ax~ + ~(yj _  y',) ay}  + (Yk  _  yk) (ay~ - 1) - AYk. \n\n(4) \n\n8Yk  ~ \n\nI \n\n8Yk  ~  J  8Yk \n\nJ \n\n8Yk \n\nIt is straightforward to show that the first three terms of equation (4) are the error signals at \nthe  kth hand neuron, which can be calculated backward from the output layer to the input \nlayer.  The  fourth  term of equation (4)  is  a suppressive  signal which is  given  to  the  hand \nneuron by itself.  When the  state of the hand neuron obeys the differential equation (3), the \ntime change E can be expressed as : \n\ndE = L  dYk  aE  = -c L(dyk )2  < O. \nds \n\nk  ds  aYk \n\nds \n\nk \n\n-\n\n(5) \n\nTherefore, the energy function E always decreases and the network comes to the equilibrium \nstate  that  is  the  (local)  minimum energy  state.  The  outputs  of the  hand  neurons  in  the \nequilibrium state represent the solution of the optimization problem which corresponds to \nthe most suitable finger configuration. \n\nThe relaxation computation of the neural network was simulated. For example, when given \nthe image of a circular cylinder whose diameter was 5 cm, the prehensile finger configura(cid:173)\ntion was computed.  After a hundred-thousand iterations for the relaxation computation, we \nhad the results shown in Figure 3.  The left sied shows the hand shape that had the minimum \nvalue  of the criterion of all  the  training data recorded when the  subject grasped a circular \ncylinder whose diameter was 5 cm.  The right side shows the  hand shape produced by re(cid:173)\nlaxation computation.  These two hand shapes  were  very similar, which indicated that the \nnetwork reproduced hand shape by using relaxation computation. \n\n\f316 \n\nUno,  Fukumura, Suzuki, and Kawato \n\nTrainnig Data \n\nResults of relaxation \n\nFigure 3:  Prehesile hand shapes for grasping a circular cylinder whose diameter was 5cm. \n\n4  VARIOUS TYPES OF PREHENSIONS \n\nIn the  sections  above,  the  subject was instructed to  grasp objects  using only one type of \nprehension.  It is,  however,  thought  that a  human  chooses  different  types  of prehensions \ndepending on the manipulation tasks.  In order to investigate the dependence of the internal \nrepresentation on the type of prehension, the second behavioral experiment was conducted. \nIn this experiment, five  circular cylinders and three  spheres which were  the  same size as \nthose  in the  first experiment were prepared.  The subject was first  instructed to grasp  the \nobjects tightly with his palm and all of his fingers, and then to grasp the same objects with \nonly his fingertips.  Iberall et al.(1988) referred to the first prehension and the second pre(cid:173)\nhension as palm opposition  and pad opposition, respectively.  The  subject  grasped eight \nobjects in two different types of prehensions twenty times each, which produced 320 pre(cid:173)\nhensile patterns.  Four  neurons  were  set in the  3rd layer of the  network  and the  network \nlearning was simulated using these prehensile patterns as training data.  Figure 4 shows the \nneuronal activation patterns formed in the 3rd layer after the network learning.  Even if the \ngrasped objects were the same, the neuron activation pattern for palm opposition was quite \ndifferent from that for palm opposition. \n\nThe neural network can reproduce different prehension,  by introducing different criteria. \nCl (y) is definded corresponding to palm opposition.  Furthermore, we defind another crite(cid:173)\nrion C2(Y), corresponding to pad opposition. \n\ni{MP,CM \n\nC2(y) =  2:  yf + 2:(1.0 - Yj)2. \n\nidP \n\n(6) \n\nMinimizing the criterion C2(y) demands that the MP joints (metacarpophalangeal joints) of \nthe  four fingers  and the eM joint (carpometacarpal joint) of the thumb be flexed as  much \nas possible and that the IP joints (interphalangeal joints) of all five fingers  be  stretched as \nmuch as possible.  The relaxation computation of the neural network was simulated, when \ngiven the image of a sphere whose diameter was 5 cm.  The results of the relaxation compu(cid:173)\ntation are shown in Figure 5.  Adopting the different criteria, the neural network reproduced \ndifferent prehensile hand configurations which corresponded to a) palm opposition and b) \npad opposition. \n\n\fVisual  &  Somatosensory Information for  Preshaping Hand in  Grasping Movements \n\n317 \n\n1.00 \n\nNeuron activity  Neuron activity \nq \n\u20ac1 \n1.00  e.-fl-,  9 \n~  , \n\\. \n'Q\\  ~ \n\\.  ,',  \\ \n. , \n\\\\~/f\\~ \nDiameter \n\\  /  ~ laD \nV \n\n~ \n\n0 \n\n000 \n\n-0- - 4cm \n\n-0- Scm \n\n--0-- 6cm \n-0' 7cm \n\n0,00 \n\nNeuron activity \n\n100 \n\nNeuron activity \n\n100 \n\nDiameter \n\n- - 4cm \n\n{> \nI  I \nI \nI \nI \nI \n'A \nI \nI \n\n0.00  A \n\n~\\ \n\n~ \n\n,100 \n\n'1.oo_~..--~~ \n\n,1.00+--___  ..--. \n\n,1 2 3 4  \n\n12 34  \n\nNeuron index of the 3rt! layer  Neuron index of the 3rd layer \na)  Palm Opposition  b)  Pad  Oppositon \n\nCircular  cylinder \n\n1234  \n\n12 34  \n\nNeuron index of the 3rd layer \nNeuron index of the 3rd layer \nC)  Palm Opposition  d) Pad  Oppositon \n\nSphere \n\nFigure 4:  Internal representations of grasped objects formed in the 3rd layer of the network_ \nGraphs a), b), c) and d) show the activation patterns of neurons for palm oppositions when \ngrasping 5 circular cylinders, for pad oppositions when grasping 5 circular cylinders,  for \npalm oppositions when grasping 3 spheres and for pad oppositions when grasping 3 spheres, \nrespectively.  See Figure 2 legend for description. \n\n5  DISCUSSION \n\nIn view of the function of neurons in the posterior parietal association cortex, we have de(cid:173)\nvised a neural network model for integrating visual and motor information.  The proposed \nneural network model is  an active sensing model, as  it learns only when an  object is  suc(cid:173)\ncessfully  grasped.  In this  paper,  tactile  information is  not treated,  as  the materials  of the \ngrasped objects are ,not considered for simplicity.  We know  that tactile information plays \nan important role in the recognition of grasped objects.  The neural network model shown \nin Figure  1 can easily be developed so as to integrate visual, motor and tactile information. \nHowever,  it is not clear how  the internal representations of grasped objects is changed by \nadding tactile information. \n\nThe critical problem in our neural network model is how many neurons should be set in the \n3rd layer to represent the shapes of grasped objects.  If there are too few  neurons in the 3rd \nlayer, the 3rd layer cannot represent enough information to restor x and y between the 3rd \nlayer and the 5th layer; that is, the network cannot learn to realize the identity map between \nthe  input  layer and the  output layer.  If there  are  too  many  neurons  in  the  3rd layer,  the \nnetwork cannot obtain useful representations of the grasped objects in the 3rd layer and the \nrelaxation computation sometimes fails.  In the present stage, we have no method to decide \nan adequate number of neurons for the 3rd layer.  This is an important task for the future. \n\nAcknowledgements \nThe main part ofthis study was done while the first author (Y.U.) was working at University \nof Tokyo.  Y.  Uno,  N.  Fukumura and R.  Suzuki was  supported by  Japanese  Ministry  of \n\n\f318 \n\nUno, Fukumura, Suzuki,  and Kawato \n\nTrainnig Data  Result  of relaxation \n\nTraining Data  Result  of relaxation \n\na)  Plam Opposition \n\nb)  Pad opposition \n\nFigure 5:  Prehensile hand configuration a) for palm opposition and prehensile hand con(cid:173)\nfiguration  b) for  pad opposition  when grasping a sphere  whose diameter was  5 cm.  The \nleft sides show the hand shapes with the minimum values of the criterions for all  training \ndata recorded when the subject grasped a sphere whose diameter was 5 cm.  The right sides \nshow the hand shapes made by the relaxation computation. \n\nEducation,  Science and Culture Grants, NO.03251102 and No.03650338.  M.  Kawato was \nsupported by Human Frontier Science Project Grant. \n\nReferences \nM. Jeannerod.  (1984) The timing of natural prehension movements, J.  Motor Behavior, 16: \n235-254. \nM.A. Arbib, T.  Iberall and D. Lyons.  (1985) Coordinated control programs for movements \nof the  hand.  Hand Function and the Neocortex.  Experimental Brain Research,  suppl.l0, \n111-129. \nN. Fukumura, Y. Uno, R. Suzuki andK. Kawato (1991) A neural network model which rec(cid:173)\nognizes shape of a grasped object and decides hand configuration. Japan IEICE Technical \nReport, NC90-104:  213-218 (in Japanese). \nKatayama and  M.  Kawato  (1990)  Neural  network  model  integrating  visual  and  somatic \ninformation. J. Robotics Society of Japan, 8:  117-125 (in Japanese). \nT.  Iberall (1998) A neural network for  planning hand shapes in human prehension.  proc. \nAutomation and controls Con/.:  2288-2293. \nB. Irie and Kawato (1991) \"Acquisition of Internal Representation by Multilayered Percep(cid:173)\ntrons.\" Electronics and Communications in Japan, Part 3, 74:  112-118. \nM.  Taira,  S.  Mine,  A.P.  Georgopoulos,  A.  Murata and S.  Sakata.  (1990) Parietal  cortex \nneurons of the monkey related to the  visual guidance of hand movement.  Exp.  Brain Res., \n83:  29-36. \n\n\f", "award": [], "sourceid": 693, "authors": [{"given_name": "Yoji", "family_name": "Uno", "institution": null}, {"given_name": "Naohiro", "family_name": "Fukumura", "institution": null}, {"given_name": "Ryoji", "family_name": "Suzuki", "institution": null}, {"given_name": "Mitsuo", "family_name": "Kawato", "institution": null}]}