{"title": "Recognition of Manipulated Objects by Motor Learning", "book": "Advances in Neural Information Processing Systems", "page_first": 547, "page_last": 554, "abstract": null, "full_text": "Recognition  of  Manipulated  Objects \n\nby  Motor  Learning \n\nHiroaki  Gomi \n\nMitsuo  Kawato \n\nA TR Auditory and Visual Perception Research Laboratories, \n\nInui-dani, Sanpei-dani, Seika-cho, Soraku-gun, Kyoto 619-02, Japan \n\nAbstract \n\nWe present two neural network controller learning schemes based on feedback(cid:173)\nerror-learning and  modular architecture for recognition and control  of multiple \nmanipulated objects. In the first  scheme, a Gating Network is  trained to acquire \nobject-specific representations for recognition of a number of objects (or sets of \nobjects).  In  the  second  scheme,  an  Estimation  Network is  trained  to  acquire \nfunction-specific, rather than object-specific, representations which directly estimate \nphysical parameters. Both  recognition networks are trained to identify manipulated \nobjects  using  somatic  and/or  visual  information.  After  learning,  appropriate \nmotor  commands  for  manipulation  of each  object  are  issued by  the  control \nnetworks. \n\n1  INTRODUCTION \nConventional feedforward neural-network controllers  (Barto et aI.,  1983;  Psaltis et al., \n1987; Kawato et aI.,  1987, 1990; Jordan,  1988; Katayama &  Kawato,  1991) can not cope \nwith  multiple  or changeable  manipulated objects  or disturbances  because  they  cannot \nchange  immediately  the  control  law  corresponding  to  the  object.  In  interaction  with \nmanipulated objects or, in more general terms, in interaction with  an environment which \ncontains  unpredictable  factor,  feedback  information  is  essential  for  control  and object \nrecognition.  From  these  considerations,  Gomi  &  Kawato  (1990)  have  examined  the \nadaptive feedback controller learning schemes using feedback-error-Iearning,  from  which \nimpedance control  (Hogan, 1985) can be obtained automatically. However, in that scheme, \nsome higher system needs to supervise the setting of the appropriate mechanical  impedance \nfor each manipulated object or environment. \nIn this paper, we introduce semi-feedforward control schemes using neural networks which \nreceive feedback and/or feedforward information for recognition of multiple manipulated \nobjects based on feedback-error-learning and modular network architecture. These schemes \nhave two advantages over previous ones as follows.  (1) Learning is  achieved without the \n547 \n\n\f548 \n\nGomi and Kawato \n\nexact target motor command vector, which is unavailable during supervised motor learning. \n(2)  Although somatic information alone  was  found to be sufficient to recognize objects, \nobject identification is predictive and more reliable when both somatic and visual information \nare used. \n\n2  RECOGNITION  OF  MANIPULATED  OBJECTS \nThe most important issues in object manipulation are (l) how to recognize the manipulated \nobject and (2) how to achieve uniform performance for different objects. There are several \nways to acquire helpful information for recognizing manipulated objects. Visual information \nand somatic information (performance by motion) are most informative for object recognition \nfor manipulation. \nThe physical characteristics  useful  for object manipulation such as  mass,  softness and \nslipperiness, can not be predicted without the experience of manipulating similar objects. \nIn  this  respect,  object recognition  for  manipulation  should  be  learned  through  object \nmanipulation. \n\n3  MODULAR  ARCHITECTURE  USING  GATING  NETWORK \nJacobs et al. (1990,  1991) and Nowlan & Hinton (1990,  1991) have proposed a competitive \nmodular  network architecture  which  is  applied  to  the  task  decomposition  problem  or \nclassification problems. Jacobs (1991) applied this network architecture to the multi-payload \nrobotics  task  in  which  each  expert network  controller is  trained for  each  category  of \nmanipulated objects in terms of the object's mass. In his scheme, the payload's identity is \nfed  to the gating network to select a suitable expert network which acts as a feedforward \ncontroller. \nWe examined modular network architecture using feedback-e\"or-learning  for simultaneous \nlearning of object recognition and control task as  shown in Fig.l. \n\nM1Tt rU.\"nbmmOOi--._,\",So \n\nv \n\n\"'::\" \n,:f, \nt:~~ ~~t~ t:~ \n\nGatmg \nNetwork  .... +----. \n\nExpert Network 1 \n\nExpert Network 2  t--=-.c::u-~ \n\nExpert Network 3 \n\n+ \n\n... -_ .. \n\nu~ \nt -....... eo{4~\"\"\"\"\"-\"l~  object \n\nControlled \n\nu \n\n1--.... \n\nFig.1  Configuration of the modular architecture using Gating Network \n\nfor object manipulation based on feedback-error-learning \n\nIn this learning scheme, the quasi-target vector for combined output of expert networks is \nemployed instead of the exact target vector.  This is because it is unlikely that the exact \ntarget motor  command vector can be provided in  learning.  The quasi-target  vector of \nfeedforward motor command, u'  is produced by : \n\n'- + \nU  - U  Ufo ' \n\n(1) \n\n\fRecognition of Manipulated Objects by Motor Learning \n\n549 \n\nHere,  U  denotes  the  previous  final  motor command and  ufo denotes the feedback motor \ncommand.  Using  this quasi-target vector, the  gating and expert networks  are  trained to \nmaximize the log-likelihood function,  In L, by using backpropagation. \n\nIn L = In i gje -IU'-u,r /2(1,2 \n\n(2) \n\nj=! \n\nHere,  uj  is  the  i  th expert network output,  (Ij is a variance  scaling parameter of the  i  th \nexpert network and gj' the  i th output of gating network, is  calculated by \n\ngj  = -1 1  - ,  \n\neS, \nLesJ \n\nj=! \n\n(3) \n\nwhere  Sj  denotes the weighted input received by the i th  output unit.  The  total  output of \nthe modular network is \nuff = ~gjUj' \n\n(4) \n\n11 \n\nj=l \n\nBy maximizing Eq.2  using  steepest ascent method,  the  gating network learns  to  choose \nthe expert network whose output is closest to  the quasi-target command, and each expert \nnetwork is tuned correctly when it is chosen by the gating network. The desired trajectory \nis fed to the expert networks so as to make them work as feedforward controllers. \n\n4  SIMULATION  OF  OBJECT  MANIPULATION \nBY  MODULAR  ARCHITECTURE  WITH  GATING  NETWORK \nWe  show  the  advantage of the  learning schemes presented above by simulation results \nbelow.  The configuration of the  controlled object and manipulated object  is  shown  in \nFig.2  in  which  M,  B,  K  respectively denote  the  mass,  viscosity  and  stiffness  of the \ncoupled object (controlled- and manipulated-object). The manipulated object is changed \nevery epoch (l [sec]) while the coupled object is controlled to track the desired trajectory. \nFig.3  shows the selected object,  the feedforward and feedback motor commands, and the \ndesired and actual trajectories before learning. \n\nx \n\n-4--j \nM \n\nFig.2  Configuration of the controlled \nobject and the manipulated object \n\n- -- - -\n--\n--\nl~-a  __   -- -\n\n~ \n\nt:~ .24,------~1------r-----\"--~~1 \n\no \n\n5 \n\ntime  [ \u2022\u2022 c] \n\n20 \n\nFig.3  Temporal  patterns  of  the  selected \nobject,  the  motor  commands,  the  desired \nand actual trajectories before learning \n\nThe desired trajectory,  xd '  was  produced  by  Ornstein-Uhlenbeck random  process.  As \nshown in Fig.3, the error between the desired trajectory and the actual trajectory remained \nbecause  the  feedback  controller in  which  the  gains  were  fixed,  was  employed in  this \ncondition. (Physical characteristics of the objects used are listed in Fig.4a) \n\n\f550 \n\nGomi and Kawato \n\n4.1  SOMATIC  INFORMATION  FOR  GATING  NETWORK \nWe  call  the  actual  trajectory  vector,  x,  and  the  final  motor  command,  U  ,  \"somatic \ninfonnation\". Somatic infonnation should be most useful for on-line (feedback) recognition \nof the  dynamical  characteristics  of manipulated objects.  The  latest four  times  data of \nsomatic  information  were  used  as  the  gating  network  inputs  for  identification  of the \ncoupled object in this simulation. s ofEq.3 is expressed as: \n\ns(t) =  '1'1 (x(t), x(t -1), x(t - 2), x(t - 3), u(t), u(t -1), u(t - 2), u(t - 3\u00bb). \n\n(5) \n\nThe dynamical  characteristics  of coupled objects  are  shown  in  Fig.4a.  The  object was \nchanged in every epoch (l [secD.  The variance scaling parameter was  (Jj  = 0.8  and the \nlearning rates were  77ga,e  = 1. 0 x 10-3 and  77expert i  = 1. 0 x 10-5 \u2022 The three-layered feedforward \nneural network (input 16, hidden 30, output 3) was employed for the  gating network and \nthe two-layered linear networks (input 3, output 1) were used for the expert networks. \nComparing  the  expert's  weights  after learning and the  coupled object characteristics  in \nFig.4a, we realize  that expert networks  No.1,  No.2, No.3 obtained the inverse dynamics \nof coupled objects  y, (3,  a, respectively. The time variation of object, the gating network \noutputs, motor commands and trajectories after learning are  shown in Fig.4b.  The  gating \nnetwork  outputs  for  the  objects  responded  correctly  in  the  most  of the  time  and  the \nfeedback motor command,  ufo'  was  almost zero.  As  a  consequence  of adaptation,  the \nactual trajectory almost perfectly corresponded with the desired trajectory. \n\nb. \n\na.  Gating Net Outputs v.s. Objects \n\nrelinal \n\nchara::loristiicsl  ,mago  D-u::-;-\"\"\"\"::'::\"::-;:;':-:=-r-\"7.:\"\"':...--i \nM  B  K \n\n'Y  - - - ---\n\n- -\n\n-\n\na  1.0  2.0  8.0  none \n\nf3  5.0  7.0  4.0  none \n\n- 20 --\\---' __  ---\"'.-__  ----'-..L\"--__  ~~---.:,; \n\nL. \n\n1~L_ ____ L:::Lll\u00a3[ili\u00b1~ .... ~ .... =::O~ \n\n8.03.01 .0  none \n\n!1~actC:al \n\nIS.  __  \n\n~d.:: \n\no \n\n5 \n\n10 \n\ntlmo  [,.cl \n\n15 \n\n20 \n\nFig.4  Somatic  information  for  gating  network,  a.  Statistical  analysis  of the \ncorrespondence  of the  expert networks  with  each  object  after  learning  (averaged \ngating  outputs),  b.  Temporal patterns of objects,  gating  outputs,  motor commands \nand trajectories after learning \n\n4.2  VISUAL  INFORMATION  FOR  GATING  NETWORK \nWe usually assume the  manipulated object's characteristics by using visual infonnation. \nVisual information might be helpful for feedforward recognition. In this case,  s of Eq.3 is \nexpressed as: \n\ns(t) = 'l'2(V(t\u00bb) \n\n. \n\n(6) \n\nWe  used  three  visual  cues  corresponding  to each  coupled object in  this  simulation as \nshown  in Fig.5a.  At each epoch  in  this  simulation,  one  of three  visual  cues  selected \nrandomly is  randomly placed at one of four possible locations  on a  4 x 4 retinal matrix. \n\n\fRecognition of Manipulated Objects by Motor Learning \n\n551 \n\nThe visual cues of each object are different, but object ex  and ex*  have the same dynamical \ncharacteristics  as  shown  in Fig.5a.  The  gating  network  should identify  the  object and \nselect a suitable expert network for feedforward control by using this visual information. \nThe learning coefficients were  O\"j  = 0.7, 17gate  = 1. 0 X 10-3 ,  17eXpert  j  = 1. 0 X 10-5 .  The  same \nnetworks used in above experiment were used in this simulation. \nAfter learning, the expert network No.2 acquired the inverse dynamics of object ex  and ex * , \nand expert network  No.3 accomplished this for object  y.  It is recognized from Fig.5b that \nthe  gating network almost perfectly selected expert network No.2 for  object  ex  and ex*, \nand almost perfectly selected expert network  No.3  for  object y.  Expert network  No.1 \nwhich did not acquire inverse dynamics corresponding to any of the three objects, was not \nselected in the test period after learning. The actual trajectory in the test period corresponded \nalmost perfectly to the desired trajectory. \n\n- - ---\n--- --\n-\n\n-\n\na.  Gating Net. Outputs V.s.  Objects \n\nb. \n\n-\n\n-\n\nFig.  5  Visual  information  for  gating  network,  a.  Statistical  analysis  of the \ncorrespondence  of the  expert networks  with  each  Object  after learning  (averaged \ngating  outputs),  b. Temporal patterns of objects, gating outputs,  motor commands \nand trajectories after learning \n\ntlma  [sac] \n\n4.3  SOMATIC  &  VISUAL  INFORMATION  FOR  GATING  NETWORK \nWe show here the simulation results by using both of somatic and visual information as \nthe gating network inputs. In this  case, s ofEq.3 is represented as: \n\ns(t)= 'l'3(x(t),\u00b7\u00b7 \u00b7,x(t-3),u(t),\u00b7\u00b7\u00b7,u(t-3),V(t)). \n\n(7) \n\nIn this  simulation,  the  object ex  and  ~* had different dynamical characteristics, but shared \nsame visual cue as  listed in Fig.6a. Thus, to identify the coupled object one by one,  it is \nnecessary for the  gating network to  utilize not only visual information but also somatic \ninformation.  The  learning  coefficients  were  O\"j  = 1. 0, \n17gale  = 1. 0 X 10-3  and \n17expert j  = 1. 0 X 10-5 .  The gating network had 32 input units, 50 hidden units, and  1 output \nunit, and the expert networks were the same as in the above experiment. \nAfter  learning,  expert networks  No.1,  No.2,  No.3  acquired  the  inverse  dynamics  of \nobjects  y,  ~*,  ex  respectively.  As  shown in Fig.6b,  the  gating  network  identified  the \nobject almost correctly. \n\n\f552 \n\nGomi and Kawato \n\na.  Gating Net. Outputs v.s. Objects \n\nObjac1 \nphysical \ncharactonstics \nM  B  K \n\nb. \n\n-\n\n- --\n- - - -\n\n- 20~------~----~ ____________ ~ \n\nLL __ .J\u00a3~2Lill1iiliJill ___ \u2022 \u2022 \u2022  \"l  0  ~ \n\n1 \n\n-\n\n:2j~ \n\nactual \n\n8 \n\nj \no \n\n5 \n\n10 \n\ntime  [.ocJ \n\n15 \n\n20 \n\nFig. 6 Somatic & Visual information for gating network, a. Statistical analysis of the \ncorrespondence of the expert networks with each object after learning (averaged gating \noutputs),  b.  Temporal  patterns  of  objects,  gating  outputs,  motor  commands  and \ntrajectories after learning \n\n4.4  UNKNOWN  OBJECT  RECOGNITION \n\nBY  USING  SOMATIC  INFORMATION \n\nFig.7b  shows  the  responses  for  unknown  objects  whose  physical  characteristics  were \nslightly different from  known  objects  (see Fig.7a and Fig.4a) in the case using  somatic \ninformation as  the gating network inputs. Even if each tested object was not the  same as \nany  of the  known  (learned)  objects,  the  closest expert network  was  selected.  (compare \nFig.4a  and Fig.7a) During some period in the test phase, the feedback command increased \nbecause of an inappropriate feedforward command. \n\n- - - - --\n- -\n---\n-\n\n- -\n-\n\nb. \n\na.  Gating Net. Outputs v.s. Objects \n\nobject \nphysical \ncharactorisbCS \nM  B  K \n\nrotinal  II---..:-..---.-----..-:~_=_r____._;:_;;__t \nImago \n\nII--'--'----'--+---'--\"---'---+--'---'--\"-t \n\na'  2.0  3.0  7.0  none \n\n20 \n\n4.0  6 .0  5.0  none \n\n9.0  2.0  2.0  none \n\n~  0 ....,..\".~~,y 'i~~~k--~ ~ \n\n__ III  \u00b72 O~--....lIt,.------':'--i:..i,-__  ~--'-_~ \n\n~ \n\nFig. 7  Unknown objects recognition  by  using Somatic information,  a.  Statistical \nanalysis  of the correspondence  of the expert networks  with  each object after learning \n(averaged  gating  outputs),  b.  Temporal  patterns  of objects,  gating  outputs,  motor \ncommands and trajectories after learning \n\ntim.  [secJ \n\n\fRecognition of Manipulated Objects  by Motor Learning \n\n553 \n\n5  MODULAR  ARCHITECTURE \n\nUSING  ESTIMATION  NETWORK \n\nThe  previous  modular  architecture  is  competitive  in  the  sense  that  expert  networks \ncompete with each other to occupy its niche in the input space.  We  here propose a new \ncooperative modular architecture where expert networks specified for different functions \ncooperate to produce the required output. In this  scheme, estimation networks are trained \nto recognize physical characteristics of manipulated objects by using feedback information. \nUsing this  method, an infinite number of manipulated objects in the  limited domain can \nbe treated by  using  a small  number of estimation  networks.  We  applied this  method to \nrecognizing the mass of the manipulated objects. (see  Fig.8) \nFig.9a  shows  the  output  of the  estimation  network  compared  to  actual  masses.  The \nrealized trajectory almost coincided with the desired trajectory as  shown in Fig.9b.  This \nlearning  scheme can  be  applied not only  to  estimating  mass  but also  to  other physical \ncharacteristics such as softness or slipperiness. \n\na.  ~  8 \n6 \n~  4 \n~  2 \n\n0'r--_\"\"\"T\"\"'_\"\"\"\"''''-_-'--_--' \n2.0 \n\n0.5 \n\n0.0 \n\n1.5 \n\n1.0 \n\n~me (sec] \n\n-1 \n-2 \n-3 \n\nb. _ \n\n2 \n~  1 \n.~  a  -\"'--__ ~ I \n8. \n'..-1 \n\nactual traJecklry ,1\"\\ \n\\ \n\\ \n\ndesired traj9Ctory  f\\. \n\nI  ~ \n\nj ,i.x \n\no \n\n5 \n\n10 \n\ntime (sec] \n\n15 \n\n20 \n\nFig. 8  Confaguration of the modular architecture using \nmass estimation network for object manipulation by \nfeedback-error-Iearning \n\nFig. 9 a. Comparison of actual & \nestimated mass,  b.  desired &  actual \ntrajectory \n\n6  DISCUSSION \nIn  the  first  scheme,  the  internal  models  for  object manipulation  (in  this  case,  inverse \ndynamics)  were  represented  not in  terms  of visual  information  but  rather,  of somatic \ninformation (see 4.2).  Although the current simulation is  primitive,  it indicates the very \nimportant  issue  that  functional  internal-representations  of objects  (or  environments), \nrather than declarative ones, were acquired by motor learning. \nThe quasi-target motor command in the first scheme and the motor command error in the \nsecond  scheme  are  not always  exactly  correct in each time  step  because  the  proposed \nlearning schemes are based on the feedback-error-learning method. Thus, the learning rates \nin  the  proposed  schemes  should  be  slower  than  those  schemes  in which  exact target \ncommands  are employed.  In our preliminary simulation,  it was  about five  times  slower. \nHowever, we emphasize that exact target motor commands are not available in supervised \nmotor learning. \nThe limited number of controlled objects which can be dealt with by the modular network \nwith  a gating network  is  a  considerable problem  (Jacobs,  1991;  Nowlan,  1990,  1991). \nThis problem depends on choosing an appropriate number of expert networks and value of \nthe variance scaling parameter, (J'  .  Once this  is  done,  the expert networks can interpolate \n\n\f554 \n\nGomi and Kawato \n\nthe  appropriate output for  a number of unknown  objects.  Our second scheme provides a \nmore satisfactory solution to  this problem. \nOn the other hand, one possible drawback of the second scheme is that it may be difficult \nto estimate many physical parameters for complicated objects, even though the learning \nscheme  which  directly  estimates  the  physical  parameters  can  handle  any  number  of \nobjects. \nWe showed here basic examinations of two types  of neural networks - a gating network \nand a direct estimation network. Both networks use feedback and/or feedforward information \nfor recognition of multiple  manipulated objects.  In  future.  we  will  attempt to  integrate \nthese two  architectures in  order to model tasks involving skilled motor coordination and \nhigh level recognition. \n\nAck nowledgmen t \nWe would like to thank  Drs.  E. Yodogawa and K.  Nakane of AlR Auditory and Visual \nPerception Research Laboratories for their continuing encouragement. Supported by HFSP \nGrant to  M.K. \nReferences \nBarto, A.G.,  Sutton.  R.S., Anderson, C.W.  (1983)  Neuronlike adaptive elements that can \nsolve  difficult learning  control problems;  IEEE  Trans.  on  Sys.  Man  and  Cybern. \nSMC-13, pp.834-846 \n\nGomi, H., Kawato,  M.  (1990) Learning control for a closed loop system using feedback(cid:173)\n\nerror-learning.  Proc.  the  29th IEEE Conference  on Decision and Control,  Hawaii. \nDec., pp.3289-3294 \n\nHogan, N.  (1985) Impedance control: An approach to manipulation: Part I - Theory, Part \nII  - Implementation, Part III  - Applications,  ASME Journal  of Dynamic  Systems, \nMeasurement,  and Control, Vol. 107, pp.1-24 \n\nJacobs, R.A., Jordan,  M.I.,  Barto, A.G.  (1990) Task decomposition through competition \nin  a  modular  connectionist architecture:  The  what and where  vision  tasks,  COINS \nTechnical Report 90-27,  pp.1-49 \n\nJacobs, R.A.,  Jordan,  M.I.  (1991)  A  competitive modular  connectionist architecture. In \n\nLippmann, R.P.  et al.,  (Eds.) NIPS 3, pp.767-773 \n\nJordan.  M.I.  (1988)  Supervised learning and systems  with excess  degrees  of freedom, \n\nCOINS Technical Report 88-27, pp.1-41 \n\nKawato,  M.,  Furukawa,  K.,  Suzuki,  R.  (1987)  A hierarchical neural-network model for \n\ncontrol and learning of voluntary movement; Bioi.  Cybern.  57, pp.169-185 \n\nKawato, M.  (1990) Computational schemes and neural network models for formation and \ncontrol of multijoint arm trajectory.  In:  Miller,  T.,  Sutton, R.S.,  Werbos,  P.J.(Eds.) \nNeural Networksfor Control, The  MIT Press, Cambridge, Massachusetts, pp.197-228 \nKatayama, M.,  Kawato,  M.  (1991) Learning trajectory  and force  control  of an artificial \nmuscle arm by parallel-hierarchical neural network model. In Lippmann, R.P. et al., \n(Eds.) NIPS 3, pp.436-442 \n\nNowlan,  S.J.  (1990)  Competing experts:  An experimental  investigation  of associative \n\nmixture models, Univ.  Toronto  Tech.  Rep. CRG-TR-90-5,  pp.I-77 \n\nNowlan,  S.1.,  Hinton,  G.E. (1991) Evaluation of adaptive mixtures of competing experts. \n\nIn Lippmann, R.P.  et al.,  (Eds.) NIPS 3,  pp.774-780 \n\nPsaltis,  D.,  Sideris,  A.,  Yamamura,  A.  (1987) Neural controllers,  Proc.  IEEE Int.  Con! \n\nNeural Networks,  Vol.4, pp.551-557 \n\n\f", "award": [], "sourceid": 543, "authors": [{"given_name": "Hiroaki", "family_name": "Gomi", "institution": null}, {"given_name": "Mitsuo", "family_name": "Kawato", "institution": null}]}