{"title": "Neurally Inspired Plasticity in Oculomotor Processes", "book": "Advances in Neural Information Processing Systems", "page_first": 290, "page_last": 297, "abstract": null, "full_text": "290 \n\nViola \n\nNeurally  Inspired  Plasticity in  Oculomotor \n\nProcesses \n\nPaul A.  Viola \n\nArtificial Intelligence  Laboratory \n\nM\"assachusetts Institute of Technology \n\nCambridge, MA  02139 \n\nABSTRACT \n\nWe  have  constructed  a  two axis  camera positioning system which \nis roughly analogous to a single human eye.  This Artificial-Eye (A(cid:173)\neye)  combines  the  signals  generated  by  two  rate gyroscopes  with \nmotion  information  extracted from  visual  analysis  to stabilize  its \ncamera.  This stabilization process is similar to the vestibulo-ocular \nresponse  (VOR);  like the  VOR, A-eye  learns a  system model that \ncan  be incrementally modified to adapt to changes in its structure, \nperformance and environment.  A-eye is an example of a robust sen(cid:173)\nsory system that performs computations that can be of significant \nuse  to the designers of mobile robots. \n\nIntroduction \n\n1 \nWe have constructed an  \"artificial eye\"  (A-eye), an  autonomous robot that incorpo(cid:173)\nrates a two axis camera positioning system (figure 1).  Like a the human oculomotor \nsystem, A-eye can estimate the rotation rate of its body with a  gyroscope and esti(cid:173)\nmate the rotation rate of its \"eye\" by measuring image slip acr~ its \"retina\".  Using \nthe  gyroscope  to sense rotation,  A-eye  attempts to stabilize its camera by driving \nthe  camera motors  to counteract  body  motion.  The  conversion of gyro output to \nmotor command is  dependent on the characteristics of the gyroscope, the structure \nof camera  lensing  system  and  the  response  of the  motors.  A  correctly  function(cid:173)\ning  stabilization  system  must  model  the  characteristics of each  of these  external \nvariables. \n\n\fNeurally Inspired Plasticity in Oculomotor Processes \n\n291 \n\nFigure 1: The construction of A-eye can  be viewed  in rough  analogy to the human \noculomotor system.  In  place  of an  eye,  A-eye  has  a  camera on  a  two axis position(cid:173)\ning  platform.  In  place  of the circular  canals  of the  inner  ear,  A-eye  has  two  rate \ngyroscopes  that measure  rotation in  perpendicular axes. \n\nSince  camera  motion  implies  stabilization  error,  A-eye  uses  a  visual  estimate  of \ncamera motion to incrementally update its system model.  \\Vhen  the camera is  cor(cid:173)\nrectly stabilized there is no statistically significant slip.  \\Vhenever a particular gyro \nmeasurement is associated  with  a  result camera motion,  A-eye makes an  incremen(cid:173)\ntal change to  its  response  to  that  particular  measurement  to  reduce  that error in \nthe future. \nA-eye  was  built  for  two  reasons:  to  facilitate  the  operation  of  complex  visually \nguided mobile  robots and to explore the applicability of simple learning techniques \nto the construction  of a  robust  robot. \n\n2  Autonomous Robots \nAn autonomous robot must function correctly  for  long periods of time without hu(cid:173)\nman intervention.  It is  certainly difficult  to create an autonomous robot or process \nthat will  function accurately, both initially and perpetually.  To achieve such a goal, \nautonomous processes must be  able to adapt both to unforeseen  aspects of the en(cid:173)\nvironment  and  inaccuracies  in  construction.  One  approach  to attaining  successful \nautonomous performance would entail the full characterization of the robot's struc(cid:173)\nture, its performance requirements, and its relationship with the environment.  Since \nclearly  both  the  robot  and  its  environment  are susceptible  to  change any  charac(cid:173)\nterization  could  not  be static.  In  contrast, our approach only  partially categorizes \nthe  robot's structure, environment,  and  task.  Without  more  detailed  information \ninitial performance is  inaccurate.  However,  by  using  a  measure of error  in  perfor(cid:173)\nmance  initially  partial  categorization  can  incrementally  improved.  In  addition,  a \nchange  to  system  performance  can  be  compensated  continually.  In  this  way  the \nextensive analysis  and  engineering  that  would  be  required  to  characterize, foresee \n\n\f292 \n\nViola \n\nand circumvent variability can be greatly reduced. \n\n3  The VOR \nThe oculomotor processes found  in  vertebrates are  well studied.  examples of adap(cid:173)\ntive,  visually  guided  processing  [Gou85].  The  three  oculomotor  processes  found \nalmost  universally  in  vertebrates  (the  vestibulo-ocular  response,  the  optokinetic \nsystem, and  the saccadic system), accurately  perform ocular positioning tasks with \nlittle or  no  conscious  direction.  The  response  times of these  systems  demonstrate \nthat  little  high  level,  \"conscious\",  processing  could  take  place.  In  a  limited  sense \nthese  processes  are  autonomous,  and  it  should  come  as  no  surprise  that  they  are \nquite plastic.  Such  plasticity is  necessary  to counteract the  foreseeable  changes  in \nthe  eye  due  to  growth  and  aging  and  the unforeseeable  changes due  to illness  and \ninjury. \n\nThe VOR works  to counteract the  motion of a  creature  in  its environment.  A  cor(cid:173)\nrectly  functioning  VOR ensures  that a  creature  \"sees\"  as  little  unintended  motion \nas  possible.  Miles  [FAM81]  and  others  have  demonstrated  that  the  VOR  is  an \nadaptive  motor  response,  capable  of significant  recalibration  in  a  matter  of days. \nAdaptation  can be  demonstrated by  the use  of inverting or  magnifying spectacles. \nWhile  wearing these  glasses  the  correct  orbital  motion  of the eye,  given  a  particu(cid:173)\nlar  head  motion,  is significantly  different  from  the  normal  response.  Initially,  the \nresponse  to head  motion  is  an  incorrect  eye  motion.  With  time  eye  motion  begins \nto  approach  the  correct  counteracting motion.  This  kind  of adaptation  allows  an \nanimal to continue functioning  in spite of injury or illness. \n\n4  The Device \nA-eye is  a  small autonomous robot  that incorporates a  CCD camera,  a  three wheel \nbase,  a  two axis  pitch/yaw camera positioning platform,  and  two  rate gyroscopes. \nOn  board  processing  includes  a  Motorola  microcontroller  and  68020  based  video \nprocessing board.  Including batteries, A-eye is  a foot high cylinder that is  12 inches \nwide.  In its present configuration A-eye can run autonomously for up to three hours \n(figure  2). \nA-eye's  goal  is  to learn  how  to  keep  its  camera stable  as  its  base  trundles  down \ncorridors.  There  are  two  sources  of information  regarding  the  motion  of A-eye's \nbase:  gyro rotation measurements and optical flow.  Rate gyroscopes  measure base \nrotation rate directly.  Visual analysis can be used  to estimate motion  by a  number \nof methods of varying complexity (see [Hil83]  for  a good overview).  By attempting \nto measure  only camera rotation from slip  complexity can be avoided.  The simple \nmethod  we have  chosen  measures the slip of images across  the  retina. \n\n4.1  Visual Rotation Estimation \n\nOur  approach  to  camera  rotation  estimation  uses  a  pre-processing  subunit  com(cid:173)\nmonly  known  as  a  \"Reichard  detector\"  which  for  clarity  we  will  call  a  shift  and \ncorrelate  .nit [PR73].  A  shift  and  cOJTelate  unit has as its inputs a  set of samples \n\n\fNeurally Inspired Plasticity in Oculomotor Processes \n\n293 \n\nFigure 2:  A photo of the current state of A-eye. \n\nfrom a  blurred area of the retina.  It shifts these inputs spatially and correlates them \nwith  a  previous,  unshifted  set of inputs.  \\\"hen  two  succeeding  images  are  identi(cid:173)\ncal  except  for  a  spatial  shift,  the  units  which  perform  that shift  respond  strongly. \nClearly the activity of a  shift  and  correlate  anit contains information about retinal \nmotion.  Due  to  the  size  and  direction  of shift,  some  detectors  will  be  sensitive \nto  small  motions,  others  large  motions,  and  each  will  be  sensitive  to  a  particular \ndirection of motion. \n\nThe  input  from  the  shift  and  correlate  units  is  used  to  build  value-unit  encoded \nretinal  velocity  map,  in  which  each  unit  is  sensitive  to  a  different  direction  and \nrange  of velocities.  The  map  has  9  units  in a  3  by  3  grid  (fig  3).  To  create  such \na  map,  each  of the  shift  and  correlate  uniu is  connected  to  every  map  unit.  By \nmoving the  camera,  displaced  images  that  are examples  of motion,  are  generated. \nThe  motor  command  that  generated  this  motion  example  corresponds  to  a  unit \nin  the  visual  velocity  map.  Connection  weights  are  updated  by  a  standard  least \nsquares learning rule.  In operation, the most  ac.tive  unit represents the estimate of \nvisual motion. \n\n4.2  Gyroscope Rotation Estimation \n\nContrary to  first  intuition,  vertebrates do  not  rely  on  visual  information  to stabi(cid:173)\nlize  their  eyes.  Instead  head  rotation  information  measured  by  the  inner  ear,  or \nthe  vestibula,  is  used keep  the eyes stable.  Animals do  not  measure ocular motion \ndirectly from  visual  information  for  two  reasons:  a)  the  response  rates of photore(cid:173)\nceptors prevent useful  visual processing during rapid eye movements [Gou85]  b) the \n\n\f294 \n\nViola \n\n(s)CD0 \n808 \n0CDG) \n\nFigure 3:  The 9 unit  velocity  map has  1 unit for  each of the  8  \"chess  moves\" . \n\nBase \nRotation \nRate \n\n--.  Gyroscope \nTransfa \nFunction \n\nInverting \nTransfer \nFunction \n\nMotor \nTransfer \nFunction \n\n...... Eye \n\nRotation \nRate \n\nFigure 4:  Open-loop  control of ocular position  based on  gyroscope  output. \n\nrequired  visual  analysis takes approximately  lOOms l .  These  difficulties  combine to \nprevent rapid response to unexpected head  and  body motions.  A-eye is  beset with \nsimilar limitations and  we  have chosen a  similar solution. \nThe output of the gyroscope  is  some  function  of head  rotation  rate.  Stabilization \nis  achieved  by  driving  the  ocular  motors  directly  in  opposition  to  the  measured \nvelocity  (fig  4).  This  counteract  rotation  of the  base  in  one  direction  by  moving \nthe camera in  the opposite direction.  Such an open-loop system is very simple  and \ncan  perform  well;  they  are  unfortunately  very  reliant  on  proper  calibration  and \nrecalibration  to maintain performance  [Oga70). \nA-eye  maintains calibration  information  in  the form  of a  function  from  gyroscope \noutput to motor velocity command.  This function  is an 8 unit gaussian radial basis \napproximation  network  (TP89).  Basis  function  approximation  has  excellent  com(cid:173)\nputational properties while representing wide  variety of smooth functions.  Weights \nare modified with a simple least squares update rule, based errors in camera motion \ndetected  visually. \n\n5  Training A-eye \nA-eye learns to perform the VOR in a two phase process.  First, the measurement of \nvisual motion is  calibrated to the generation of camera motion commands.  Second, \n\nlOcular following, the tendency to follow  the motion of a  !Cene in the &beence of head motion, \n\nhas a  typical latency of lOOms  [FM87). \n\n\fNeurally Inspired Plasticity in Oculomotor Processes \n\n295 \n\nthe stabilizing motor responses to gyroscope measurements are approximated.  This \napproximation is  modified  based on a  visual estimate of camera motion. \n\nBy observing motor commands and comparing them to the resulting visual motion, \na  map from  visual  motion  to appropriate motor command can  be learned.  To train \nthe visual  motion map, A-eye performs a set  of characteristic motions and observes \nthe  results.  Each  motor  command  is  categorized  as  one  of the  9  distinct  motions \nencoded  by  the  visual  motion  map.  With each  motion,  the  connections  from  shift \nand  correlate  units to  the  visual  motion  map  are  updated  so  that  issuing  a  motor \ncommand  results in  activity in  the correct visual  motion  unit.  Because  no reference \nis  made  to external  variables,  this  measure  of visual  motion  is  completely  relative \nto the function  of the camera motors.  The visual  motion  map plays the role of ertor \nsignal  for  later  learning. \nBy observing both the gyroscope output and the visual response from head motion, \nA-eye  learns  the  appropriate  compensating eye  motion  for  all  head  motions.  Eye \ncompensation motions are  the result of motor  commands generated by  the approx(cid:173)\nimation  network  applied  to  the  gyroscope  output.  Incorrect  responses  will  cause \nvisual mot.ion.  This motion, as measured by the visual motion detector,  is  the error \nsignal that  drives the modification of the approximation network.  This is  the heart \nof the adaptation in  the  VOR. \n\n5.1  Results \n\nWhile training the motion detector and approximation network there are 5 training \nevents per second  (the visual  analysis  takes  about  200  msec).  Training  the visual \nmotion detector can take up to 10 minutes (in a few environments the weights refuse \nto settle on  the  correct  values).  While  it is  possible  to  hand  wire  a  detector  that \nis  95%  accurate,  most  learned  detectors  worked  well,  attaining  85%  accuracy.  In \nboth  cases,  the  detectors  have  the  desirable  capability  of rejecting  object  motion \nwhenever  there  is  actual  camera  motion  (this  is  due  to  the  global  nature  of the \nanalysis). \nThe approximation  network converges to a function  that performs  well  in  minutes \n(figure  5) .  Analysis  of the  images generated  by  the camera leads  us  to bound  the \ncumulative  error  in  rotation  over  a  1  minute  trial  at  5  degrees  (we  believe  this \napproaches the accuracy limitations inherent in  the gyroscope). \n\nAn  approach to reducing this gyroscope error involves yet another oculomotor pr~ \ncess:  optokinetic  nystagmus  (OKN).  This  is  the  tendency  for  an  otherwise  undi(cid:173)\nrected eye to follow  visual  motion  in  the absence  of vestibular cues.  A-eye's visual \nmotion map  is  in  motor  coordinates.  By  directing the  camera in  the  opposite di(cid:173)\nrection from  observed motion,  residual errors in  VOR can be  reduced. \n\n6  Application \nWe  claim  that  the  stabilization  that  results  from  a  correctly  calibrated  VOR is \nuseful  both  for  navigation  and  scene  analysis.  A  stable  inertial  reference  can  act \nto  assist  tactical  navigation  when  traversing  rough  terrain.  Large  body  attitude \n\n\f296 \n\nViola \n\n4000. \n\n2000. \n\n-60. \n\n-40. \n\n20. \n\n40. \n\n60. \n\n-2000. \n\n-4000. \n\nFigure 5:  A  correct transfer function  (rough)  and  the learned (smoother)  approx(cid:173)\nimation. \n\nchanges, that can result from such travel, make it difficult to maintain a navigational \nbearing.  However,  when  there exists a  relatively stable inertial reference  fr :lme  less \nanalysis  need  be performed  to predict or sense  changes  in  bearing by other  means. \nThe VOR is especially applicable  to legged vehicles,  where the terrain and the form \nof locomotion  can  cause  constant rapid  changes  in  attitude  [Rai89]  [Ang89].  The \ntask  of  adapting  conventional  vision  systems  to  such  vehicles  is  formidable.  As \nthe  rate of pitching increases.  the quality of video  images  degrade,  while  the  task \nof finding  a  correspondence  between successive  images  will  increase  in  complexity. \nWith  the  addition  of the visual stabilization  that  A-eye  can  provide,  an  otherwise \ncomplex  visual  analysis task can  be  much simplified. \n\n7  Conclusions \nA-eye is  in part a  response to the observation  that static calibration is  a  disastrous \nweakness.  Static calibration not only forces  an engineer to expend additional effort \nat  design  time.  it requires  constant performance  monitoring and recalibration.  By \ncreating a  device that monitors its own  performance and adapts to changes.  signif(cid:173)\nicant work can be saved in design  and at numerous times during the lifetime of the \ndevice. \n\nA-eye  is  also  in  part  a  confirmation  that  simple.  tractable  and  reliable  learning \n. mechanisms are  sufficient to perform useful  motor learning. \n\nFinally. A-eye  is  in  part a  demonstration  that useful  visual  processing can  be per(cid:173)\nformed  in  real-time  with  an  reasonable  amount of computation.  This  processing \nyields  the additional side-benefit of simplifying the complex task  of visual  recogni(cid:173)\ntion. \n\n\fNeurally Inspired Plasticity in Oculomotor Processes \n\n297 \n\nAcknowledgements \n\nThis report describes research  done at the Artificial Intelligence Laboratory of the \nMassachusetts Institute of Technology.  Support  for  this  research  was  provided  by \nHughes  Artificial  Intelligence  Center  contract  #SI-804475-D,  the  Office  of Naval \nResearch contract NOOOI4-86-K-0685,  and the  Defense  Advanced Research Projects \nAgency  under Office  of Naval  Research contract  ~OOOI4-85-K-0124. \n\nReferences \n[Ang89]  Colin  Angle.  Genghis,  a  six  legged  autonomous  walking  robot.  Masterts \n\nthesis,  MIT,  1989. \n\n[FAM81]  S.  G.  Lisberger  F.  A.  Miles.  Plasticity in  the  vestibulo-ocular  reflex:  A \n\nnew  hypothesis.  Ann.  Rev.  Neurosci., 4:273-299,1981. \n\n[FM87]  K.  Kawano F.A.  Miles.  Visual stabilization of the eyes.  TINS, 4(10):153-\n\n158,  1987.  Reference on  Opto-kinetic nystagmus latency. \n\n[Gou85]  Peter Gouras.  Oculomotor system.  In  James Schwartz  Eric  Kandel,  ed(cid:173)\n\nitor,  Principles  of Neuroscience,  chapter 34.  Elsevier  Science  Publishing, \n1985. \n\n[Hil83] \n\nEllen  C.  Hildreth.  The  Measurement  of Visual  Motion.  The MIT Press, \n1983.  Good  book on  the extraction of motion  from edges. \n\n[Oga70]  Katsuhiko  Ogata.  Modem  Control  Engineering.  Prentice-Hall,  Engle(cid:173)\n\nwood  Cliffs,  N.J.,  1970.  Steady State Frequency Response  (page 372). \n\n[PR73]  T. Poggio and  W.  Reichard.  Considerations on  models of movement  de(cid:173)\n\ntection.  Kybernetic,  13:223-227,  1973. \n\n[Rai89]  Marc  H.  Raibert.  Trotting, pacing, and bounding by a quadruped robot. \n\nJournal  of Biomechanics,  1989. \n\n[TP89] \n\nFederico Girosi  Tomaso  Poggio.  A  theory  of networks for  approximation \nand learning.  AI  Memo  1140,  MIT,  1989. \n\n\f", "award": [], "sourceid": 266, "authors": [{"given_name": "Paul", "family_name": "Viola", "institution": null}]}