{"title": "Visual Speech Recognition with Stochastic Networks", "book": "Advances in Neural Information Processing Systems", "page_first": 851, "page_last": 858, "abstract": null, "full_text": "Visual Speech Recognition with \n\nStochastic Networks \n\nJavier R.  Movellan \n\nDepartment of Cognitive Science \nUniversity of California San Diego \n\nLa Jolla, Ca 92093-0515 \n\nAbstract \n\nThis  paper presents  ongoing  work  on  a  speaker  independent  visual \nspeech recognition system. The work presented here builds on  previous \nresearch  efforts \nin  this  area  and  explores  the  potential  use  of simple \nhidden  Markov  models  for  limited  vocabulary,  speaker  independent \nvisual  speech recognition.  The  task  at  hand  is  recognition  of the  first \nfour  English  digits,  a  task  with  possible  applications  in  car-phone \nimages  were  modeled  as  mixtures  of  independent \ndialing.  The \nGaussian  distributions,  and  the \ntemporal  dependencies  were  captured \nwith standard left-to-right hidden Markov  models.  The  results  indicate \nthat \nsimple  hidden  Markov  models  may  be  used  to  successfully \nrecognize relatively unprocessed image sequences. The system  achieved \nperformance  levels  equivalent  to  untrained  humans  when  asked  to \nrecognize the fIrst four English digits. \n\n1  INTRODUCTION \n\nVisual articulation is an important source of information in face to  face  speech  perception. \nLaboratory studies have shown that visual information allows subjects to tolerate an  extra \n4-dB of noise in the acoustic signal. This  is  particularly important  considering that  each \ndecibel  of  signal  to  noise  ratio  translates  into  a  10-15%  error  reduction  in  the \nintelligibility  of entire  sentences  (McCleod  and  SummerfIeld,  1990). Lip  reading  alone \nprovides a basis for understanding for  a  large  majority  of the  hearing  impaired and  when \nsupplemented by acoustic or electrical signals it allows  fluent  understanding  of speech  in \nhighly  trained  subjects.  However  visual \ninformation  plays  more  than  a  simple \ncompensatory  role  in  speech  perception.  From  early  on  humans  are  predisposed  to \nintegrate acoustic and visual information.  Sensitivity  to  correspondences  in  auditory  and \nvisual  information  for  speech  events  has  been  shown  in  4  month  old  infants  (Spelke, \n1976; Kuhl &  Meltzoff,  1982). By 6  years  of age,  humans  consistently  use  audio  visual \ncontingencies  to  understand  speech (Massaro,  1987).  By  adulthood,  visual  articulation \nautomatically modulates perception of the acoustic signal.  Under  laboratory  conditions  it \nis  possible to create  powerful  illusions  in  which  subjects  mistakenly  hear sounds  which \nare biased by visual articulations.  Subjects  in  these  experiments  are  typically  unaware  cf \n\n\f852 \n\nJavier  Movellan \n\nthe  discrepancy  between  the  visual  and  auditory tracks  and their experience  is  that  of a \nunified auditory percept (McGurk &  McDonnald,  1976). \nRecent years  have  seen  a  revival  of interest  in  audiovisual  speech  perception  both  in \npsychology  and  in  the  pattern recognition  literature.  There  have  been  isolated  efforts  to \nbuild  synthetic  models  of visual  and  audio-visual  speech  recognition  (Petahan,  1985; \nNishida,  1986;  Yuhas, Goldstein,  Sejnowski &  Jenkins,  1988;  Bregler,  Manke,  Hild  & \nWaibel,  1993; Wolff, Prassad,  Stork, & Hennecke,  1994). The  main  goal  of these  efforts \nto  explore  different  architectures  and  visual  processing  techniques  and  to \nhas  been \nillustrate  the  potential  use  of visual  information  to  improve  the \nrobustness  of current \nspeech  recognition  systems.  Cognitive  psychologists  have  also  developed  high  level \nmodels  of audio-visual  speech  perception  that  describe  regularities  in  the  way  humans \nintegrate  visual  and  acoustic  information  (Massaro,  1987).  In  general  these  studies \nsupport the  idea  that  human  responses  to  visual  and  acoustic  stimuli  are  conditional \nindependent.  This  regularity  has  been  used  in  some  synthetic  systems  to  simplify  the \ntask  of integrating  visual  and  acoustic  signals  (Wolff,  Prassad,  Stork,  &  Hennecke, \n1994). Overall, multimodal speech perception is  still an emerging  field  in  which  a  lot  cf \nexploration needs to  be  done.  The  work presented here  builds  on  the  previous  research \nefforts  in  this  area  and  explores  the  potential  use  of simple  hidden  Markov models  1ir \nlimited  vocabulary,  speaker independent visual  speech recognition.  The  task  at  hand  is \nrecognition of the first four English digits, a task  with  possible  applications  in  car-phone \ndialing. \n\n2  TRAINING SAMPLE \n\nThe  training  sample  consisted  of 96  digitized  movies  of 12  undergraduate  students  (9 \nmales,3 females) from the Cognitive Science Department at UCSD. Video  capturing was \nperformed  in  a  windowless  room  at  the  Center  for  Research  in  Language  at  UCSD. \nSubjects were asked to talk into a video camera and to say  the first four digits in  English \ntwice.  Subjects  could  monitor  the  digitized  images  in  a  small  display  conveniently \nlocated in front of them. They were asked to position themselves so that that their lips  be \nroughly centered  in  the  feed-back  display.  Gray  scale video  images  were  digitized  at  30 \nfps,  100x75 pixels,  8 bits per pixel. The video tracks were hand segmented by selecting a \nfew relevant frames before and after the beginning and end of activity in the  acoustic  track. \nStatistics of the entire training sample are shown in table  1. \n\nTable 1: Frame number statistics. \n\nDigit \n\"One\" \n\"Two\" \n\"Three\" \n\"Four\" \n\nAverage \n\n8.9 \n9.6 \n9.7 \n10.6 \n\nS.D. \n2.1 \n2.1 \n2.3 \n2.2 \n\n3  IMAGE PREPROCESSING \n\nThere are two different approaches to visual preprocessing in the visual speech recognition \nliterature (Bregler, Manke,  Hild &  Waibel,  1993). The  first  approach,  represented by  the \nwork  of Wolff  and  colleagues  (Wolff,  Prassad,  Stork,  &  Hennecke,  1994)  favors \nsophisticated  image  preprocessing  techniques  to  extract  a  limited  set  of hand-crafted \nfeatures  (e.g.,  height  and  width  of the  lips).  The  advantage  of this  approach  is  that  it \n\n\fVisual  Speech  Recognition  with  Stochastic  Networks \n\n853 \n\ndrastically reduces the number of input dimensions. This translates  into  lower variability \nof the signal, potentially improved  generalization,  and  large  savings  in  computing  time. \nThe disadvantage is that vital information may be lost when  compressing the  image  into \na  limited  set  of hand-crafted features.  Variability  is  reduced  at  the  possible  expense  cf \nbias. Moreover, tests have  shown  that  subtle  holistic  features  such  as  the  wrinkling  and \nprotrusion of the lips may play  an  important role  in  human  lip-reading  (Montgomery & \nJackson,  1983).  The  second approach  to  visual  preprocessing  emphasizes  preserving the \noriginal  images  as  much  as  possible  and  letting  the  recognition  engine  discover  the \nrelevant  features  in  the  images.  In  most  cases,  images  are  low-pass  filtered  and \ndimension-reduced by  using  principal  component  analysis.  The  results  in  this  papers \nindicate that  good results can be obtained even without the use  of principal  components. \nIn this investigation image preprocessing consisted of the following phases: \n\n1.  Symmetry  enforcement:  At  each  time  frame  the  raw  images  were  symmetrized  by \naveraging  pixel by pixel the left and right side of each  image,  using  the  vertical  midline \nas the axis of symmetry. For convenience from  now on we will refer to the  raw  images  as \n\"rho-images\" and the symmetrized  images  as  \"sigma-images.\"  The  potential  benefits  cf \nsigma-images  are  robustness,  and  compression,  since  the  number  of relevant  pixels  is \nreduced by half. \n\n2.  Temporal differentiaion: At each time frame we calculated the pixel by  pixel  differences \nbetween present sigma-images and  immediately past sigma-images.  For  convenience  we \nrefer to  the resulting images as \"delta-images.\" One  of the  potential  advantages  of delta(cid:173)\nimages  in  the  visual  domain  is  their robustness  to  changes  in  illumination  and the  fuct \nthat they emphasize the dynamic aspects of the visual track. \n\n3.  Low pass filtering and subsampling: The sigma and delta images were compressed and \nsubsampled  using  20x 15  equidistant  Gaussian  filters.  Different  values  of the  standard \ndeviation of the Gaussian filters were tested. \n\n4.  Logistic  thresholding  and  scaling:  The  sigma  and  delta  images  were  independently \nthresholded  by  feeding  the  output  of the  Gaussian  filters  through  a  according  to  the \nfollowing equation \n\n7r \nY = 256 j(K  r;; \n\n(x - J.L)) \n\n-v3a \n\nwhere  f is  the  logistic  function,  and  J.L, a,  are  respectively  the  average  and  standard \ndeviation  of the  gray  level  distribution  of  entire  image  sequences.  The  constant  K \ncontrols  the  sharpness  of the  logistic  function.  Assuming  an  approximately  Gaussian \ndistribution  of gray  levels  when  K=1  the  thresholding  function  approximates  histogram \nequalization,  a  standard  technique  in  visual  processing.  Three  different  K  values  were \ntried:  0.3, 0.6  and  1.2. \n\n5.  Composites of the relevant portions of the blurred sigma  and  delta  images  were  fed  to \nthe  recognition  network.  The  number of pixels  of each  processed  image  was  300  (150 \nfrom the blurred  sigma  images  and  150  from  the  blurred delta  images).  Figure  1 shows \nthe effect of the different preprocessing stages. \n\n\f854 \n\nJavier  Movellan \n\nFigure  1:  Image Preprocessing.  1) Rho-Image. 2) Sigma-Image. \n\n3) Delta-Image. 4) Filtered and Sharpened Composite. \n\n4  RECOGNITION NETWORK \nWe used the standard approach in  limited vocabulary systems: a bank of hidden  Markov \nmodels,  one  per  word  category,  independently  trained  on  the  corresponding  word \ncategories. The images were modeled as mixtures of continuous  probability  distributions \nin  pixel  space.  We  tried  mixtures  of Gaussians  and  mixtures  of Cauchy  distributions. \nThe mixtures of Cauchy distributions were very stable  numerically  but  they  did  perform \nvery  poorly  when compared to  the  Gaussian mixtures.  We  believe  the  reason  for  their \npoor  performance  is  the  tendency  of Cauchy-based  maximum-likelihood  estimates  to \nfocus on  individual  exemplars.  Gaussian-based estimates  are  much  more  prone to  blend \nexemplars  that  belong  to  the  same  cluster.  The  initial  state  probabilities,  transition \nprobabilities, mixture coefficients, mixture centroids and variance  parameters  were  trained \nusing the E-M algorithm. \nWe  initially  encountered  severe  numerical  underflow  problems  when  using  the  E-M \nalgorithm  with  Gaussian  mixtures.  These  instabilities  were  due  to  the  fact  that  the \nprobability densities of images rapidly went to zero due to the large dimensionality of the \nimages. Trimming the outputs of the Gaussian and using  very  small  Gaussian  gains  did \nnot work well. We solved the numerical problems in the following  way:  1)  Constraining \nall  the  variance  parameters  for  all  the  states  and  mixtures  to  be  equal.  This  allowed \npulling  out  a  constant  in  the  likelihood-function  of  the  mixtures,  avoiding  most \nnumerical  problems.  2)  Initializing  the  mixture  centroids  using  linear  segmentation \nfollowed by the K-means clustering algorithm. For example, ifthere were  4  visual  frames \nand 2 states, the first 2 frames were assigned to state  1 and the last 2 frames to state  2.  K(cid:173)\nmeans was then used independently on each of the  states  and their  assigned  frames.  This \nis  a  standard  initialization  method  in  the  acoustic  domain  (Rabiner  &  Bing-Hwang, \n1993).  Since  K-means  can  be  trapped  in  local  minima,  the  algorithm  was  repeated  20 \ntimes with different starting point  and the  best  solution  was  fed  as  the  starting  point  for \nthe E-M algorithm. \n\n5  RESULTS \nThe main purpose of this  study  was  to  fmd  simple  image  preprocessing techniques that \nwould work well with hidden Markov  models.  We  tested  a wide  variety  of architectures \nand  preprocessing  parameters.  In  all  cases  the  results  were  evaluated  in  terms  cf \ngeneralization  to  new  speakers.  Since  the  training  sample  is  small,  generalization \nperformance  was  estimated using  the jackknife  procedure.  Models  were  trained  with  11 \n\n\fVisual  Speech  Recognition  with  Stochastic  Networks \n\n855 \n\nsubjects,  leaving  one  subject  out  for  generalization  testing.  The  entire  procedure  was \nrepeated 12  times,  each time  leaving a  different  subject out  for  testing.  Results  are  thus \nbased on 96 generalization trials (4  digits x  12 subjects  x  2  observations per  subject).  In \nall cases  we tested  several  preprocessing techniques  using  20  different  architectures  with \ndifferent number of states  (1,3,5,7,9)  and mixtures  per  state  (1,3,5,7).  To  compare  the \neffect of each processing technique  we used the  average  generalization performance  of the \nbest 4 architectures out of the 20 architectures tested. \n\n85 \n\n65 \n\n45 \n\n25  +---L. __ \n\nRho  Sigma Delta+Sigma \n\nFigure 2: Average performance with the rho, sigma, and delta images. \n\nFigure  2  shows  the  effects  of  symmetry  enforcement  and  temporal  differentiation. \nSymmetry enforcement had the benefit of reducing the  input  dimensionality  by  half and, \nas  the  figure  show  it  did  not  hinder recognition  performance.  Using  delta  images  had  a \nvery positive effect on recognition performance, as the  figure  shows.  Figure  3  shows  the \neffect of  varying the  thresholding  constant  and the  standard deviation  of the  Gaussian \nfilters.  Best  performance was  obtained with  blurring  windows  about 4  pixel  wide  and \nwith thresholding just about histogram equalization. \n\n84 \n83 \n82 \n81 \n80 \n\n---'-K=i.2 \n---K=O.6 \nK=O.3 \n\n3  4 \n\n5 \n\nStamdard Deviation of \n\nGaussian Filter. \n\nFigure 3: Effect of blurring and sharpening. \n\n\f856 \n\nJavier  Movellan \n\nTable 2 shows the effects of variations in the number of states  (S)  and  Gaussian  mixtures \n(G) per state.  The  number within  each cell  is  the  percentage  of simulations  for  which  a \nparticular combination  of states  and mixtures  performed best  out  of the  20  architectures \ntested. \n\nTable 2:  Effect of varying the the number of states (S) and Gaussian mixtures (G). \n\nGl \n\n0.00% \n0.00% \n3.12% \n6.25% \n6.25% \n\nG3 \n\n0.00% \n21.87% \n9.37% \n12.5% \n0.00% \n\nG5 \n\n0.00% \n12.5% \n15.62% \n0.00% \n3.12% \n\nG7 \n\n0.00% \n6.25% \n0.00% \n3.12% \n0.00% \n\nSI \nS3 \nS5 \nS7 \nS9 \n\nBest  overall performance  was  obtained with  about  3  states  and  3  mixtures  per  state. \nPeak performance was also obtained with  a  3-state,  3-mixture per state  network,  with  a \ngeneralization rate of 89.58% correct. \nTo  compare these results  with  human  performance,  9  subjects  were  tested  on  the  same \nsample.  Six  subjects  were  normal  hearing  adults  who  were  not  trained  in  lip-reading. \nThree were hearing impaired with profound hearing loss and had received  training  in  lip \nreading at 2 to 8 years of age. The mean correct  response  for  normal  subjects  was  89.93 \n% correct, just about the  same  rate  as  the  best artificial  network.  The  hearing  impaired \nhad an average  performance of 95.49% correct, significantly better than our  network. \n\nTable 3:  Confusion matrix of the best artificial system. \n\n\"One\" \n\"Two\" \n\"Three\" \n\"Four\" \n\n1 \n\n100.00% \n4.17% \n12.5% \n8.33% \n\n2 \n\n0.00% \n87.50% \n0.00% \n4.17% \n\n3 \n\n0.00% \n4.17% \n83.33% \n0.00% \n\n4 \n\n0.00% \n4.17% \n4.17% \n87.50% \n\nTable 4: Average human confusion matrix. \n\n\"One\" \n\"Two\" \n\"Three\" \n\"Four\" \n\n1 \n\n89.36% \n1.39% \n9.25% \n4.17% \n\n2 \n\n0.46% \n98.61% \n3.24% \n0.46% \n\n3 \n\n8.33% \n0.00% \n85.64% \n1.85% \n\n4 \n\n1.85% \n0.00% \n1.87% \n93.52% \n\nthe  average \nTables  3  and  4  show  the  confusion  matrices  for  the  best  network  and \nconfusion  matrix  with  all  9  subjects  combined.  The \ncorrelation  between  these  two \nmatrices  was  0.99.  This  means  that  98%  of the  variance  in  human  confusions  can  be \naccounted  for  by  the  artificial  model.  This  suggests  that  the  representational  space \nlearned by the artificial system may be a reasonable  model  of  the  representational  space \nused by humans. Figure 5  shows the representations  learned by  a network with  6  states \nand 1 mixture per  state. Each column is a  different  digit,  starting  with  \"one.\"  Each row \n\n\fVisual  Speech  Recognition  with  Stochastic  Networks \n\n857 \n\nis a different temporal state. The two pictures within each  cell  are  sigma  and  delta  image \nidentity  of individual  exemplars  is  lost  but  the \ncentroids.  As  the  figure  shows,  the \nunderlying  dynamics  of the  digits  are  preserved.  The  digits  can be  easily \nrecognized \nwhen played as a movie. \n\nFigure 4: Dynamic representations learned by a simple network. \n\n6  CONCLUSIONS \n\nThis  paper shows  that  simple  stochastic  networks,  like  hidden  Markov models,  can be \nsuccessfully  applied for visual  speech  recognition  using  relatively  unprocessed  images. \nThe  performance  level  obtained with  these networks  roughly  matches  untrained  human \nperformance.  Moreover,  the  representational  space  learned  by  these  networks  may  be  a \nreasonable model of the representations used by humans. More research should be done to \nbetter  understand  how  humans  integrate  visual  and  acoustic  information  in  speech \nperception and to develop practical models for robust audio-visual speech recognition. \n\nReferences \n\nBregler c., Manke S.,  Hild H.  &  Waibel  A.  (1993) Bimodal  Sensor  Integration  on  the \nExample of \"Speech-Reading\". Proc ICNN-93,  11,667-677. \n\nKuhl P. &  Meltzoff  A (1982) The Bimodal  Development of Speech  in  Infancy. Science, \n218,  1138-1141. \n\nMacLeod A. &  Summerfield Q. (1990) A  Procedure  for  Measuring Auditory  and Audio(cid:173)\nvisual  Speech-Reception  Measuring  Thresholds  for  Sentences  in  Noise:  Rationale, \nEvaluation and Recommendations for Use. British Journal of Audiology. 24,29-43. \n\nMassaro D.  (1987) Speech Perception by Ear and Eye.  In  Dodd B. &  Campbell  R.  (ed.) \nHearing by Eye:  The Psychology of Lip-Reading. London, LEA, 53-83. \n\nMassaro D.,  Cohen  M  &  Getsi  (1993) Long-Term  Training,  Transfer and  Retention  in \nLearning to Lip-read. Perception and Psychophysics,  53,549-562. \n\n\f858 \n\nJavier  Movellan \n\nMcGurk H. &  MacDonald J.  (1976) Hearing Lips and Seeing Voices. Nature, 264,  126-\n130. \n\nMontgomery  A.  &  Jackson P.  (1983)  Physical  Characteristics  of the  Lips  Underlying \nVowel Lipreading Performance. Journal  of the Acoustical Society of America,  73,2134-\n2144. \n\nNishida S.  (1986) Speech Recognition Enhancement by Lip  Information  Proceedings  of \nACMlCHI86,  198-204. \n\nPetajan E.  (1985) Automatic Lip  Reading to  Enhance  Speech Recognition.  IEEE  CVPR \n85,  40-47. \n\nRabiner  L.,  Bing-Hwang J.  (1993) Fundamentals  of Speech Recognition.  New  Jersey, \nPrentice Hall. \n\nSpelke  E.  (1976) \n533-560. \n\nInfant's  Intermodal  Perception  of Events.  Cognitive  Psychology,  8, \n\nYuhas  B.,  Goldstein  T.,  Sejnowski  T.,  Jenkins  R.  (1988)  Neural  Network  Models  cf \nSensory  Integration  for  Improved  Vowel  Recognition.  Proceedings \nIEEE  78,  1655-\n1668. \n\nWolff G.,  Prassad L.,  Stork  D.,  Hennecke M.  (1994)  Lipreading  by  Neural  Networks: \nVisual  Preprocessing,  Learning  and  Sensory  Integration.  In  J.  Cowan,  G.  Tesauro,  J. \nAlspector (ed.), Advances in Neural Information Processing Systems 6,  1027-1035.  San \nMateo, CA: Morgan Kaufinann. \n\n\f", "award": [], "sourceid": 993, "authors": [{"given_name": "Javier", "family_name": "Movellan", "institution": null}]}