{"title": "Word Space", "book": "Advances in Neural Information Processing Systems", "page_first": 895, "page_last": 902, "abstract": null, "full_text": "Word  Space \n\nCenter for  the Study of Language and  Information \n\nHinrich Schiitze \n\nVentura Hall \n\nStanford, CA 94305-4115 \n\nAbstract \n\nRepresentations  for  semantic  information about  words  are  neces(cid:173)\nsary  for  many applications of neural  networks  in natural language \nprocessing.  This paper describes  an efficient,  corpus-based method \nfor  inducing distributed semantic  representations  for  a  large num(cid:173)\nber  of words  (50,000)  from  lexical  coccurrence  statistics by means \nof a  large-scale  linear  regression.  The  representations  are  success(cid:173)\nfully applied to word sense disambiguation using a nearest neighbor \nmethod . \n\n1 \n\nIntroduction \n\nMany  tasks  in  natural  language processing  require  access  to  semantic  information \nabout lexical items and text segments.  For example, a system processing  the sound \nsequence:  /rE.k~maisbi:tJ/ needs to know the topic of the discourse in order to decide \nwhich  of the  plausible hypotheses  for  analysis  is  the  right  one:  e.g.  \"wreck  a  nice \nbeach\"  or  \"recognize  speech\" .  Similarly,  a  mail filtering  program has  to  know  the \ntopical significance of words  to  do its job properly. \n\nTraditional semantic representations are ill-suited for artificial neural networks since \nthey  presume  a  varying  number of elements  in  representations  for  different  words \nwhich  is  incompatible with  a  fixed  input  window.  Their  localist  nature  also  poses \nproblems  because  semantic  similarity  (for  example  between  dog  and  cat)  may  be \nhidden  in  inheritance  hierarchies  and  complicated feature  structures.  Neural  net(cid:173)\nworks  perform  best  when  similarity of targets  corresponds  to similarity of inputs; \ntraditional symbolic representations  do not have this  property.  Microfeatures  have \nbeen  widely  used  to  overcome  these  problems.  However,  microfeature  representa-\n\n895 \n\n\f896 \n\nSchutze \n\ntions  have  to be  encoded  by  hand  and  don't scale  up  to large vocabularies. \n\nThis paper presents an efficient method for deriving vector representations for words \nfrom  lexical cooccurrence  counts  in  a  large text corpus.  Proximity of vectors  in  the \nspace  (measured by the  normalized correlation coefficient)  corresponds  to semantic \nsimilarity.  Lexical  coocurrence  can  be  easily  measured.  However,  for  a  vocabulary \nof 50,000  words,  there  are  2,500,000,000 possible  coo currence  counts  to keep  track \nof.  While many of these  are  zero,  the number of non-zero  counts  is  still huge.  On \nthe  other  hand ,  in  any  document  collection  most  of  these  counts  are  small  and \ntherefore  unreliable.  Therefore,  letter fourgrams are  used  here  to bootstrap  the \nrepresentations.  Cooccurrence  statistics  are  collected  for  5,000 selected  fourgrams. \nSince  each  of the  5000  fourgrams  is  frequent,  counts  are  more  reliable  than  cooc(cid:173)\ncurrence  counts  for  rare  words.  The  5000-by-5000  matrix used  for  this  purpose  is \nmanageable.  A vector for  a lexical item is  computed as  the sum of fourgram vectors \nthat occur  close  to it  in  the  text .  This process  of confusion yields  representations \nof words  that  are  fine-grained  enough  to  reflect  semantic  differences  between  the \nvarious  case  and inflectional forms  a  word  may have  in  the  corpus. \n\nThe  paper  is  organized  as  follows.  Section  2  discusses  related  work.  Section  3 \ndescribes  the  derivation  of the  vector  representations .  Section  4  performs  an  eval(cid:173)\nuation.  The final  section concludes. \n\n2  Related Work \n\nTwo kinds of semantic representations  commonly used  in connectionism are  micro(cid:173)\nfeatures  (e.g . \\Valtz  and Pollack  1985,  McClelland and  Kawamoto 1986)  and local(cid:173)\nist  schemes  in  which  there  is  a  separate  node  for  each  word  (e.g .  Cottrell  1989). \nNeither  approach scales  up  well enough in its original form to be applicable to large \nvocabularies  and  a  wide  variety  of  topics.  Gallant (1991),  Gallant et a1.  (1992) \npresent  a  less  labor-intensive  method based  on  microfeatures,  but  the  features  for \ncore  stems still have  to be  encoded  by hand for each  new  document collection.  The \nderivation  of the  Word  Space  presented  here  is  fully  automatic.  It also  uses  fea(cid:173)\nture  vectors  to  represent  words,  but  the  features  cannot  be  interpreted  on  their \nown.  Vector  similarity is  the only information present  in  Word Space:  semantically \nrelated  words  are  close,  unrelated  words  are  distant.  The  emphasis  on  seman(cid:173)\ntic  similarity  rather  than  decomposition  into  interpretable  features  is  similar  to \nKawamoto (1988) .  Scholtes  (1991)  uses  a  two-dimensional  Kohonen  map  to  rep(cid:173)\nresent  semantic  similarity.  While  a  Kohonen  map  can  deal  with  non-linea.rities \n(in  contrast  to  the  singular  value  decomposition  used  below),  a  space  of  much \nhigher  dimensionality  is  likely  to  capture  more  of the  complexity  of semantic  re(cid:173)\nlatedness  present  in  natural  language.  Scholtes '  idea  to  use  n-gl'ams  to  reduce \nthe  number of initial features  for  the semantic representations  is  extended  here  by \nlooking  at  n-gram (oocurrence  statistics  rather  than  occurrence  in  documents  (cf. \n(Kimbrell  1988)  for  the  use  of n-grams in information retrieval). \n\nAn  important goal of many schemes of semantic represent.ation  is  to find  a  limited \nnumber  of semantic  classes  (e.g .  classical  thesauri  such  as  Roget's ,  Crouch  1990, \nBrown  et  a1.  1990).  Instead, a multidimensional space is constructed here,  in which \neach  word  has  its own individual representation.  Any  clustering  into  classes  intro(cid:173)\nduces  artificial  boundaries  that cut off words from part of their semantic neighbol'-\n\n\fWord  Space \n\n897 \n\ngovernor  quits  knights  of  columbus  over  bishop's  abortion gag  rule \nGOVE \n\nNIGH \n\nHTS \n\nOLUM \nLUMB \n\n_QUI \nQUIT \n\nVERN \nERNO \nRNOR \n\nSHOP \nHOP \n\nABOR \nBORT \nORTI \nRTIO \n\nRUL \nRULE \nULE_ \n\nFigure  1:  A  line from  the  New  York  Times with selected  fourgrams. \n\nhood.  In large classes,  there will be members \"from opposite sides of the class\"  that \nare  only  distantly  related.  So  any  class  size  is  problematic, since  words  are  either \nseparated from close  neighbors or lumped together  with distant terms.  Conversely, \na  multidimensional space  does  not  make such  an  arbitrary  classification necessary. \n\n3  Derivation of the Vector Representations \n\nFourgram selection.  There  are  about.  600,000  possible  fourgrams  if the  empty \nspace,  numbers  and  non-alphanumeric characters  are  included  as  \"special letters\" . \nOf these,  95,000 occurred  in  5  months of the  New  York  Times.  They  were  reduced \nto 5000 by first  deleting all rare ones (frequency less  than 1000) and then redundant \nand  uninformative fourgrams as  described  below. \nIf there is  a group of fourgrams tha.t occurs in only one word, all but.  one is  delet.ed. \nFor  instance,  the  fourgrams  BAGH,  AGHD,  GHDA,  HDAD  tend  to  occur  together  in \nBaghdad,  so  three  of  them  will  be  deleted.  The  rationale  for  this  move  is  that \ncooccurrence  information  about  one  of  the  fourgrams  can  be  fully  derived  from \neach  of the  others,  so  that  an  index  in  the  matrix  would  be  wasted  if more  than \none of them was  included.  The  relative frequency  of one  fourgram  occurring  after \nanother was calculated with fivegrams.  For instance,  the relative frequency  of AGHD \nfollowing BAGH  is  the frequency  of the fivegram  BAGHD  divided  by  the frequency  of \nthe fourgram  BAGH. \n\nMost  fourgrams  occur  predominantly  in  three  or  four  stems  or  words.  U ninfor(cid:173)\nmative  fourgrams  are  sequences  such  as  RET!  or  TION  that  are  part  of so  many \ndifferent  words  (resigned,  residents,  retirements,  resisted,  . .. ;  abortion,  despera(cid:173)\ntion,  construction,  detention,  ... )  that  knowledge  about  coocurrence  with  them \ncarries  almost  no  semantic  information.  Such  fourgrams  are  therefore  useless  and \nare  deleted.  Again,  fivegrams  were  used  to  identify  fourgrams  that  occurred  fre(cid:173)\nquently  in many stems. \n\nA set of 6290 fourgrams remained after these  deletions.  To reduce  it to the required \nsize  of 5000,  t.he  most  frequent  300  and  the  least  frequent.  990  were  also  delet.ed. \nFigure  1  shows  a  line  from  the  New  York  Times  and  which  of the  5000  selecteo \nfourgrams occurred  in it. \n\nComputation  of  fourgram  vectors.  The  computation  of  word  vectors  de(cid:173)\nscribed  below  depends  on  fourgram  vectors  that  accurately  reflect  semantic  sim(cid:173)\nilarity in  the sense of being  used  to  describe  the same contents.  Consequently,  one \nneeds  to  be  able  to  compare  the sets  of contexts  two  fourgrams occur  in.  For  this \npurpose,  a collocation matrix for fourgrams was collected such that the entry ai,j \n\n\f898 \n\nSchiitze \n\ncounts  the  number  of times that  fourgram  i  occurs  at  most  200  fourgrams  to  the \nleft of fourgram j.  Two columns in this matrix are similar if the contexts the corre(cid:173)\nsponding fourgrams are  used  in are similar.  The counts were  determined  using five \nmonths of the  New  York  Times  (June - October  1990).  The  resulting  collocation \nmatrix is  dense:  only  2%  of entries  are  zeros,  because  almost  any  two  fourgrams \ncooccur.  Only  10%  of  entries  are  smaller  than  10,  so  that  culling  small  counts \nwould  not  increase  the  sparseness  of the  matrix.  Consequently,  any  computation \nthat  employs the fourgram vectors  directly  would  be inefficient.  For  this  reason,  a \nsingular  value  decomposition  was  performed  and  97  singular values extracted  (cf. \nDeerwester  et  al.  1990)  using  an  algorithm  from  SVDPACK  (Berry  1992).  Each \nfourgram  can  then  be  represented  by  a  vector  of 97  real  values.  Since  the singular \nvalue decomposition finds  the best least-square approximation of the original space \nin  97  dimensions,  two  fourgram  vectors  will  be  similar if their  original  vectors  in \nthe  collocation matrix are similar.  The reduced  fourgram  vectors  can  be efficiently \nused  for  confusion as described  in  the following section. \n\nComputation of word vectors.  We can think of fourgrams as highly ambiguous \nterms.  Therefore,  they are inadequate if used  directly  as input  to a  neural net.  We \nhave  to  get  back  from  fourgrams  to  words.  For  the  experiment  reported  here, \ncooccurrence  information was  used  for  a  second  time to  achieve  this  goal:  in  this \ncase  coo currence  of a  target  word  with  any  of the  5000  fourgrams.  For  each  of \nthe  selected  words  (see  below),  a  context  vector  was  computed for  every  position \nat  which  it  occurred  in  the  text.  A  context  vector  was  defined  as  the  sum  of all \ndefined fourgram vectors  in a window of 1001 fourgrams centered  around the target \nword.  The context  vectors  were then normalized and summed.  This sum of vectors \nis  the  vector  representation  of the  target  word.  It is  the  confusion of all  its  uses \nin  the  corpus.  More formally,  if C( w)  is  the set  of positions in the  corpus  at which \nw  occurs  and  if 'P(f)  is  the  vector  representation  for  fourgram  f,  then  the  vector \nrepresentation  r( w)  of w  is  defined  as:  (the  dot stands for  normalization) \n\n\u2022 \nr(w) =  L  (  L \n\ni\u20acC(w)  J  close  to  i \n\n'P(f)) \n\nThe  treatment of words  is  case-sensitive.  The following  terminology  will  be  used: \na  surface form  is the string of characters  as  it occurs in  the text;  a  lemma  is either \nlower  case  or  upper  case:  all  letters  are  lower  case  with  the  possible  exception  of \nthe  first;  word  is  used  as  a  case-insensitive  term.  So  every  word  has  exactly  two \nlemmas.  A  lemma of length n  has  up  to 2n  surface forms.  Almost every  lower  case \nlemma can  be  realized  as  an  upper  case  surface form.  But  upper  case  lemmas are \nhardly ever  realized  as  lower  case  surface forms. \n\nThe  confusion  vectors  were  computed  for  all  54366  lemmas that occurred  at  least \n10  times in  18  months of the  New  York  Times News  Service  (May  1989 - October \n1990, about 50  million words).  Table 1 lists the percentage of lower  case  and upper \ncase  lemmas, and  the  distribution of lemmas with  respect  to words. \n\n\fWord  Space \n\n899 \n\nlemmas \nlower  ca.se \nupper  case \ntotal \n\nnumber \n32549 \n21817 \n54366 \n\npercent \n60  0 \n40% \n100  0 \n\nwords \nlower  case lemma only \nupper  case lemma only \nboth lemmas \ntotal \n\nnumber \n23766 \n13034 \n8783 \n45583 \n\npercent \n52  0 \n29% \n19% \n100  0 \n\nTable  1:  The distribution of lower  and  upper  case  in words  and  lemmas. \n\nword \nburglar \ndisable \ndisenchantment \ndomestically \nDour \ngrunts \nkid \nS.O.B. \nSte. \nworkforce \nkeepmg \n\nI nearest  neighbors \nburglars  thief rob  mugging  stray  robbing  lookout  chase  C) ate  thieves \ndeter intercept  repel  halting  surveillance  shield  maneuvers \ndisenchanted  sentiment  resentment  grudging  mindful  unenthusiastic \ndomestic  auto/-s importers/-ed  threefold  inventories  drastically  cars \nmelodies/-dic  Jazzie  danceable  reggae  synthesizers  Soul  funk  tunes \nheap  into  ragged  goose  neatly  pulls  buzzing  rake  odd rough \ndad  kidding  mom  ok  buddies  Mom Oh  Hey  hey  mama \nConfessions  Jill  Julie  biography  Judith  Novak  Lois  Learned  Pulitzer \ndry  oyster  whisky  hot  filling  rolls  lean  float  bottle ice \njobs employ /-s/-ed/-ing  attrition  workers  clerical  labor  hourly \nI hopmg  brmg  wlpmg  could  some  would  other here rest  have \n\n.. \n\nTable 2:  Ten  random and  one selected  word  and their  nearest  neighbors. \n\n4  Evaluation \n\nTable 2 shows a  random sample of 10 words and their ten nearest neighbors in Word \nSpace  (or  less  depending  on  how  many would  fit  in  the  table).  The  neighbors  are \nlisted  in  order  of proximity to  the  head  word.  burglar,  disenchantment,  kid,  and \nworkforce  are  closely  related  to  almost  all  of their  nearest  neighbors.  The same is \ntrue for  disable,  dom esticaUy,  and  Dour,  if we  regard  as  the goal to  come up  with \na  characterization  of semantic  similarity in  a  corpus  (as  opposed  to  the  language \nin  general).  In  the  New  York  Times,  the  military use  of disable  dominates,  Iraq's \nmilitary, oil  pipelines  and ships  are  disabled.  Similarly,  domestic  usually  refers  to \nthe domestic market, and only one person named Dour occurs in the newspaper:  the \nSenegalese jazz musician Youssou N'Dour.  So these three cases  can also be  counted \nas successes.  The topic/ content of grunts  is  moderately well  characterized  by other \nobjects  like  goose  and  rake  that  one  would  also  expect  on  a  farm.  Finally,  little \nuseful  information can  be  extracted  for  S. D.B.  and  Ste.  S. D.B.  mainly  occurs  in \narticles about.  the bestseller  \"Confessions of an S.O.B.\"  Since it is not.  used literally, \nits semantics  don't come out very  well.  The neighbors of Ste  are for  the  most part \nwords  associated  with  water,  because  the  name  of the  river  \"Ste.-Marguerite\"  in \nQuebec  (popular  for  salmon  fishing)  is  the  most  frequent  context  for  Ste.  Since \nthe significance of Ste  depends  heavily on  the name it occurs  in,  its  usefulness  a.s  a. \ncontributor of semantic informa.tion is  limited, so  its  poor  characterization  should \nprobably not be seen as problematic.  The word keeping  has been added to the table \nto show  that the vector  representations  of words  that can be used  in a  wide variety \nof contexts  are  not.  interesting. \n\nTable  3  shows  that it is  important for  many  words  to  make a  distinction  between \n\n\f900 \n\nSchiitze \n\nword \npinch  (.41) \nPinch \nkappa  (.49) \nKappa \nroe  (.54) \nRoe \ncompletion  (.73) \ncompletions \nok  (.60) \noks \ntriad  (.52) \ntriads \n\nnearest  neighbors \nouts pitch  Cone  hitting  Cary  strikeout  Whitehurst  Teufel  Dykstra mound \nunsalted  grated cloves  pepper  teaspoons  coarsely  parsley  Combine  cumin \ncasein  protein/-s  synthesize liposomes  recombinant  enzymes  amino  dna \nPhi Wesleyan  graduate cum dean  graduating  nyu  Amherst  College  Yale \ncod  squid  fish  salmon  flounder  lobster  haddock lobsters  crab chilled \nWade  v overturn/-ing  uphold/-ing  abortion  Reproductive  overrule \ncomplete/-~/-s/-ing complex  phase/-s  uncompleted  incomplete \ntouchdown/-s interception/-s  td yardage yarder  tds fumble  sacked \nd  me  I  m wouldn  t  crazy  you  ain  anymore \napprove/-s/-d/-ing  Senate Waxman  bill  appropriations  omnibus \nwarhea~/-s ballistic  missile[-s  ss  bombers intercontinental  silos \nTriads  Organized  Interpol  Cosa Crips gangs  trafficking  smuggling \n\nTable 3:  Words for  which  case  or inflection matter. \n\nword \n\nI senses \n\ngoodsLseat  of government \nspecial  attention/financial \n\ncap ita ljs \ninterestjs \nmotionjs  movement/proposal \nfactory /living  being \nplantjs \ndecision/to  exert control \n\"uling \nspace \narea,  volume/outer  space \nlegal  action/garments \nsuitjs \ncombat  vehicle/receptacle \ntankjs \nrailroad  cars/to  teach \ntrainjs \nship/blood  vessel/hollow  utensil \nvesse1js \n\n96 \n94 \n92 \n94 \n90 \n89 \n94 \n97 \n94 \n93 \n\n3 \n\n% correct \n2 \n92 \n92 \n91 \n88 \n91 \n90 \n95 \n85 \n69 \n91 \n\n86 \n\nsum \n95 \n93 \n92 \n92 \n90 \n90 \n95 \n95 \n89 \n92 \n\nTable 4:  Ten  disambiguation experiments using  the vector  representations. \n\nlower  case  and  upper  case  and  between  different  inflections.  The  normalized  cor(cid:173)\nrelation  coefficient  between  the two  case/inflectional forms of the word is indicated \nin each  example. \n\nWord  sense disambiguation.  Word  sense  disambiguation is  a  task  that  many \nsemantic  phenomena  bear  on  and  therefore  well  suited  to  evaluate  the  quality  of \nsemantic  representations.  One  can  use  the  vector  representations  for  disambigua(cid:173)\ntion  in  the  following  way.  The  context  vector  of the  occurrence  of an  ambiguous \nword  is  defined  as  the  sum  of all  word  vectors  ocurring  in  a  window  around  it . \nThe  set  of context  vectors  of the  word  in  the  training  set  can  be  clustered.  The \nclustering  programs  used  were  AutoClass  (Cheeseman et al.  1988)  and  Buckshot \n(Cutting et  al.  1992).  The  clusters  found  (between  2  and  13)  were  assigned  senses \nby  inspecting a  few  of its  members  (10-20) .  An  occurrence  of an  ambiguous word \nin the test set  was then  disambiguated by assigning the sense of the training cluster \nthat  was  closest  to  its  context  vector.  Note  that  this  method  is  unsupervised  in \nthat the structure of the  \"sense space\"  is analyzed automatically by clustering.  See \nSchiitze  (1992)  for  a  more detailed  description . \n\nTable  4  lists  the  results  for  ten  disambiguation experiments  that  were  performed \n\n\fWord  Space \n\n901 \n\nusing  the above algorithm.  Each  line shows  the ambiguous words,  its major senses \nand the success  rate of disambiguation for  the individual senses and all major senses \ntogether.  Training  and  test  sets  were  taken  from  the  New  York  Times  newswire \nand were  disjoint for  each  word.  These  disambiguation results  are  among the  best \nreported in  the literature (e.g.  Yarowsky  1992).  Apparently,  the  vector  representa(cid:173)\ntions  respect  fine  sense  distinctions. \n\nAn interesting question is to what degree the vector representations are  distributed. \nUsing  the  algorithm for  disambiguation described  above,  a  set  of contexts  of suit \nwas  clustered  and  applied  to  a  test  text.  When  the  first  30  dimensions  were  used \nfor  clustering  the training set,  the error rate was 9%  in  the test set.  \\\\Then  only the \nodd dimensions were  used  (1,3,5, ... ,27,29)  the error  was  14%.  With only  the even \ndimensions  (2,4,6, ... ,28,30),  13%  of occurrences  in  the  test  set  were  misclassified. \nThis graceful  degradation indicates that the vector  representations  are  distributed. \n\n5  Discussion  and  Conclusion \n\nThe linear dimensionality reduction performed here  could be a  useful  preprocessing \nstep  for  other  applications  as  well .  Each  of the  fourgram  features  carries  a  small \namount of information.  Neglecting  individual features  degrades  performance,  but \nthere  are  so  many that they  cannot  be used  directly  as  input  to  a  neural  network. \nThe  word  sense  disambiguation  results  suggest  that  no  information  is  lost  when \nonly axes of variations extracted by the singular value decomposition are considered \ninstead  of the  original 5000-dimensional fourgram  vectors.  Schiitze  (Forthcoming) \nuses the same methodology for the derivation of syntactic representations for words \n(so that verbs and nouns occupy different regions in syntactic word space).  Problems \nin  pattern  recognition  often  have  the same characteristics:  uniform distribution  of \ninformation  over  all  input  features  or  pixels  and  a  high-dimensional  input  space \nthat  causes  problems in  training if the features  are  used  directly.  A  singular  value \ndecomposition  could  be  a  useful  preprocessing  step  for  data  of this  nature  that \nmakes neural nets applicable to high-dimensional problems for which training would \notherwise  be slow  if possible  at all. \n\nThis  paper  presents  Word  Space,  a  new  approach  to  representing  semantic infor(cid:173)\nmation  about  words  derived  from  lexical  cooccurrence  statistics.  In  contrast  to \nmicrofeature  representations,  these  semantic representations  can  be  summed for  a \ngiven  context  to  compute  a  representation  of the  topic  of a  text  segment.  It was \nshown  that semantically  related  words  are  close  in  Word  Space  and  that  the  vec(cid:173)\ntor  representations  can  be  used  for  word  sense  disambiguation.  Word  Space  could \ntherefore be a promising input representation for  applications of neural nets in natu(cid:173)\nrallanguage processing such as information filtering or language modeling in speech \nrecognition. \n\nAcknowledgements \n\nI'm  indebted  to  Mike  Berry  for  SVDPACK,  to  NASA  and  RIACS  for  AutoClass \nand  to  the  San  Diego  Supercomputer  Center  for  computing resources.  Thanks  to \nMartin Kay, Julian Kupiec, Jan Pedersen,  Martin Roscheisen, and Andreas Weigend \nfor  help  and  discussions. \n\n\f902 \n\nSchutze \n\nReferences \n\nBerry,  M.  W.  1992.  Large-scale sparse  singular  value  computations.  The  Interna(cid:173)\ntional  Journal  of Supercomputer Applications 6(1):13-49. \nBrown,  P.  F.,  V.  J.  D.  Pietra,  P.  V.  deSouza,  J.  C.  Lai,  and  R.  L.  Mercer.  1990. \nClass-based  n-gram models of natural language.  Manuscript,  IBM. \nCheeseman,  P.,  J.  Kelly,  M.  Self,  J.  Stutz,  W.  Taylor,  and  D.  Freeman.  1988.  Au(cid:173)\ntoClass:  A  Bayesian  classification system.  In  Proceedings  of the  Fifth  International \nConference  on  Machine  Learning. \nCottrell,  G.  ,,y.  1989.  A  Connectionist  Approach  to  Word  Sense  Disambiguation. \nLondon:  Pitman. \nCrouch,  C.  J.  1990.  An  approach  to the automatic construction of global thesauri. \nInformation  Processing  {3  Management  26(5):629-640. \nCutting,  D.,  D.  Karger, J.  Pedersen,  and J.  Thkey.  1992.  Scatter-gather:  A  cluster(cid:173)\nbased approach to browsing large document collections. In Proceedings  of SIGIR '92. \nDeerwester,  S.,  S.  T.  Dumais,  G.  W.  Furnas,  T.  K.  Landauer,  and  R.  Harshman. \n1990.  Indexing  by  latent  semantic  analysis.  Journal  of the  American  Society  for \nInformation  Science 41(6):391-407. \nGallant, S.  I.  1991. A practical approach for representing context and for performing \nword  sense  disambiguation using  neural  networks.  Neural  Computation  3(3) :293-\n309. \nGallant, S.  I.,  W.  R.  Caid, J.  Carleton,  R.  Hecht-Nielsen,  K.  P.  Qing,  and  D.  Sud(cid:173)\nbeck.  1992.  HNC's matchplus system.  In  Proceedings  of TREC. \nKawamoto, A.  H.  1988.  Distributed  representations  of ambiguous words  and  their \nresolution  in  a  connectionist  network.  In  S.  1.  Small,  G.  W.  Cottrell,  and  M.  K. \nTanenhaus  (Eds.),  Lexical  A mbiguity  Resolution:  Perspectives from  Psycholinguis(cid:173)\ntics,  Neuropsychology,  and  Artificial  Intelligence.  San  Mateo  CA:  Morgan  Kauf(cid:173)\nmann. \nKimbrell,  R.  E.  1988.  Searching  for  text?  Send  an  N-gram!  Byte  Magazine \nMay:297-312. \nMcClelland, J.  L.,  and A.  H.  Kawamoto.  1986.  Mechanisms of sentence  processing: \nAssigning roles  to constituents of sentences.  In J.  L.  McClelland,  D.  E.  Rumelhart, \nand the  PDP  Research  Group  (Eds.),  Parallel Distributed Processing.  Explorations \nin  the  Microstructure  of Cognition.  Volume  2:  Psychological  and  Biological M ode/s, \n272-325.  Cambridge MA:  The  MIT  Press. \nScholtes,  J.  C.  1991.  Unsupervised  learning and the information retrieval  problem. \nIn  Proceedings  of the  International  Joint  Conference  on  Neural  Networks. \nSchiitze,  H.  1992.  Dimensions of meaning.  In  Proceedings  of Supercomputing  '92. \nSchiitze,  H.  Forthcoming.  Sublexical  tagging.  In  Proceedings  of the  IEEE Interna(cid:173)\ntional  Conference  on  N euml Networks. \nWaltz,  D.  L.,  and  J.  B.  Pollack.  1985.  A  strongly  interactive  model  of natural \nlanguage interpretation.  Cognitive  Science  9:51-74. \nYarowsky,  D.  1992.  Word-sense  disambiguation using statistical models of Roget's \ncategories  trained on large corpora.  In  Proceedings  of Coling-92. \n\n\f", "award": [], "sourceid": 603, "authors": [{"given_name": "Hinrich", "family_name": "Sch\u00fctze", "institution": null}]}