{"title": "A Connectionist Learning Approach to Analyzing Linguistic Stress", "book": "Advances in Neural Information Processing Systems", "page_first": 225, "page_last": 232, "abstract": null, "full_text": "A Connectionist Learning Approach to Analyzing \n\nLinguistic Stress \n\nPrahlad Gupta \n\nDepartment of Psychology \nCarnegie Mellon University \n\nPittsburgh, PA  15213 \n\nDavid S. Touretzky \n\nSchool of Computer Science \nCarnegie Mellon University \n\nPittsburgh, PA  15213 \n\nAbstract \n\nWe use connectionist modeling to develop an analysis of stress systems in terms \nof ease  of learnability.  In traditional linguistic analyses,  learnability arguments \ndetermine default parameter settings based on the feasibilty of logicall y deducing \ncorrect settings from an initial state.  Our approach provides an empirical alter(cid:173)\nnative to such arguments.  Based on perceptron learning experiments using data \nfrom  nineteen  human  languages,  we  develop a  novel  characterization of stress \npatterns in terms of six parameters.  These provide both a partial description of the \nstress pattern itself and a prediction of its learnability, without invoking abstract \ntheoretical constructs  such  as  metrical  feet.  This  work  demonstrates  that ma(cid:173)\nchine learning methods can provide a fresh  approach to understanding linguistic \nphenomena. \n\n1  LINGUISTIC STRESS \n\nThe domain of stress systems in language is considered to have a relatively good linguistic \ntheory,  called metrical phonologyl.  In this  theory,  the stress  patterns of many  languages \ncan be described concisely, and characterized in terms of a set of linguistic \"parameters,\" \nsuch as  bounded vs.  unbounded metrical feet,  left vs.  right dominant feet,  etc.2  In many \nlanguages, stress tends to be placed on certain kinds of syllables rather than on others; the \nformer are termed heavy syllables, and the latter light syllables. Languages that distinguish \n\nlFor an overview of the theory, see [Goldsmith 90, chapter 4]. \n2See [Dresher 90]  for one such parameter scheme. \n\n225 \n\n\f226 \n\nGupta and Touretzky \n\nOUTPUT UNIT \n(PERCEPTRON) \n\nInput \n\nINPUT LAYER  (2 x 13  units) \n\nFigure 1:  Perceptron model used in simulations. \n\nbetween heavy and light syllables are termed quantity-sensitive (QS), while languages that \ndo not make this distinction are termed quantity-insensitive (QI). In some QS  languages, \nwhat counts as  a heavy syllable is  a closed syllable (a syllable that ends in a consonant), \nwhile in others  it is  a  syllable with  a  long vowel.  We  examined  the  stress  patterns  of \nnineteen QI and QS systems, summarized and exemplified in Table 1.  The data were drawn \nprimarily from descriptions in [Hayes 80]. \n\n2  PERCEPTRON SIMULATIONS \n\nIn separate experiments,  we trained a perceptron to  produce the stress  pattern of each of \nthese languages.  1\\\\10 input representations were used.  In the syllabic representation, used \nfor QI patterns only, a syllable was represented as  a [11] vector, and [00] represented no \nsyllable.  In the weight-string representation, which  was  necessary  for QS  languages,  the \ninput patterns used were [1  0]  for a heavy  syllable, [0  1]  for a light syllable, and [00] for \nno syllable.  For stress systems  with up to two levels of stress,  the output targets  used in \ntraining were 1.0 for primary stress, 0.5 for secondary stress, and 0 for no stress.  For stress \nsystems with three levels of stress, the output targets were 0.6 for secondary stress, 0.35 for \ntertiary stress, and 1.0 and 0 respectively for primary stress and no stress.  The input data set \nfor all stress systems consisted of all word-forms of up to seven syllables.  With the syllabic \ninput representation there are 7 of these, and with the weight -string representation, there are \n255.  The perceptron's input array  was a buffer of 13 syllables;  each  word was  processed \none syllable at a time by sliding it through the buffer (see Figure 1).  The desired output at \neach step was the stress level of the middle syllable of the buffer.  Connection weights were \nadjusted at each step using the back-propagation learning algorithm [Rumelhart 86].  One \nepoch consisted of one presentation of the entire training set.  The network was  trained for \nas many epochs as necessary to ensure that the stress value produced by the perceptron was \nwithin 0.1 of the target value, for each syllable of the word, for all words in the training set. \nA learning rate of 0.05 and momentum of 0.90 was used in all simulations.  Initial weights \nwere uniformly distributed random values in the range \u00b10.5.  Each  simulation was  run at \nleast three times, and the learning times averaged. \n\n\fConnectionist Learning and Linguistic Stress \n\n227 \n\nLANGUAGE \n\nDESCRIPTION OF S1RESS PATIERN \n\nREF \nQuantity-Insensitive Languages: \nLl \nL2 \nL3  Maranungku \n\nLatvian \nFrench \n\nFixed word-initial stress. \nFixed word-final stress. \nPrimary stress on first syllable, secondary stress on alternate \nsucceeding syllables. \nPrimary stress on last syllable, secondary stress on alternate \npreceding syllables. \nPrimary stress on first  syllable, secondary stress on penulti-\nmate syllable, tertiary stress on alternate syllables preceding \nthe penUlt, no stress on second syllable. \nPrimary stress On second syllable. \nPrimary stress on penultimate syllable. \nPrimary stress on second syllable, secondary stress on alter-\nnate succeeding syllables. \nPrimary stress on penultimate syllable.  secondary stress on \nalternate preceding syllables. \n\nEXAMPLES \n\nSlSOSOSOSOSoSo \nSOSoSoSOSOSOSI \nSlSOS2S0S2S0S2 \n\nS2S0S2S0S2S0S1 \n\nSlSOSOS3S0S2S0 \n\nSOSI SOSOSoSOSo \nSOSOSOSOSOSlSO \nSOSlSOS2SOS2S0 \n\nSOS2S0S2SOS1S0 \n\nQuantity-Sensitive Languages: \nLlO  Koya \n\nL4  Weri \n\nL5 \n\nGarawa \n\nL6 \nL7 \nL8 \n\nLakota \nSwahili \nPaiute \n\nL9  Warao \n\nL11 \n\nEskimo \n\nL12  Gurkhali \n\nL13  Yapese \n\nL14  Ossetic \n\nL15  Rotuman \n\nL16  Komi \n\nL17  Cheremis \n\nL18  Mongolian \n\nL19  Mayan \n\nPrimary  stress  on  first  syllable,  secondary stress  on heavy  LILoLoH2LoLoLo \nsyllables. \nLILoLoLoLoLoLo \n(Heavy = closed syllable or syllable with long vowel.) \n(Primary) stress on final and heavy syllables. \nLOLoLoHILoLoLl \n(Heavy = closed syllable.) \nLOLoLoLoLoLoLI \nPrimary stress on first syllable except when first syllable light  LILoL\u00b0J-fJLoLoLo \nL \u00b0 HI L 0J-fJ L \u00b0LoLo \nand second syllable heavy. \n(Heavy = long vowel.) \nPrimary stress on last syllable except when last is  light and  LOLoL\u00b0J-fJLoLoLl \nL\u00b0J-fJL \u00b0J-fJL \u00b0HIL \u00b0 \npenultimate heavy. \n(Heavy = long vowel.) \nPrimary stress on first syllable if heavy.  else on  second syl- HI L \u00b0 L \u00b0 J-fJ L \u00b0 L \u00b0 L \u00b0 \nlable. \nLOLILoLoLoLoLo \n(Heavy = long vowel.) \nPrimary stress on last syllable if heavy. else on penultimate  LOLoL\u00b0J-fJLoLoHl \nsyllable. \nLOLoLoLoLoLlLo \nJHeavy = long vowel.) \nPrimary  stress  on  first  heavy syllable.  or on last syllable if  L \u00b0 L \u00b0H1 L \u00b0L\u00b0J-fJL \u00b0 \nnone heavy. \n(Heavy = long vowel.) \nPrimary  stress on  last heavy syllable.  or on first  syllable if  LOL \u00b0J-fJLoLoH1L \u00b0 \nnone heavy. \nLILoLoLoLoLoLo \n(Heavy = long vowel.) \nPrimary stress on first  heavy syllable.  or on first  syllable if  L \u00b0L \u00b0H1LoL\u00b0J-fJL \u00b0 \nnone heavy. \nLILoLoLoLoLoLo \n(Heavy = long vowel.) \nPrimary  stress on last heavy syllable.  or on last syllable if  LOL\u00b0J-fJLoLoHlLo \nnone heavy. \nLOLoLoLoLoLoLI \n(Heavy = long vowel.) \n\nLOLoLoLoLoLoLl \n\nTable  1:  Stress  patterns:  description and  example  stress  assignment.  Examples  are  of \nstress  assignment in seven-syllable words.  Primary stress is  denoted by the superscript 1 \n(e.g., Sl), secondary stress by the superscript 2, tertiary stress by the superscript 3, and no \nstress by the superscript O.  \"S\" indicates an arbitrary syllable, and is used for the QI stress \npatterns.  For QS  stress patterns, \"H\" and \"L\" are used to denote Heavy and Light syllables, \nrespectivel y. \n\n\f228 \n\nGupta and Touretzky \n\n3  PRELIMINARY ANALYSIS OF LEARNABILITY OF STRESS \n\nThe  learning  times  differ  considerably  for  {Latvian,  French},  {Maranungku,  Weri} , \n{Lakota, Polish} and Garawa, as  shown in the last column of Table 2.  Moreover, Paiute \nand Warao  were unlearnable with this mode1.3  Differences  in learning times  for the var(cid:173)\nious  stress  patterns suggested that the factors  (\"parameters\")  listed below are relevant in \ndetermining learnability. \n\n1. Inconsistent Primary Stress (IPS):  it is computationally expensive to learn the pattern \nif neither edge receives  primary stress except in mono- and di-syllables;  this can  be \nregarded as an index of computational complexity that takes the values {O,  I}:  1 if an \nedge receives primary stress inconsistently, and 0, otherwise. \n\n2.  Stress clash avoidance (SeA):  if the  components  of a  stress  pattern  can  potentially \nlead to stress clash4, then the language may either actually permit such stress clash, or \nit may avoid it.  This index takes the values  {O,  I}:  0 if stress clash is permitted, and \n1 if stress clash is avoided. \n\n3. Alternation (AIt):  an  index of learnability with value 0 if there is  no  alternation, and \nvalue  1 if there  is.  Alternation refers  to  a  stress  pattern  that repeats  on alternate \nsyllables. \n\n4.  Multiple Primary Stresses (MPS):  has  value 0 if there is exactly one primary stress, \nand value 1 if there is more then one primary stress.  It has been assumed that a repeating \npattern of primary stresses  will be on alternate, rather than adjacent syllables.  Thus, \n[Alternation=O] implies [MPS=O].  Some of the hypothetical stress patterns examined \nbelow include ones  with more than one primary stress;  however,  as  far as  is known, \nno actually occurring QI stress pattern has more than one primary stress. \n\n5. Multiple Stress Levels (MSL):  has  value 0 if there is a single level of stress (primary \n\nstress only), and value 1 otherwise. \n\nNote that it is  possible to  order these factors  with  respect  to each  other  to form  a  five(cid:173)\ndigit binary string characterizing the ease/difficulty of learning.  That is, the computational \ncomplexity of learning a stress pattern can be characterized as a 5-bit binary number whose \nbits represent the five  factors  above,  in decreasing  order of significance.  Table 2  shows \nthat this characterization captures the learning times of the QI patterns quite accurately.  As \nan  example of how  to read  Thble 2,  note that Garawa takes  longer to learn  than  Latvian \n(165 vs.  17  epochs).  This is reflected in the parameter setting for Garawa, \"01101\", being \nlexicographically greater  than  that  for Latvian,  \"00000\".  A  further  noteworthy point is \nthat this framework provides an account of the non-learnability of Paiute and Warao,  viz,. \nthat stress  patterns  whose parameter string is  lexicographically greater than  \"10000\" are \nunlearnable by the perceptron. \n\n4  TESTING THE QI LEARNABILITY PREDICTIONS \n\nWe devised a series of thirty artificial QI stress patterns (each a variation on some language \nin  Table  1)  to  examine our parameter scheme in  more detail.  The details of the patterns \n\n3They were learnable in a three-layer model, which exhibited a similar ordering of learning times \n\n[Gupta 92]. \n\n4Placement of stress on adjacent syllables. \n\n\fConnectionist Learning and Linguistic Stress \n\n229 \n\nIPS  SCA  Alt  MPS  MSL  QI LANGUAGES  REF \n\n0 \n\n0 \n\n0 \n1 \n\n1 \n\n0 \n\n0 \n\n1 \n0 \n\n0 \n\n0 \n\n1 \n\n1 \n0 \n\n1 \n\n0 \n\n0 \n\n0 \n0 \n\n0 \n\n0 \n\n1 \n\n1 \n0 \n\n1 \n\nLatvian \nFrench \nMaranungku \nWeri \nGarawa \nLakota \nSwahili \nPaiute \nWarao \n\nLl \nL2 \nL3 \nL4 \nL5 \nL6 \nL7 \nL8 \nL9 \n\nEPOCHS \n(syllabic) \n17 \n16 \n37 \n34 \n165 \n255 \n254 \n** \n** \n\nTable  2:  Preliminary analysis  of learning times  for  QI stress  systems.  using  the syllabic \ninput  representation. \nIPS=Inconsistent  Primary  Stress;  SCA=Stress  Clash  Avoidance; \nAlt=Altemation;  MPS=Multiple Primary  Stresses;  MSL=Multiple Stress Levels.  Refer(cid:173)\nences LI-L9 refer to Table 1. \n\nII  Agg  I IPS \n\n0 \n\n0 \n0 \n0 \n\n0 \n0 \n\n0 \n\n0 \n\n0 \n\n1 \n\n2 \n\n0 \n\n0 \n0 \n0 \n\n0 \n0.25 \n\n0.50 \n\n1 \n\n1 \n\n0 \n\n0 \n\n0 \n0 \n0 \n\n1 \n0 \n\n0 \n\n0 \n\n0 \n\n0 \n\n0 \n\n0 \n0 \n1 \n\n1 \n0 \n\n0 \n\n0 \n\n1 \n\n0 \n\n0 \n\n0 \n1 \n0 \n\n0 \n0 \n\n0 \n\n0 \n\n0 \n\n0 \n\n0 \n\n1 \n0 \n1 \n\n1 \n0 \n\n0 \n\n0 \n\n1 \n\n0 \n\n0 \n\nSCA  I Alt  I MPS  I MSL  \"  QI LANGS \n0 \n\n0 \n\nLatvian \nFrench \n\n0 \n\n0 \n\nI REF  I TIME  \"  QS LANGS  I REF  I TIME  II \nLl \nL2 \n\n2 \n2 \n\nKoya \nEskimo \n\nLI0 \nLll \n\n2 \n3 \n\nMaranungku  L3 \nL4 \nWeri \nL5 \nGarawa \n\nLakota \nSwahili \nPaiute \nWarao \n\nL6 \nL7 \nL8 \nL9 \n\n3 \n3 \n7 \n\n10 \n10 \n** \n** \n\nGurkhali \nYapese \nOssetic \nRotuman \n\nL12 \nL13 \nL14 \nL15 \n\n19 \n19 \n30 \n29 \n\nL16 \nKomi \nCheremis \nL17 \nMongolian  U8 \nMayan \nU9 \n\n216 \n212 \n2306 \n2298 \n\nTable  3:  Summary  of  results  and  analysis  of  QI  and  QS  learning  (using  weight(cid:173)\nstring  input representations).  Agg=Aggregative Information;  IPS=Inconsistent Primary \nStress;  SCA=Stress  Clash Avoidance;  Alt=Altemation; MPS=Multiple Primary  Stresses; \nMSL=Multiple Stress  Levels.  References  index  into Table  1.  Time is  learning time  in \nepochs. \n\n\f230 \n\nGupta and Touretzky \n\nare  not  crucial  for  present purposes  (see  [Gupta 92]  for  details).  What  is  important  to \nnote  is  that the  learnability predictions generated  by  the  analytical  scheme  described  in \nthe previous section show good agreement with actual perceptron learning experiments on \nthese patterns. \nThe learning results are summarized in Table 4.  It can be seen that the 5-bit characterization \nfits the learning times of various actual and hypothetical patterns reasonably well (although \nthere are exceptions - for example, the hypothetical stress patterns with reference numbers \nh21  through h25  have a higher 5-bit characterization than other stress patterns, but lower \nlearning  times.)  Thus,  the  \"complexity  measure\"  suggested  here  appears  to  identify  a \nnumber of factors  relevant to the learnability of QI stress patterns  within a minimal two(cid:173)\nlayer connectionist architecture.  It also  assesses  their relative  impacts.  The  analysis  is \nundoubtedly a Simplification, but it provides a completely novel framework within which \nto  relate  the  various  learning  results.  The  important point to  note is  that this  analytical \nframework  arises from  a consideration of (a)  the nature of the stress systems,  and (b)  the \nlearning results from simulations.  That is, this framework is empirically based, and makes \nno reference to abstract constructs of the kind that linguistic theory employs. Nevertheless, \nit provides a descriptive framework, much as the linguistic theory does. \n\n5 \n\nINCORPORATING QS SYSTEMS INTO THE ANALYSIS \n\nConsideration of the  QS  stress  patterns  led  to  refinement  of the  IPS  parameter  without \nchanging its setting for the QI patterns.  This parameter is modified so that its value indicates \nthe  proportion of cases  in  which primary  stress  is  not assigned at  the  edge  of a  word. \nAdditionally,  through analysis  of connection weights  for QS  patterns,  a  sixth parameter, \nAggregative Information, is  added as a further index of computational complexity. \n\n6.  Aggregative Information (Agg)  : has value 0 if no aggregative information is required \n(single-positional information suffices);  1 if one kind of aggregative information is \nrequired; and 2 if two kinds of aggregative information are required. \n\nDetailed  discussion of the  analysis  leading  to  these  refinements  is  beyond  the  scope  of \nthis  paper;  the  interested  reader  is  referred  to  [Gupta 92].  The  point  we  wish  to  make \nhere  is  that,  with  these  modifications,  the  same  parameter  scheme  can  be  used  for  both \nthe QI and QS  language classes,  with good learnability predictions within each  class,  as \nshown in Table 3.  Note that in this table,  learning times for all languages are reported in \nterms  of the  weight-string representation (255 input patterns) rather than the unweighted \nsyllabic representation (7 input patterns) used for the initial QI studies.  Both the QI and QS \nresults fall into a single analysis within this generalized parameter scheme and weight-string \nrepresentation, but with a less perfect fit than the within-class results. \n\n6  DISCUSSION \n\nTraditional linguistic analysis has devised abstract theoretical constructs such as \"metrical \nfoot\"  to  describe  linguistic  stress  systems.  Learnability  arguments  were  then  used  to \ndetermine  default  parameter  settings  (e.g.,  whether  feet  should  by  default  be  assumed \nto  be  bounded or unbounded,  left  or  right  dominant,  etc.)  based  on  the  feasibility  of \nlogically deducing correct settings  from  an  initial state.  As  an  example,  in one  analysis \n\n\fConnectionist Learning and Linguistic Stress \n\n231 \n\nIPS  SCA  Alt  MPS  MSL \n\nLANGUAGE \n\nREF \n\n0 \n\n0 \n\n0 \n0 \n1 \n1 \n\n0 \n\n0 \n\n1 \n1 \n0 \n0 \n\n0 \n\n1 \n\n0 \n1 \n0 \n1 \n\n1 \n\n1 \n\n0 \n\n1 \n0 \n0 \n\n0 \n0 \n1 \n\n1 \n\n1 \n0 \n\n0 \n0 \n0 \n1 \n\n1 \n1 \n\n1 \n0 \n0 \n\n1 \n1 \n0 \n\n1 \n\n1 \n0 \n\n0 \n1 \n1 \n0 \n\n1 \n1 \n\n1 \n0 \n1 \n\n0 \n1 \n1 \n\n0 \n\n1 \n0 \n\n1 \n0 \n1 \n1 \n\n0 \n1 \n\n1 \n1 \n\n1 \n1 \n1 \n\n1 \n\n1 \n\n1 \n\n1 \n1 \n1 \n1 \n\n1 \n1 \n\nLl \nL2 \nhI \nh2 \nh3 \nh4 \nh5 \nh6 \n\nL3 \nL4 \nh7 \nh8 \nh9 \nhl0 \nhll \nh12 \nh13 \nh14 \nh15 \nh16 \n\nLatvian \nFrench \nLatvian2stress \nLatvian3stress \nFrench2stress \nFrench 3 stress \nLatvian2edge \nLatvian2edge2stress \nimpossible \nMaranungku \nWeri \nMaranungku3stress \nWeri3stress \nLatvian2edge2stress-alt \nGarawa-SC \nGarawa2stress-SC \nMaranungku 1 stress \nWeri 1 stress \nLatvian2edge-alt \nGarawal stress-SC \nLatvian2edge2stress-lalt \nimpossible \nGarawa-non-alt \nh17 \nh18 \nLatvian3stress2edge-SCA \nLatvian2edge-SCA \nh19 \nLatvian2edge2stress-SCA \nh20 \nL5 \nGarawa \nh21 \nGarawa2stress \nh22 \nLatvian2edge2stress-alt -SCA \nh23 \nGarawalstress \nLatvian2edge-alt-SCA \nh24 \nLatvian2edge2stress-lalt-SCA  h25 \nL6 \nLakota \nL7 \nSwahili \nLakota2stress \nh26 \nh27 \nLakota2edge \nh28 \nLakota2edge2stress \nPaiute \nL8 \nL9 \nWarao \nh29 \nLakota-alt \nh30 \nLakota2stress-alt \n\nEPOCHS \n(syllabic) \n17 \n16 \n21 \n11 \n23 \n14 \n30 \n37 \n\n37 \n34 \n43 \n41 \n58 \n38 \n50 \n61 \n65 \n78 \n88 \n85 \n\n164 \n163 \n194 \n206 \n165 \n71 \n91 \n121 \n126 \n129 \n255 \n254 \n** \n** \n** \n** \n** \n** \n** \n\nTable  4:  Analysis  of Quantity-Insensitive learning  using  the  syllabic  input representa(cid:173)\ntion.  IPS=Inconsistent Primary  Stress;  SCA=Stress  Clash  Avoidance;  Alt=Altemation; \nMPS=Multiple Primary Stresses;  MSL=Multiple Stress Levels.  References  LI-L9 index \ninto Table 1. \n\n\f232 \n\nGupta and Touretzky \n\n[Dresher 90, p.  191], \"metrical feet\"  are taken  to be \"iterative\" by default, since there is \nevidence that can  cause revision of this  default if it turns  out to be the  incorrect setting, \nbut there might not be such disconfirming evidence if the feet were by default taken to be \n\"non-iterative\".  We provide an alternative to logical deduction arguments for determining \n\"markedness\"  of parameter  values,  by  measuring  learnability  (and  hence  markedness) \nempirically.  The parameters  of our novel analysis  generate both a partial description of \neach stress pattern and a prediction of its learnability.  Furthermore, our parameters encode \nlinguistically salient concepts  (e.g.,  stress clash avoidance) as  well as  concepts that have \ncomputational significance (single-positional vs. aggregative information.)  Although our \nanalyses do not explicitly invoke theoretical linguistic constructs such as metrical feet, there \nare suggestive similarities between such constructs and the weight patterns the perceptron \ndevelops [Gupta 91]. \n\nIn conclusion, this work offers a fresh perspective on a well-studied linguistic domain, and \nsuggests that machine learning techniques in conjunction with more traditional tools might \nprovide the basis for a new approach to the investigation of language. \n\nAcknowledgements \n\nWe  would  like  to  acknowledge  the  feedback  prOvided  by  Deirdre  Wheeler  throughout \nthe  course  of this  work.  The  first  author  would  like  to  thank  David Evans  for  access \nto  exceptional  computing  facilities  at  Carnegie  Mellon's  Laboratory  for  Computational \nLinguistics,  and  Dan  Everett,  Brian  MacWhinney,  Jay  McClelland,  Eric  Nyberg,  Brad \nPritchett and Steve Small for helpful discussion of earlier versions of this paper.  Of course, \nnone of them is responsible for any errors. \n\nThe second author was supported by a grant from Hughes Aircraft Corporation, and by the \nOffice of Naval Research under contract number NOOOI4-86-K-0678. \n\nReferences \n[Dresher 90]  Dresher,  B.,  &  Kaye,  J.,  A  Computational  Learning  Model  for  Metrical \n\nPhonology, Cognition 34, 137-195. \n\n[Goldsmith 90]  Goldsmith, J.,  Autosegmental and Metrical Phonology,  Basil Blackwell, \n\nOxford, England, 1990. \n\n[Gupta 91]  Gupta, P. & Touretzky, D., What a perceptron reveals about metrical phonology. \nProceedings of the Thirteenth Annual Conference of the Cognitive Science Society, 334-\n339. Lawrence Erlbaum, Hillsdale, NJ,  1991. \n\n[Gupta 92]  Gupta, P.  & Touretzky,  D., Connectionist Models and Linguistic Theory:  In(cid:173)\n\nvestigations of Stress Systems in Language. Manuscript. \n\n[Hayes 80]  Hayes,  B.,  A  Metrical  Theory  of Stress  Rules,  doctoral  dissertation,  Mas(cid:173)\n\nsachusetts  Institute of Technology,  Cambridge,  MA,  1980.  Circulated by  the  Indiana \nUniversity Linguistics Club, 1981. \n\n[Rumelhart 86]  Rumelhart, D., Hinton, G., & Williams, R, Learning Internal Representa(cid:173)\n\ntions by Error Propagation, in D. Rumelhart, J. McClelland & the PDP Research Group. \nParallel Distributed Processing. Volume 1: Foundations, MIT Press, Cambridge, MA, \n1986. \n\n\f", "award": [], "sourceid": 458, "authors": [{"given_name": "Prahlad", "family_name": "Gupta", "institution": null}, {"given_name": "David", "family_name": "Touretzky", "institution": null}]}