{"title": "Learning to localise sounds with spiking neural networks", "book": "Advances in Neural Information Processing Systems", "page_first": 784, "page_last": 792, "abstract": "To localise the source of a sound, we use location-specific properties of the signals received at the two ears caused by the asymmetric filtering of the original sound by our head and pinnae, the head-related transfer functions (HRTFs). These HRTFs change throughout an organism's lifetime, during development for example, and so the required neural circuitry cannot be entirely hardwired. Since HRTFs are not directly accessible from perceptual experience, they can only be inferred from filtered sounds. We present a spiking neural network model of sound localisation based on extracting location-specific synchrony patterns, and a simple supervised algorithm to learn the mapping between synchrony patterns and locations from a set of example sounds, with no previous knowledge of HRTFs. After learning, our model was able to accurately localise new sounds in both azimuth and elevation, including the difficult task of distinguishing sounds coming from the front and back.", "full_text": "Learning to localise sounds with spiking neural\n\nnetworks\n\nDan F. M. Goodman\n\nD\u00b4epartment d\u2019Etudes Cognitive\n\nEcole Normale Sup\u00b4erieure\n\n29 Rue d\u2019Ulm\n\nParis 75005, France\n\ndan.goodman@ens.fr\n\nRomain Brette\n\nD\u00b4epartment d\u2019Etudes Cognitive\n\nEcole Normale Sup\u00b4erieure\n\n29 Rue d\u2019Ulm\n\nParis 75005, France\n\nromain.brette@ens.fr\n\nAbstract\n\nTo localise the source of a sound, we use location-speci\ufb01c properties of the signals\nreceived at the two ears caused by the asymmetric \ufb01ltering of the original sound by\nour head and pinnae, the head-related transfer functions (HRTFs). These HRTFs\nchange throughout an organism\u2019s lifetime, during development for example, and\nso the required neural circuitry cannot be entirely hardwired. Since HRTFs are\nnot directly accessible from perceptual experience, they can only be inferred from\n\ufb01ltered sounds. We present a spiking neural network model of sound localisation\nbased on extracting location-speci\ufb01c synchrony patterns, and a simple supervised\nalgorithm to learn the mapping between synchrony patterns and locations from a\nset of example sounds, with no previous knowledge of HRTFs. After learning, our\nmodel was able to accurately localise new sounds in both azimuth and elevation,\nincluding the dif\ufb01cult task of distinguishing sounds coming from the front and\nback.\n\nKeywords: Auditory Perception & Modeling (Primary); Computational Neural Models, Neuro-\nscience, Supervised Learning (Secondary)\n\n1\n\nIntroduction\n\nFor many animals, it is vital to be able to quickly locate the source of an unexpected sound, for\nexample to escape a predator or locate a prey. For humans, localisation cues are also used to isolate a\nspeaker in a noisy environment. Psychophysical studies have shown that source localisation relies on\na variety of acoustic cues such as interaural time and level differences (ITDs and ILDs) and spectral\ncues (Blauert 1997). These cues are highly dependent on the geometry of the head, body and pinnae,\nand can change signi\ufb01cantly during an animal\u2019s lifetime, notably during its development but also in\nmature animals (which are known to be able to adapt to these changes, for example Hofman et al.\n1998). Previous neural models addressed the mechanisms of cue extraction, in particular neural\nmechanisms underlying ITD sensitivity, using simpli\ufb01ed binaural stimuli such as tones or noise\nbursts with arti\ufb01cially induced ITDs (Colburn, 1973; Reed and Blum, 1990; Gerstner et al., 1996;\nHarper and McAlpine, 2004; Zhou et al., 2005; Liu et al., 2008), but did not address the problem of\nlearning to localise natural sounds in realistic acoustical environments.\nSince the physical laws of sound propagation are linear, the sound S produced by a source is received\nat any point x of an acoustical environment as a linearly \ufb01ltered version Fx \u2217 S (linear convolution),\nwhere the \ufb01lter is speci\ufb01c of the location x of the listener, the location of the source and the acoustical\nenvironment (ground, wall, objects, etc.). For binaural hearing, the acoustical environment includes\nthe head, body and pinnae, and the sounds received at the two ears are FL \u2217 S and FR \u2217 S, where\n(FL, FR) is a pair of location-speci\ufb01c \ufb01lters. Because the two sounds originate from the same\n\n1\n\n\fsignal, the binaural stimulus has a speci\ufb01c structure, which should result in synchrony patterns in\nthe encoding neurons. Speci\ufb01cally, we modelled the response of monaural neurons by a linear\n\ufb01ltering of the sound followed by a spiking nonlinearity. Two neurons A and B responding to two\ndifferent sides (left and right), with receptive \ufb01elds NA and NB, transform the signals NA \u2217 FL \u2217 S\nand NB \u2217 FR \u2217 S into spike trains. Thus, synchrony between A and B occurs whenever NA \u2217 FL =\nNB \u2217 FR, i.e., for a speci\ufb01c set of \ufb01lter pairs (FL, FR). Thus, in our model, sounds presented\nat a given location induce speci\ufb01c synchrony patterns, which then activate a speci\ufb01c assembly of\npostsynaptic neurons (coincidence detection neurons), in a way that is independent of the source\nsignal (see Goodman and Brette, in press). Learning a new location consists in assigning a label to\nthe activated assembly, using a teacher signal (for example visual input).\nWe used measured human HRTFs to generate binaural signals at different source locations from a\nset of various sounds. These signals were used to train the model and we tested the localisation\naccuracy with new sounds. After learning, the model was able to accurately locate unknown sounds\nin both azimuth and elevation.\n\n2 Methods\n\n2.1 Virtual acoustics\n\nSound sources used were: broadband white noise; recordings of instruments and voices from the\nRWC Music Database (http://staff.aist.go.jp/m.goto/RWC-MDB/); and recordings\nof vowel-consonant-vowel sounds (Lorenzi et al., 1999). All sounds were of 1 second duration and\nwere presented at 80 dB SPL. Sounds were \ufb01ltered by head-related impulse responses (HRIRs)\nfrom the IRCAM LISTEN HRTF Database (http://recherche.ircam.fr/equipes/\nsalles/listen/index.html). This database includes 187 approximately evenly spaced lo-\ncations at all azimuths in 15 degree increments (except for high elevations) and elevations from -45\nto 90 degrees in 15 degree increments. HRIRs from this and other databases do not provide suf\ufb01-\nciently accurate timing information at frequencies below around 150Hz, and so subsequent cochlear\n\ufb01ltering was restricted to frequencies above this point.\n\n2.2 Mathematical principle\n\nConsider two sets of neurons which respond monaurally to sounds from the left ear and from the\nright ear by \ufb01ltering sounds through a linear \ufb01lter N (modeling their receptive \ufb01eld, corresponding\nto cochlear and neural transformations on the pathway between the ear and the neuron) followed by\nspiking. Each neuron has a different \ufb01lter. Spiking is modeled by an integrate-and-\ufb01re description\nor some other spiking model. Consider two neurons A and B which respond to sounds from the\nleft and right ear, respectively. When a sound S is produced by a source at a given location, it\narrives at the two ears as the binaural signal (FL \u2217 S, FR \u2217 S) (convolution), where (FL, FR) is\nthe location-speci\ufb01c pair of acoustical \ufb01lters. The \ufb01ltered inputs to the two spiking models A and\nB are then NA \u2217 FL \u2217 S and NB \u2217 FR \u2217 S. These will be identical for any sound S whenever\nNA \u2217 FL = NB \u2217 FR, implying that the two neurons \ufb01re synchronously. For each location indicated\nby its \ufb01lter pair (FL, FR), we de\ufb01ne the synchrony pattern as the set of binaural pairs of neurons\n(A, B) such that NA \u2217 FL = NB \u2217 FR. This pattern is location-speci\ufb01c and independent of the\nsource signal S. Therefore, the identity of the synchrony pattern induced by a binaural stimulus\nindicates the location of the source. Learning consists in assigning a synchrony pattern induced by\na sound to the location of the source.\nTo have a better idea of these synchrony patterns, consider a pair of \ufb01lters (F \u2217\nL, F \u2217\nR) that corresponds\nto a particular location x (azimuth, elevation, distance, and possibly also position of the listener in\nthe acoustical environment), and suppose neuron A has receptive \ufb01eld NA = F \u2217\nR and neuron B has\nR \u2217 FL = F \u2217\nL \u2217 FR,\nreceptive \ufb01eld NB = F \u2217\nL. Then neurons A and B \ufb01re in synchrony whenever F \u2217\nin particular when FL = F \u2217\nR, that is, at location x (since convolution is commutative).\nMore generally, if U is a band-pass \ufb01lter and the receptive \ufb01elds of neurons A and B are U \u2217 F \u2217\nR and\nU \u2217 F \u2217\nL, respectively, then the neurons \ufb01re synchronously at location x. The same property applies if\na nonlinearity (e.g. compression) is applied after \ufb01ltering. If the bandwidth of U is very small, then\nU \u2217 F \u2217\nR is essentially the \ufb01lter U followed by a delay and gain. Therefore, to represent all possible\n\nL and FR = F \u2217\n\n2\n\n\flocations in pairs of neuron \ufb01lters, we consider that the set of neural transformations N is a bank of\nband-pass \ufb01lters followed by a set of delays and gains.\nTo decode synchrony patterns, we de\ufb01ne a set of binaural neurons which receive input spike trains\nfrom monaural neurons on both sides (two inputs per neuron). A binaural neuron responds prefer-\nentially when its two inputs are synchronous, so that synchrony patterns are mapped to assemblies\nof binaural neurons. Each location-speci\ufb01c assembly is the set of binaural neurons for which the\ninput neurons \ufb01re synchronously at that location. This is conceptually similar to the Jeffress model\n(Jeffress, 1948), where a neuron is maximally activated when acoustical and axonal delays match,\nand related models (Lindemann, 1986; Gaik, 1993). However, the Jeffress model is restricted to az-\nimuth estimation and it is dif\ufb01cult to implement it directly with neuron models because ILDs always\nco-occur with ITDs and disturb spike-timing.\n\n2.3\n\nImplementation with spiking neuron models\n\nFigure 1: Implementation of the model. The source signal arrives at the two ears after acoustical \ufb01l-\ntering by HRTFs. The two monaural signals are \ufb01ltered by a set of gammatone \ufb01lters \u03b3i with central\nfrequencies between 150 Hz and 5 kHz (cochlear \ufb01ltering). In each band (3 bands shown, between\ndashed lines), various gains and delays are applied to the signal (neural \ufb01ltering F L\nj ) and\nspiking neuron models transform the resulting signals into spike trains, which converge from each\nside on a coincidence detector neuron (same neuron model). The neural assembly corresponding to\na particular location is the set of coincidence detector neurons for which their input neurons \ufb01re in\nsynchrony at that location (one pair for each frequency channel).\n\nj and F R\n\nThe overall structure and architecture of the model is illustrated in Figure 1. All programming\nwas done in the Python programming language, using the \u201cBrian\u201d spiking neural network simulator\npackage (Goodman and Brette, 2009). Simulations were performed on Intel i7 Core processors. The\nlargest model involved approximately one million neurons.\nCochlear and neural \ufb01ltering. Head-\ufb01ltered sounds were passed through a bank of fourth-order\ngammatone \ufb01lters with center frequencies distributed on the ERB scale (central frequencies from\n150 Hz to 5 kHz), modeling cochlear \ufb01ltering (Glasberg and Moore, 1990). Linear \ufb01ltering was\ncarried out in parallel with a custom algorithm designed for large \ufb01lterbanks (around 30,000 \ufb01lters\nin our simulations). Gains and delays were then applied, with delays at most 1 ms and gains at most\n\u00b110 dB.\nNeuron model. The \ufb01ltered sounds were half-wave recti\ufb01ed and compressed by a 1/3 power law\nI = k([x]+)1/3 (where x is the sound pressure in pascals). The resulting signal was used as an\ninput current to a leaky integrate-and-\ufb01re neuron with noise. The membrane potential V evolves\naccording to the equation:\n\n\u03c4m\n\ndV\ndt\n\n= V0 \u2212 V + I(t) + \u03c3\n\n3\n\n\u221a\n\n2\u03c4m\u03be(t)\n\n(30\u00b0, 15\u00b0)(45\u00b0, 15\u00b0)(90\u00b0, 15\u00b0)Neural(cid:31)lteringNeural(cid:31)lteringCochlear(cid:31)lteringCochlear(cid:31)lteringCoincidencedetection\u03b3iFjRFjL\u03b3iLRHRTFHRTF\fParameter Value\n\nDescription\n\nTable 1: Neuron model parameters\n\nVr\nV0\nVt\ntrefrac\n\n\u03c3\n\u03c4m\nW\nk\n\n-60 mV\n-60 mV\n-50 mV\n5 ms\n0 ms\n1 mV\n1 ms\n5 mV\n0.2 V/Pa1/3 Acoustic scaling constant\n\nReset potential\nResting potential\nThreshold potential\nAbsolute refractory period\n(for binaural neurons)\nStandard deviation of membrane potential due to noise\nMembrane time constant\nSynaptic weight for coincidence detectors\n\nwhere \u03c4m is the membrane time constant, V0 is the resting potential, \u03be(t) is Gaussian noise (such\nthat (cid:104)\u03be(t), \u03be(s)(cid:105) = \u03b4(t\u2212s)) and \u03c3 is the standard deviation of the membrane potential in the absence\nof spikes. When V crosses the threshold Vt a spike is emitted and V is reset to Vr and held there\nfor an absolute refractory period trefrac. These neurons make synaptic connections with binaural\nneurons in a second layer (two presynaptic neurons for each binaural neuron). These coincidence\ndetector neurons are leaky integrate-and-\ufb01re neurons with the same equations but their inputs are\nsynaptic. Spikes arriving at these neurons cause an instantaneous increase W in V (where W is the\nsynaptic weight). Parameter values are given in Table 1.\nEstimating location from neural activation. Each location is assigned an assembly of coincidence\ndetector neurons, one in each frequency channel. When a sound is presented to the model, the total\n\ufb01ring rate of all neurons in each assembly is computed. The estimated location is the one assigned to\nthe maximally activated assembly. Figure 2 shows the activation of all location-speci\ufb01c assemblies\nin an example where a sound was presented to the model, after learning.\nComputing assemblies from HRTFs. In the hardwired model, we de\ufb01ned the location-speci\ufb01c as-\nsemblies from the knowledge of HRTFs (the learning algorithm is explained in section 2.4). For a\ngiven location (\ufb01lter pair (FL, FR)) and frequency channel (gammatone \ufb01lter G), we choose the bin-\naural neuron for which the the gains (gL, gR) and delays (dL, dR) of the two presynaptic monaural\nneurons minimize the RMS difference\n\n(cid:115)(cid:90)\n\n\u2206 =\n\n(gL(G \u2217 FL)(t \u2212 dL) \u2212 gR(G \u2217 FR)(t \u2212 dR))2dt,\n\nthat is, the RMS difference between the inputs of the two neurons for a sound impulse at that loca-\ntion. We also impose max(gL, gR) = 1 and max(dL, dR) = 0 (so that one delay is null and the\nother is positive). The RMS difference is minimized when the delays correspond to the maximum of\n\nthe cross-correlation between L and R, C(s) =(cid:82) (G\u2217FL)(t)\u00b7(G\u2217FR)(t+s)dt, so that C(dR\u2212dL)\nis the maximum, and gR/gL = C(dR \u2212 dL)/(cid:82) R(t)2dt.\n\n2.4 Learning\n\nIn the hardwired model, the knowledge of the full set of HRTFs is used to estimate source location.\nBut HRTFs are never directly accessible to the auditory system, because they are always convolved\nwith the source signal. They cannot be genetically wired either, because they depend on the geome-\ntry of the head (which changes during development). In our model, when HRTFs are not explicitly\nknown, location-speci\ufb01c assemblies are learned by presenting unknown sounds at different locations\nto the model, where there is one coincidence detector neuron for each choice of frequency, relative\ndelay and relative gain. Relative delays were uniformly chosen between \u22120.8 ms and 0.8 ms, and\nrelative gains between \u22128 dB and 8 dB uniformly on a dB scale. In total 69 relative delays were\nchosen and 61 relative gains. With 80 frequency channels, this gives a total of roughly 106 neurons\nin the model. When a sound is presented at a given location, we de\ufb01ne the assembly for this location\nby picking the maximally activated neuron in each frequency channel, as would be expected from a\nsupervised Hebbian learning process with a teacher signal (e.g. visual cues). For practical reasons,\n\n4\n\n\fFigure 2: Activation of all location-speci\ufb01c assemblies in response to a sound coming from a par-\nticular location indicated by a black +. The white x shows the model estimate (maximally activated\nassembly). The mapping from assemblies to locations were learned from a set of sounds.\n\nwe did not implement this supervised learning with spiking models, but supervised learning with\nspiking neurons has been described in several previous studies (Song and Abbott, 2001; Davison\nand Frgnac, 2006; Witten et al., 2008).\n\n3 Results\n\nWhen the model is \u201chardwired\u201d using the explicit knowledge of HRTFs, it can accurately localise\na wide range of sounds (Figure 3A-C): for the maximum number of channels we tested (80), we\nobtained an average error of between 2 and 8 degrees for azimuth and 5 to 20 degrees for elevation\n(depending on sound type), and with more channels this error is likely to further decrease, as it\ndid not appear to have reached an asymptote at 80 channels. Performance is better for sounds with\nbroader spectrums, as each channel provides additional information. The model was also able to\ndistinguish between sounds coming from the left and right (with an accuracy of almost 100%), and\nperformed well for the more dif\ufb01cult tasks of distinguishing between front and back (80-85%) and\nbetween up and down (70-90%).\nFigure 3D-F show the results using the learned best delays and gains, using the full training data\nset (seven sounds presented at each location, each of one second duration) and different test sounds.\nPerformance is comparable to the hardwired model. Average azimuth errors for 80 channels are 4-8\ndegrees, and elevation errors are 10-27 degrees. Distinguishing left and right is done with close to\n100% accuracy, front and back with 75-85% and up and down with 65-90%. Figure 4 shows how\nthe localisation accuracy improves with more training data. With only a single sound of one second\nduration at each location, the performance is already very good. Increasing the training to three sec-\nonds of training data at each location improves the accuracy, but including further training data does\nnot appear to lead to any signi\ufb01cant improvement. Although it is close, the performance does not\nseem to converge to that of the hardwired model, which might be due to a limited sampling of delays\nand gains (69 relative delays and 61 relative gains), or perhaps to the presence of physiological noise\nin our models (Goodman and Brette, in press).\nFigure 5 shows the properties of neurons in a location-speci\ufb01c assembly: interaural delay (Figure\n5A) and interaural gain difference (Figure 5B) for each frequency channel. For this location, the\nassemblies in the hardwired model and with learning were very similar, which indicates that the\nlearning procedure was indeed able to catch the binaural cues associated with that location. The\ndistributions of delays and gain differences were similar in the hardwired model and with learning.\nIn the hardwired model, these interaural delays and gains correspond to the ITDs and ILDs in \ufb01ne\nfrequency bands. To each location corresponds a speci\ufb01c frequency-dependent pattern of ITDs and\nILDs, which is informative of both azimuth and elevation. In particular, these patterns are different\nwhen the location is reversed between front and back (not shown), and this difference is exploited\nby the model to distinguish between these two cases.\n\n5\n\n15010050050100150Azimuth (deg)4020020406080Elevation (deg)\fFigure 3: Performance of the hard-wired model (A-C) and with learning (D-F). A, D, Mean error in\nazimuth estimates as a function of the number of frequency channels (i.e., assembly size) for white\nnoise (red), vowel-consonant-vowel (blue) and musical instruments (green). Front-back reversed\nlocations were considered as having the same azimuth. The channels were selected at random be-\ntween 150 Hz and 5 kHz and results were averaged over many random choices. B, E, Mean error in\nelevation estimates. C, F, Categorization performance discriminating left and right (solid), front and\nback (dashed) and up and down (dotted).\n\nFigure 4: Performance improvement with training (80 frequency channels). A, Average estimation\nerror in azimuth (blue) and elevation (green) as a function of the number of sounds presented at\neach location during learning (each sound lasts 1 second). The error bars represent 95% con\ufb01dence\nintervals. The dashed lines indicate the estimation error in the hardwired model (when HRTFs are\nexplicitly known). B, Categorization performance vs. number of sounds per location for discrimi-\nnating left and right (green), front and back (blue) and up and down (red).\n\n4 Discussion\n\nThe sound produced by a source propagates to the ears according to linear laws. Thus the ears\nreceive two differently \ufb01ltered versions of the same signal, which induce a location-speci\ufb01c structure\n\n6\n\nABCDEFAB\fFigure 5: Location-speci\ufb01c assembly in the hardwired model and with learning. A, Preferred in-\nteraural delay vs. preferred frequency for neurons in an assembly corresponding to one particular\nlocation, in the hardwired model (white circles) and with learning (black circles). The colored\nbackground shows the distribution of preferred delays in all neurons in the hardwired model. B,\nInteraural gain difference vs. preferred frequency for the same assemblies.\n\nin the binaural stimulus. When binaural signals are transformed by a heterogeneous population of\nneurons, this structure is mapped to synchrony patterns, which are location-speci\ufb01c. We designed a\nsimple spiking neuron model which exploits this property to estimate the location of sound sources\nin a way that is independent of the source signal. In the model, each location activates a speci\ufb01c\nassembly. We showed that the mapping between assemblies and locations can be directly learned\nin a supervised way from the presentation of a set of sounds at different locations, with no previous\nknowledge of the HRTFs or the sounds. With 80 frequency channels, we found that 1 second of\ntraining data per location was enough to estimate the azimuth of a new sound with mean error 6\ndegrees and the elevation with error 18 degrees.\nHumans can learn to localise sound sources when their acoustical cues change, for example when\nmolds are inserted into their ears (Hofman et al., 1998; Zahorik et al., 2006). Learning a new\nmapping can take a long time (several weeks in the \ufb01rst study), which is consistent with the idea\nthat the new mapping is learned from exposure to sounds from known locations.\nInterestingly,\nthe previous mapping is instantly recovered when the ear molds are removed, meaning that the\nrepresentations of the two acoustical environments do not interfere. This is consistent with our\nmodel, in which two acoustical environments would be represented by two possibly overlapping\nsets of neural assemblies.\nIn our model, we assumed that the receptive \ufb01eld of monaural neurons can be modeled as a band-pass\n\ufb01lter with various gains and delays. Differences in input gains could simply arise from differences in\nmembrane resistance, or in the number and strength of the synapses made by auditory nerve \ufb01bers.\nDelays could arise from many causes: axonal delays (either presynaptic or postsynaptic), cochlear\ndelays (Joris et al., 2006), inhibitory delays (Brand et al., 2002). The distribution of best delays\nof the binaural neurons in our model re\ufb02ect the distribution of ITDs in the acoustical environment.\nThis contradicts the observation in many species that the best delays are always smaller than half the\ncharacteristic period, i.e., they are within the \u03c0-limit (Joris and Yin, 2007). However, we checked\nthat the model performed almost equally well with this constraint (Goodman and Brette, in press),\nwhich is not very surprising since best delays above the \u03c0-limit are mostly redundant. In small\nmammals (guinea pigs, gerbils), it has been shown that the best phases of binaural neurons in the\nMSO and IC are in fact even more constrained, since they are scattered around \u00b1\u03c0/4, in constrast\nwith birds (e.g. barn owl) where the best phases are continuously distributed (Wagner et al., 2007).\nHowever, in larger mammals such as cats, best IPDs in the MSO are more continuously distributed\n(Yin and Chan, 1990), with a larger proportion close to 0 (Figure 18 in Yin and Chan, 1990). It\nhas not been measured in humans, but the same optimal coding theory that predicts the discrete\ndistribution of phases in small mammals predicts that best delays should be continuously distributed\nabove 400 Hz (80% of the frequency channels in our model). In addition, psychophysical results also\n\n7\n\nAB\fimply that humans can estimate both the azimuth and elevation of low-pass \ufb01ltered sound sources\n(< 3 kHz) (Algazi et al., 2001), which only contain binaural cues. This is contradictory with the two-\nchannel model (best delays at \u00b1\u03c0/4) and in agreement with ours (including the fact that elevation\ncould only be estimated away from the median plane in these experiments).\nOur model is conceptually similar to a recent signal processing method (with no neural implemen-\ntation) to localize sound sources in the horizontal plane (Macdonald, 2008), where coincidence de-\ntection is replaced by Pearson correlation between the two transformed monaural broadband signals\n(no \ufb01lterbank). However, that method requires explicit knowledge of the HRTFs, so that it cannot\nbe directly learned from natural exposure to sounds.\nThe HRTFs used in our virtual acoustic environment were recorded at a constant distance, so that we\ncould only test the model performance in estimating the azimuth and elevation of a sound source.\nHowever, in principle, it should also be able to estimate the distance when the source is close.\nIt should also apply equally well to non-anechoic environments, because our model only relies\non the linearity of sound propagation. However, a dif\ufb01cult task, which we have not addressed, is\nto locate sounds in a new environment, because re\ufb02ections would change the binaural cues and\ntherefore the location-speci\ufb01c assemblies. One possibility would be to isolate the direct sound from\nthe re\ufb02ections, but this requires additional mechanisms, which probably underlie the precedence\neffect (Litovsky et al., 1999).\n\nReferences\nAlgazi, V. R., C. Avendano, and R. O. Duda (2001, March). Elevation localization and head-related transfer\nfunction analysis at low frequencies. The Journal of the Acoustical Society of America 109(3), 1110\u20131122.\nBrand, A., O. Behrend, T. Marquardt, D. McAlpine, and B. Grothe (2002). Precise inhibition is essential for\n\nmicrosecond interaural time difference coding. Nature 417(6888), 543.\n\nColburn, H. S. (1973, December). Theory of binaural interaction based on auditory-nerve data. i. general\nstrategy and preliminary results on interaural discrimination. The Journal of the Acoustical Society of Amer-\nica 54(6), 1458\u20131470.\n\nDavison, A. P. and Y. Frgnac (2006, May). Learning Cross-Modal spatial transformations through spike\n\nTiming-Dependent plasticity. J. Neurosci. 26(21), 5604\u20135615.\n\nGaik, W. (1993, July). Combined evaluation of interaural time and intensity differences: Psychoacoustic results\n\nand computer modeling. The Journal of the Acoustical Society of America 94(1), 98\u2013110.\n\nGerstner, W., R. Kempter, J. L. van Hemmen, and H. Wagner (1996). A neuronal learning rule for sub-\n\nmillisecond temporal coding. Nature 383(6595), 76.\n\nGlasberg, B. R. and B. C. Moore (1990, August). Derivation of auditory \ufb01lter shapes from notched-noise data.\n\nHearing Research 47(1-2), 103\u2013138. PMID: 2228789.\n\nGoodman, D. F. M. and R. Brette (2009). The Brian simulator. Frontiers in Neuroscience 3(2), 192\u2013197.\nGoodman, D. F. M. and R. Brette (in press). Spike-timing-based computation in sound localization. PLoS\n\nComp. Biol..\n\nHarper, N. S. and D. McAlpine (2004). Optimal neural population coding of an auditory spatial cue. Na-\n\nture 430(7000), 682\u2013686.\n\nHofman, P. M., J. G. V. Riswick, and A. J. V. Opstal (1998). Relearning sound localization with new ears. Nat\n\nNeurosci 1(5), 417\u2013421.\n\nJeffress, L. A. (1948, February). A place theory of sound localization. Journal of Comparative and Physiolog-\n\nical Psychology 41(1), 35\u20139. PMID: 18904764.\n\nJoris, P. and T. C. T. Yin (2007, February). A matter of time: internal delays in binaural processing. Trends in\n\nNeurosciences 30(2), 70\u20138. PMID: 17188761.\n\nJoris, P. X., B. V. de Sande, D. H. Louage, and M. van der Heijden (2006). Binaural and cochlear disparities.\n\nProceedings of the National Academy of Sciences 103(34), 12917.\n\nLindemann, W. (1986, December). Extension of a binaural cross-correlation model by contralateral inhibition.\ni. simulation of lateralization for stationary signals. The Journal of the Acoustical Society of America 80(6),\n1608\u20131622.\n\nLitovsky, R. Y., H. S. Colburn, W. A. Yost, and S. J. Guzman (1999, October). The precedence effect. The\n\nJournal of the Acoustical Society of America 106(4), 1633\u20131654.\n\nLiu, J., H. Erwin, S. Wermter, and M. Elsaid (2008). A biologically inspired spiking neural network for sound\n\nlocalisation by the inferior colliculus. In Arti\ufb01cial Neural Networks - ICANN 2008, pp. 396\u2013405.\n\n8\n\n\fLorenzi, C., F. Berthommier, F. Apoux, and N. Bacri (1999, October). Effects of envelope expansion on speech\n\nrecognition. Hearing Research 136(1-2), 131\u2013138.\n\nMacdonald, J. A. (2008, June). A localization algorithm based on head-related transfer functions. The Journal\n\nof the Acoustical Society of America 123(6), 4290\u20134296. PMID: 18537380.\n\nReed, M. C. and J. J. Blum (1990, September). A model for the computation and encoding of azimuthal\ninformation by the lateral superior olive. The Journal of the Acoustical Society of America 88(3), 1442\u2013\n1453. PMID: 2229677.\n\nSong, S. and L. F. Abbott (2001, October). Cortical development and remapping through spike Timing-\n\nDependent plasticity. Neuron 32(2), 339\u2013350.\n\nWagner, H., A. Asadollahi, P. Bremen, F. Endler, K. Vonderschen, and M. von Campenhausen (2007). Dis-\ntribution of interaural time difference in the barn owl\u2019s inferior colliculus in the low- and High-Frequency\nranges. J. Neurosci. 27(15), 4191\u20134200.\n\nWitten, I. B., E. I. Knudsen, and H. Sompolinsky (2008, August). A hebbian learning rule mediates asymmetric\n\nplasticity in aligning sensory representations. J Neurophysiol 100(2), 1067\u20131079.\n\nYin, T. C. and J. C. Chan (1990). Interaural time sensitivity in medial superior olive of cat. J Neurophysiol 64(2),\n\n465\u2013488.\n\nZahorik, P., P. Bangayan, V. Sundareswaran, K. Wang, and C. Tam (2006, July). Perceptual recalibration in\nhuman sound localization: Learning to remediate front-back reversals. The Journal of the Acoustical Society\nof America 120(1), 343\u2013359.\n\nZhou, Y., L. H. Carney, and H. S. Colburn (2005, March). A model for interaural time difference sensitivity\nin the medial superior olive: Interaction of excitatory and inhibitory synaptic inputs, channel dynamics, and\ncellular morphology. J. Neurosci. 25(12), 3046\u20133058.\n\n9\n\n\f", "award": [], "sourceid": 136, "authors": [{"given_name": "Dan", "family_name": "Goodman", "institution": null}, {"given_name": "Romain", "family_name": "Brette", "institution": null}]}