{"title": "Discriminative Densities from Maximum Contrast Estimation", "book": "Advances in Neural Information Processing Systems", "page_first": 1009, "page_last": 1016, "abstract": null, "full_text": "Discriminative Densities from Maximum\n\nContrast Estimation\n\nPeter Meinicke\n\nNeuroinformatics Group\nUniversity of Bielefeld\n\nBielefeld, Germany\n\nThorsten Twellmann\nNeuroinformatics Group\nUniversity of Bielefeld\n\nBielefeld, Germany\n\npmeinick@techfak.uni-bielefeld.de\n\nttwellma@techfak.uni-bielefeld.de\n\nHelge Ritter\n\nhelge@techfak.uni-bielefeld.de\n\nNeuroinformatics Group\nUniversity of Bielefeld\n\nBielefeld, Germany\n\nAbstract\n\nWe propose a framework for classi\ufb01er design based on discriminative\ndensities for representation of the differences of the class-conditional dis-\ntributions in a way that is optimal for classi\ufb01cation. The densities are\nselected from a parametrized set by constrained maximization of some\nobjective function which measures the average (bounded) difference, i.e.\nthe contrast between discriminative densities. We show that maximiza-\ntion of the contrast is equivalent to minimization of an approximation\nof the Bayes risk. Therefore using suitable classes of probability den-\nsity functions, the resulting maximum contrast classi\ufb01ers (MCCs) can\napproximate the Bayes rule for the general multiclass case. In particular\nfor a certain parametrization of the density functions we obtain MCCs\nwhich have the same functional form as the well-known Support Vec-\ntor Machines (SVMs). We show that MCC-training in general requires\nsome nonlinear optimization but under certain conditions the problem\nis concave and can be tackled by a single linear program. We indicate\nthe close relation between SVM- and MCC-training and in particular we\nshow that Linear Programming Machines can be viewed as an approxi-\nmate realization of MCCs. In the experiments on benchmark data sets,\nthe MCC shows a competitive classi\ufb01cation performance.\n\n1 Introduction\nIn the Bayesian framework of classi\ufb01cation the ultimate goal of a classi\ufb01er \u0002\u0001\u0004\u0003\u0006\u0005\b\u0007\n\t\f\u000b\u000e\r\n\u000f\u0011\u0010\u0013\u0012\u0015\u0014\u0016\u0014\u0015\u0014\u0017\u0012\u0019\u0018\u000e\u001a\n\u0005 which\ndenotes the loss for assigning a given feature vector to class \u001d , while it actually belongs to\nclass \u001c , with \u0018\n\u0001\u0004\u0003! \u001e\u001c\"\u0005 being the class-conditional prob-\nability density functions (PDFs) and #%$ denoting the corresponding apriori probabilities of\n\nis to minimize the expected risk of misclassi\ufb01cation measured by \u001b\n\nbeing the number of classes. With \u001f\n\n\u0001\u0004\u001c\n\n\u0012\u001e\u001d\n\n\fclass-membership we have the risk\n\n(1)\n\n$\u001c\u001a denotes the Kro-\n\n\u0001\u0004\u001c\n\n\u0002\u0001\n\n\u0003\u0006\u0005\n\n\u0010\u0017\u0016\u0019\u0018\n\n\u0001\u0004\u0003! \u001e\u001c\"\u0005\u0015\u0014\n$\u001b\u001a , where\u0018\n\n\u0001\b\u0007\n\n\u0005\n\t\n\n\u0002\u0001\n\n\u0003\u0006\u0005\n\n\u0002\u0001\u0004\u0003\u0006\u0005\n\n\u0001\f\u000b\n$\u000f\u000e\u0011\u0010\nWith the standard \u201czero-one\u201d loss function \u001b\n\u0001\u0004\u001d\u001f\u001e! #\"\u0006\u001d%$\n\n$\u0013\u0012\n\nnecker delta, it is easy to show (see e.g. [3]) that the expected risk is minimized, if one\nchooses the classi\ufb01er\n\n\u0001\u0004\u0003! \u001e\u001c\"\u0005\n\n\u0002\u0001\n\u0003\u0006\u0005\nThe resulting lower bound on\nis known as the Bayes risk which limits the average perfor-\nmance of the classi\ufb01er \u0002\u0001\u0004\u0003\u0006\u0005 . Because the class-conditional densities are usually unknown,\none way to realize the above classi\ufb01er is to use estimates of these densities instead. This\nleads to the so-called plug-in classi\ufb01ers, which are Bayes-consistent if the density estima-\ntors are consistent (e.g. [9]). Due to the notoriously slow convergence of density estimates\nthe plug-in scheme usually isn\u2019t the best recipe for classi\ufb01er design and as an alternative\nmany discriminant functions including Neural Networks (see [1, 9] for an overview) and\nSupport Vector Machines (SVMs) [2, 12] have been proposed which are trained directly to\nminimize the empirical classi\ufb01cation error.\n\n(2)\n\n(3)\n\n\u0001\u0004\u0003\n\n\u00070/\n\n'\u0013(#*,+\n\u0003;/0<\n\n\"87:9\n\u0003! \u001e\u001c\"\u0005\n\n\u0005\n\t with\n\n\u001a\u0011\u0014\n\ninative) densities \u001f\n\nWe recently proposed a method for the design of density-based classi\ufb01ers without resort-\ning to the usual density estimation schemes of the plug-in approach [6]. Instead we utilized\ndiscriminative densities with parameters optimized to solve the classi\ufb01cation problem. The\napproach requires maximization of the average bounded difference between class (discrim-\n\u0005 , which we refer to as the contrast of the underlying \u201ctrue\u201d dis-\n\n\u0001\u0004\u0003\ntributions. The')(#*,+ -bounded contrast is the expectation\u0003-\u0005\n\u0003;/0<\n\u00070/\n\u001a43\n\u000e65\nThe idea is to \ufb01nd \u0018\n\u0001\u0004\u0003;/0<\n\u0005 , which represent the underlying\nin a way, that is optimal for classi\ufb01cation. When\ndistributions with \u201ctrue\u201d densities \u001f\nmaximizing the contrast with respect to the parameters <\n$ of the discriminative densities\nthe upper bound'\u0013(#*,+ plays a central role because it prevents the learning algorithm from\n\nincreasing the differences between discriminative densities where the differences between\nthe true densities are already large.\n\n(#*,+\n(#*2+\ndiscriminative densities \u001f\n\nIn this paper we show that with some slight modi\ufb01cation the contrast can be viewed as\nan approximation of the negative Bayes risk (up to some constant shift and scaling) which\nis valid for the binary as well as for the general multiclass case. Therefore for certain\nparametrizations of the discriminative densities MCCs allow to \ufb01nd an optimal trade-off\nbetween the classical plug-in Bayes-consistency and the consistency which arises from di-\nrect minimization of the approximate Bayes risk. Furthermore, for a particular parametriza-\ntion of the PDFs, we obtain certain kinds of Linear Programming Machines (LPMs) [4] as\n(in general) approximate solutions of maximum contrast estimation. In that way MCCs\nprovide a Bayes-consistent approach to realize multiclass LPMs / SVMs and they suggest\nan interpretation of the magnitude of the LPM / SVM classi\ufb01cation function in terms of\ndensity differences which provide a probabilistic measure of con\ufb01dence. For the case of\nLPMs we propose an extended optimization procedure for maximization of the contrast\nvia iteration of linear optimizations. Inspired by the MCC-framework, for the resulting\nSequential Linear Programming Machines (SLPM) we propose a new regularizer which\nallows to \ufb01nd an optimal trade-off between the above mentioned two approaches to Bayes\nconsistency. In the experiments we analyse the performance of the SLPM on simulated and\nreal world data.\n\n\u001b\n\u0012\n\n#\n\u001b\n\u0012\n\u0005\n\u001f\n\u0003\n\u0014\n\u0001\n\u001c\n\u0012\n\u001d\n\u0005\n\u0001\n$\n#\n$\n\u001f\n$\n \n&\n\u001a\n.\n\u0012\n1\n.\n\u0001\n\u0003\n\u0012\n'\n\u0005\n\u0001\n\n\u000f\n'\n\u0012\n#\n5\n\u001f\n\u0001\n5\n\u0005\n\u0016\n#\n\u001a\n\u001f\n\u0001\n\u001a\n\u0005\n$\n\u0001\n\f2 Maximum Contrast Estimation\n\nFor the design of MCCs the \ufb01rst step, which is the same as for the plug-in concept, requires\nto replace the unknown class-conditional densities of the Bayes classi\ufb01er (2) by suitably\nparametrized PDFs. Then, instead of choosing the parameters for an approximation of\nthe original (true) densities (e.g. by maximum likelihood estimation) as with the plug-in\nscheme, the density parameters are choosen to maximize the so-called contrast which is\n\n\u001f\u0006\u0002\n\n \t\u0007\u0013\u0005\n\n\u0003\u0006\u0005\n\n\u001f\u0003\u0002\n\n\u0001\u0004\u0003\u0006\u0005\n\u0001\u0004\u0003! \n\"\u0006\u001d%$\f\u000b\n\n\u001e, \n\n\u001e, #\"\u0006\u001d\u001f$\n\n\u0001\n\n'\u0015(#*,+\n\nthe expected value of the'\n\nFor the case of an unbounded contrast, i.e.\n, the general maximum contrast\nsolution can be found analytically and for notational simplicity we derive it for the binary\ncase with equal apriori probabilities, where the contrast can be written as\n\u0003! \b\u0007\u0013\u0005\n\n(#*,+ -bounded density differences as de\ufb01ned in (3).\n\u0014\u0013\u0003\n\u0001\u0004\u0003\u0006\u0005\u0015\u0014\n\n\u0001\u0004\u0003\u0006\u0005\n\u0001\u0004\u0003! \b\u0007\n\u0001\u0004\u0003\u0006\u0005\nThus the unbounded contrast is maximized for \n\nwith the peaks of the Delta (Dirac) functions located at \u0003\nand \u0003\n\n \t\u0007\u0013\u0005\n\u0005 , respectively. Obviously, these are not the best dis-\n\n\u0005\u0015\u0014\n\u0001\u0004\u0003\u0006\u0005\u0015\u0014\n\n\u0001\u0004\u0003\u0006\u0005\n\u0001\u0004\u0003! \n\n\u0003\u0005\u0004\n\u0003\u0005\u0004\n\n\u0001\u0004\u0003\u0006\u0005\n\u0001\u0004\u0003! \n\n\u0001\u0004\u0003! \t\u0007\u0013\u0005\n\n\u0001\u0004\u0003! \n\n\u0001\u0004\u0003! \n\ncriminative densities we may think of and therefore we require an appropriate bound'\n(#*2+ .\nFor \ufb01nite')(#*,+ , maximization of the contrast enforces a redistribution of the estimated\n\nprobability mass and gives rise to a constrained linear optimization problem in the space of\ndiscriminative densities which may be solved by variational methods in some cases.\n\nwith scale factor \r\nscaling:\n\nThe relation between contrast and Bayes risk becomes more convenient when we slightly\nmodify the above de\ufb01nition (3) by a unit upper bound and by adding a lower bound on the\n\r -scaled density differences:\n\u0010\u0013\u0012\n\n\u0001\u0004\u0003;/0<\n#\u0015\u001a\n')(#*,+ . Therefore, for an in\ufb01nite scale factor \r\n\u0005\n\t approaches the negative Bayes risk up to constant shift and\n\u0003\u0006\u0005\n\u0014\u0016\u0015\u0019\u0017\nThus the scale factor de\ufb01nes a subset of the input-space, which includes the decision bound-\nary and which becomes increasingly focused in their vicinity as \r\n. The extent of the\n\r on the difference between discriminative densities.\nregion is de\ufb01ned by the bounds \u001b\nIn terms of the contrast function it can be de\ufb01ned as\n\ncontrast.\n\n\"\u0006\u001d\u001f$\n\u0001\u0004\u0003\n\n\u0001\u0013\u0012:7:\"\n\nthe (expected)\n\n\"87:9\n\n\u001a43\n\u000e65\n\n\u0003\u0006\u0005\n\n\u0003;/\n\n\u001a\n\n\u000f\u0013\u0016\n\n\u0010\u0011\u0010\n\n\u00070/\n\n\u00070/\n\n\u001a\u0013\u001a\n\n(4)\n\n\u000f\u000e\n\n5\u0019\u001f\n\n\u0014\u0016\u0015\u0018\u0017\n\n\u0010\u0011\u0010\n\n\u0001\u0004\u0003\n\n\u0012:7\n\n(5)\n\n\u0005\u0016 \u001f\n\n\u001a\u0011\u0014\n\n(6)\n\n\u0001\u0004\u0003\n\n\u0007\u001e\u001d\u0013\u0007!\u0007% \n\n\u00070/\n\n\u00070/\n\naverage of.\n\nSince for MCC-training we maximize the empirical contrast, i.e. the corresponding sample\n\u0005 , the scale factor then de\ufb01nes a subset of the training data which has\nimpact on learning of the decision boundary. Thus for increasing scale factor the relative\nsize of that subset is shrinking. However for increasing size of the training set the scale fac-\ntor can be gradually increased and then, for suitable classes of PDFs, MCCs can approach\nthe Bayes rule. In other words, \r acts as a regularization parameter such that, for particular\nchoices of the PDF class convergence to the Bayes classi\ufb01er can be achieved if the quality\nof the approximation of the loss function is gradually increased for increasing sample sizes.\nIn the following section we shall consider such a class of PDFs which is \ufb02exible enough\nand which turns out to include a certain kind of SVMs.\n\n\u0012\n\u0001\n\u001f\n\u0010\n\u0016\n\u0001\n\u0005\n\u001f\n\u0010\n\u0012\n\u0001\n\u0016\n\u001f\n\u0010\n\u0005\n\u001f\n\u0001\n\u0001\n\u0012\n\u0001\n\u001f\n\u0010\n\u0005\n\u0016\n\u001f\n\u0001\n\u0003\n\u0005\n\u001f\n\u0010\n\u0012\n\u0001\n\u001f\n\u0005\n\u0016\n\u001f\n\u0010\n\u0005\n\u0005\n\u001f\n\u0002\n\u0003\n\u0014\n\u001f\n\u0010\n\u0001\n\u0018\n\u0001\n\u0003\n\u0016\n\u0003\n\u0010\n\u0005\n\u0012\n\n\u001f\n\u0002\n\u0001\n\u0018\n\u0001\n\u0003\n\u0016\n\u0003\n\u0002\n\u0005\n\u0010\n\u0001\n\u001d\n\u0001\n\u001f\n\u0010\n\u0005\n\u0016\n\u001f\n\u0001\n\u0003\n\u0005\n\u0002\n\u0001\n\u001d\n\u000b\n\u0001\n\u001f\n\u0016\n\u001f\n\u0010\n\u0005\n.\n\u0012\n\n\u0005\n\u0001\n\u0010\n\u0018\n\u0016\n\u0010\n\n\u000f\n\u0010\n\u0012\n\u0005\n#\n\u0001\n<\n5\n\u0005\n\u0016\n\u001f\n\u001a\n\u0005\n\t\n\u0001\n\u0001\n\n\u0005\n\u0001\n.\n\u0012\n\u0007\n/\n\n\n\u0010\n\u0007\n\u0016\n\u0010\n\u0007\n.\n\u0001\n\u0003\n\u0012\n\n\u0005\n\t\n\u0001\n\u0010\n\u0007\n\u0016\n\u0010\n\u0007\n\"\n.\n\u0001\n\n\u0005\n\u0014\n\u001c\n\u0001\n\u000f\n\u0003\n.\n\u0012\n\n\u0010\n\u0001\n\u0003\n\u0012\n\n\f3 MCC-Realizations\n\nIn the following we shall \ufb01rst consider a particularly useful parametrization of the dis-\ncriminative densities which gives rise to classi\ufb01ers which in the binary case have the same\nfunctional form as SVMs up to a \u201cmissing\u201d bias term in the MCC-case. For training of\nthese MCCs we derive a suitable objective function which can be maximized by sequential\nlinear programming where we show the close relation to training of Linear Programming\nMachines.\n\n3.1 Density Parametrization\n\nWe \ufb01rst have to choose a set of candidate functions from which we select the required PDF.\nBecause this set should provide some \ufb02exibility with respect to contrast maximization the\nusual kernel density estimator (KDE)[11]\n\n\u0001\u0004\u0003! \n\n\u0001\u0004\u0003\n\n(7)\n\n$\u0002\u0001\u0004\u0003\u0002\u0005\n\n\u0001\u0004\u0003\n\nwith index set\nfunctions according to \u0007\n\n\u001a containing indices of examples from class \u001d and with normalized kernel\n\nisn\u2019t a quite good choice, since the only free\nparameter is the kernel bandwidth which doesn\u2019t allow for any local adaptation. On the\nother hand if we allow for local variation of the bandwidth we get a complicated contrast\nwhich is dif\ufb01cult to maximize due to nonlinear dependencies on the parameters. The same\nis true if we treat the kernel centers as free parameters. However, if we modify the kernel\ndensity estimator to have \ufb02exible mixing weights according to\n\n\u00052\u0014\u0013\u0003\n\n\u0001\u0004\u0003\n\n\u0001\u0004\u0003;/\n\n\u0010\u0013\u0012\nwe get an objective function, which is linear in the mixing parameters \n\nconditions. Thus we have class-speci\ufb01c densities with mixing weights \n\nthe contribution of a single training example to the PDF.\n\n\u0001\u0004\u0003\u0006\u0005 with \u0011\n\n\u0001\f\u000b\u000e\r\n\u001a\u0010\u000f\n\n$\b\u0001\u0004\u0003\t\u0005\n\n$\u001c\u001a\n\nWith that choice we achieve plug-in Bayes-consistency for the case of equal mixing\nweights, since then we have the usual kernel density estimator (KDE), which, besides\nsome mild assumptions about the distributions, requires a vanishing kernel bandwidth for\n\n(8)\n\n\u001a\u0013\u0012\u0015\u0014\n$\u001b\u001a under certain\n$\u001c\u001a which control\n\n\u001a\n\n.\n\n\u0003;/\n\nexamples, as:\n\n\u001d\u0013\u0012! \n$\u001c\u001a\n\n\u0012\u0015\u0014\u0015\u0014\u0016\u0014\u0015\u0012\n\u0001\u0004\u0003\u0006\u0005\n\n. Further we de\ufb01ne the scaled density difference\n\n3.2 Objective Function\nFor notational simplicity in the following we shall incorporate the scale factor \r and the\nmixing weigths\u000b\nand \u0011\u0018\u0017\n\u0011\u0018\u0017\u000e\u0019\u001a\u0011\nso that we can write the empirical contrast.$#\n\n with\u0017\ninto a common parameter vector\u0017\n\u0010\u001c\u001b\u0015\u001d\u001f\u001e\n\u0001\u0004\u0003\u0006\u0005\n\u0005 , i.e. the sample average over\u0016\n\u0010-,\n\"87:9\n\n$\u001f)\n$/.\nis concave and maximization with respect to\u0017 gives rise\n\n&('\n$\u0002\u0001\u0004\u0003\n\u001a,\u000e\u0011\u0010\nwhere the assignment variables '\n\ufb01xed assignment variables'\n$ ,.%#\n\nrealize the maximum function in (4). With\n\ntraining\n\n.%#\n\n\u001a+*\n\n(10)\n\n\u000f\u0011\u0010\n\n\u0001\u0004\u0003\n\n(9)\n\n\u0012\u0016\u0010\n\n\u001f\n\u001d\n\u0005\n\u0001\n\u0010\n \n\n\u001a\n \n\n\u0006\n\u0012\n\u0003\n$\n\u0005\n\u0006\n\u0012\n\u000e\n\u0001\n\u0010\n\u001f\n<\n\u001a\n\u0005\n\u0001\n\n\u0006\n\u0012\n\u0003\n$\n\u0005\n\u001a\n\u000b\n\u001a\n\u0011\n\u0010\n\u0001\n\u000b\n\u0016\n\u001a\n\u0001\n\u0001\n\u0017\n\n\u0010\n\u0017\n\n\u000b\n\u0005\n\u001a\n\u0001\n\n\u000b\n\u001a\n\u001a\n\u0011\n\u0010\n\u0001\n\"\n\u0001\n\u0017\n\u0005\n\u0001\n#\n$\n\u0017\n\n$\n\u000f\n$\n\u0016\n#\n\u001a\n\u0017\n\n\u001a\n\u000f\n\u001a\n\u0014\n\u0001\n\u0017\n\u0001\n\u0017\n\u0005\n\u0001\n\u0010\n\u0016\n\u000b\n\n\u0005\n\u0010\n\u0004\n\u0010\n\u0018\n\u0016\n\u0010\n\n\u0019\n3\n\u000e\n\u001a\n\u0012\n\"\n\u001a\n\u0019\n$\n/\n\u0017\n\u0005\n\u0016\n\u000f\n\u001d\n\u001a\n\fis\n\n, \n\nobjective function can be de\ufb01ned by\n\n$ is achieved by setting'\n\nfor negative terms. This suggests a sequential linear\noptimization strategy for overall maximization of the contrast which shall be introduced in\ndetail in the following section.\n\nto a linear optimization problem. On the other hand, for \ufb01xed \u0017 maximization with respect\nto the'\nSince we have already incorporated \r as a scaling factor into the parameter vector \u0017\n\u0010 . Therefore the scale factor can be adjusted implicitly\nnow identi\ufb01ed with the norm \u0011\u0018\u0017\nby a regularization term which penalizes some suitable norm of the \u0017\n\u001a . Thus a suitable\n.%#\n.%#\n(11)\nwith \u0001 determining the weight of the penalty, i.e. the degree of regularization. We now\nconsider several instances of the case where the penalty corresponds to some \u001f -norm of\u0017\n.\nWith the \u0010 -norm, for \u0001\nthe probability mass of the discriminative densities is con-\ncentrated on those two kernel-functions which yield the highest average density difference.\nAlthough that property forces the sparsest solution for large enough \u0001\n, clearly, that solution\nisn\u2019t Bayes-consistent in general because as pointed out in Sec.2, for \u0001\nall proba-\nbility mass of the discriminative densities is concentrated at the two points with maximum\naverage density difference.\nConversely taking \u0002\n\u0002 , which resembles the standard SVM regularizer [10],\n\u0011\u0018\u0017/\u0011\nyields the KDE with equal mixing weights for \u0001\n. Indeed, it is easy to see that all\n\u0010 share this convenient property, which guarantees \u201cplug-in\u201d\n\u001f -norm penalties with \u001f\nBayes consistency in the case where the solution is totally determined by the regularizer.\nIn that case kernel density estimators are achieved as the \u201cdefault\u201d solution. Therefore we\nchose a combination of the \u0010 -norm with the maximum-norm\n\n\u0001\u0003\u0002\n\n(12)\n\n\u0011-\u0017\n\n\u001a,\u000e\u0011\u0010\n\n\u0011\u0018\u0017\nwhich is easily incorporated into a linear program, as to be shown in the following. For that\n we achieve an equal distribution of the weights\nkind of penalty in the limiting case \u0001\nwhich corresponds to the kernel density estimator (KDE) solution. In that way we have a\nnice trade-off between two kinds of Bayes consistency: for increasing \u0001\nthe class-speci\ufb01c\nthe\ndensities converge to the KDE with equal mixing weights, whereas for decreasing \u0001\nprobability mass of the discriminative densities is more and more concentrated near the\nBayes-optimal decision boundary. By a suitable choice of the kernel width and the scale\nof the weights, e.g. via cross-validation, the solution with fastest convergence to the Bayes\nrule may be selected.\n\nWith an 1-norm penalty on the weights and on the vector\nwe get the Linear Programming Machine which requires to minimize\n\nof soft margin slack variables\n\n\u0011-\u0017\n\u000f\u0013\u0016\nwith \u0007\nsubtracting\u0016\n\nsubject to\n\n\u0004\u0006\u0005\n\u001a\t\b6\u001a\n\u0011\u0007\u0004\n\u001a and with the above constraints on \u0017\n\u0010\u0013\u0012\u0016\u0010\n, setting \u0010\u0015\u0016\n\n. Dividing the objective by \u0005\nobjective shows that LPM training corresponds to a special case of MCC training with \ufb01xed\n\n'\u0011$ and turning minimization to maximization of the negative\n\n\u0010 and \u0010 -norm regularizer with \u0001\n\n\u0012\f\n\n(13)\n\n\u001a,\u000e\u0011\u0010\n\n\u0016\u000b\n\n\u0001\u0004\u0003\n\n$\u000f\u000e\n\n\u0010\u0011\u0010\n\n.\n\n,\n\n3.3 Sequential Linear Programming\n\nEstimation of mixing weights is now achieved by maximizing the sample contrast with\nrespect to the\niterative optimization scheme:\n\n$ . This can be achieved by the following\n\n$\u001c\u001a and the assignment variables'\n\n$\n\u0001\n\u001d\n\u001a\n\u0011\n\u0001\n\u0017\n\u0012\n\u0001\n\u0005\n\u0001\n\u0001\n\u0017\n\u0005\n\u0016\n\u0001\n\u0017\n\u0005\n\u0012\n\u0001\n\u001b\n\u001d\n\n\n\n\n\u0001\n\u0017\n\u0005\n\u0001\n\u0002\n\n\n\u001b\n\u0002\n\u0001\n\u0017\n\u0005\n\u0001\n\u000b\n\n\u0001\n\u001a\n\u0011\n\u0010\n\u0004\n\u001a\n\u0011\n\u0017\n\u0005\n\n\u0004\n\u0011\n\u0010\n\u0011\n\u0010\n\u0007\n$\n#\n\n\u0007\n\u0006\n$\n\u0012\n\u0003\n\u001a\n\u0005\n\u0012\n\u0010\n$\n$\n\u0012\n\u001d\n.\n'\n$\n\u0001\n\u0001\n\u0005\n\b\n\ffor \ufb01xed\n\n1. Initialization:'\n2. Maximization w.r.t. \u0017\nmaximize \u000b\n$\b\u0001\u0004\u0003\n\u001a,\u000e\u0011\u0010\nsubject to'\u0011$\u001c\u001a\n\u0019\u0006\u0005\nfor \ufb01xed\u0017\n\u0012\r\f\n\u0001\u000b\n\n3. Maximization w.r.t.\n\n:\n\n\u000e0\u001a\n\n:\n\n$\u001b\u001a\n\u001d\u0013\u0012! \n\notherwise.\n\n\u001a43\n\u000e65\u000f\u000e\n\n$\u001c\u001a\n\n$\b\u0001\u0004\u0003\n\n \b\u0007\n\n\u0012\u001e\u001d\n\n\u0012\u0015\u001d\n\n\u001a\u0004\u0003\n(#*,+\n\b\u0002\u0001\n'\u0011$\u001c\u001a\n$\u001b\u001a\n\n\u001a,\u000e\u0011\u0010\n\u001a\u0004\u0003\n(#*,+\n\b\t\u0001\n\n\u0001\u0004\u0003\n\n4. If convergence in contrast then stop else proceed with step 2.\n\nWhere'\u0011$\u001c\u001a\nwhich can be charged to the objective function. The constraint \u0011\u0018\u0017\nprogram was chosen in order to prevent the trivial solution \u0017\n\n\u0019 are slack variables, measuring the part of the density difference \"\n\u0014 which may otherwise\n. Since we used unnormalized Gaussian kernel functions with\n\u0010 , i.e. we excluded all multiplicative density constants, that constraint doesn\u2019t\n\nappear for larger values of \u0001\nexclude any useful solutions for the weights.\n\nin the linear\n\n\u0003\u0006\u0005\n\n4 Experiments\n\nIn the following section we consider the task of solving binary classi\ufb01cation problems\nwithin the MCC-framework, using the above SLPM with Gaussian kernel function. The\n\ufb01rst experiment illustrates the behaviour of the MCC for different values for the regu-\nlarization \u0001 by means of a simple two-dimensional toy dataset. The second experiment\ncompares the classi\ufb01cation performance of the MCC with those of the SVM and Kernel-\nDensity-Classi\ufb01er (KDC) which is a special case of the MCC with equal weighting of each\nkernel function. To this end, we selected four frequently used benchmark datasets from the\nUCI Machine Learning Repository.\n\n\u0014\u0011\u0010 and standard deviation\n\nisotropic normal distributions with a mutual distance of\u0014\n\u001d\u0015\u0014\u0017\u0016\nthere are regions of high contrast \n\nThe two-dimensional toy dataset consists of 300 data points, sampled from two overlapping\n.\n(only data points\nFigure 1 shows the solution of the MCC for two different values of \u0001\nare marked by symbols). In\nwith non-zero weights according the criterion\nboth \ufb01gures, data points with large mixing weights are located near the decision border. In\n alongside the decision function\nparticular for small \u0001\nin-\n(illustrated by isolines). For increasing \u0001\ncreases. At the same time, one can note a decrease of the difference between the weights.\nRegions with contrast \n, these regions\nare nearer to the decision border than for large values. This illustrates that for increasing\nthe quality of the approximation of the loss function decreases. In both \ufb01gures, several\n\u0010 . The MCC identi\ufb01ed those data points\ndata points are misclassi\ufb01ed with a contrast \nas outliers and deactivated them during the training (encircled symbols).\n\n\u0010 are highlighted gray. For small values of \u0001\n\nthe number of data points with non-zero\n\nThe second experiment demonstrates the performance of the MCC in comparison with\nthose of a Support Vector Machine, as one of the state-of-the-art binary classi\ufb01ers, and\nwith the KDC. For this experiment we selected the Pima Indian Diabetes, Breast-Cancer,\nHeart and Thyroid dataset from the UCI Machine Learning repository. The Support Vector\nMachine was trained using the Sequential Minimal Optimization algorithm by J. Platt[7]\nadjusted according to the modi\ufb01cation proposed by S.S. Keerthi [5].\n\n$\n\u0001\n\u0010\n\u001e\n\u001c\n\n\n\u0005\n'\n$\n\n\u0019\n3\n'\n\u0019\n\u0016\n\u0001\n\u000b\n\n\u0001\n\u0004\n\n\u0005\n\b\n\u0005\n\u0010\n\u0012\n\"\n\u001a\n\u0019\n\u0001\n\u0003\n$\n\u0012\n\u0017\n\u0005\n\u0016\n\u0019\n\u0012\n\u0001\n\u001d\n\u0012\n\u0011\n\b\n\u001a\n\u0011\n\u0010\n\u0001\n\u0011\n\b\n\u0019\n\u0011\n\u0010\n\u0012\n\u0010\n\u001e\n\u0012\n\b\n\u0012\n\u001d\n\u001e\n\u001c\n\n'\n$\n\u0010\n\"\n\u001a\n\u0019\n$\n\u0012\n\u0017\n\u0005\n\u001b\n\u0010\n\u0016\n\u0018\n\u001d\n\u0012\n\u001a\n\u0019\n\u0001\n\u0003\n$\n\u0012\n\u0017\n\u0005\n\u001a\n\u0011\n\u0010\n\u0012\n\u0010\n\u0001\n\u0006\n\u0001\n\u0003\n\u0012\n\u0001\n\u0001\n\u0007\n\u0012\n\u0013\n\b\n$\n\u001b\n\u0010\n.\n\b\n$\n.\n \n\u0012\n\u0001\n.\n \n\u0012\n\f 300 datapoints / l = 0.2\n\n 300 datapoints / l = 4.2\n\n0 . 5\n\n\u2212\n\n\u22121\n\n\u22121.5\n\n\u22122\n\n\u2212 2 . 5\n\n0\n\n2.5\n2\n\n1.5\n\n\u2212 0.5\n\n1\n\n\u2212\n\n\u22121.5\n\n\u22122\n\n0\n\n0.5\n\n1\n.\n5\n\n\u22120.5\n\n0\n\n1\n\n0.5\n\n0\n\n1\n\n0.5\n\nand\n\n\u0007 ). The symbols\n\nFigure 1: Two MCC solutions for the two-dimensional toy dataset for different values of \u0001\n\u0007 , right: \u0001\n\u0001\u0001\ndepict the positions of data points\n(left: \u0001\nwith with non-zero\n$ . The size of each symbol is scaled according the value of the cor-\nresponding\n$ . Encircled symbols have been deactivated during the training (symbols for\nis zero). The\ndeactivated data points are not scaled according to\nabsolute value of the contrast is illustrated by the isolines while the sign of the contrast de-\n\u0010 which corresponds\npicts the binary classi\ufb01cation of the classi\ufb01er. The region with \nto \u001c\nas de\ufb01ned in (6) is colored white and the complement colored gray. The percentage\n(left \ufb01gure) and \u0007\n(right \ufb01gure) of the\ndataset.\n\n$ , since in most cases\n\n\u0014\u0005\u0004\u0007\u0006\n\nof data points that de\ufb01ne the solution is \u0010\n\n:\n\nThe experimental setup was comparable with that in [8]: After normalization to zero\nmean and unit standard deviation, each dataset was divided 100 times in different\n(provided by G. R\u00a8atsch at\npairs of disjoint train- and testsets with a ratio of\nraetsch/data/benchmarks.htm). Since we used for all classi\ufb01ers the\nhttp://ida.\ufb01rst.gmd.de/\n. Addi-\nGaussian kernel function, all three algorithms are parametrized by the bandwidth\ntionally, for the SVM and MCC the regularization value \u0001 had to be chosen. The optimal\nparametrization was chosen by estimating the generalization performance for different val-\nues of bandwidth and regularization by means of the average test error on the \ufb01rst \ufb01ve\ndataset partitions. More precisely, a \ufb01rst coarse scan was performed, followed by a \ufb01ne\nscan in the interval near the optimal values of the \ufb01rst one. Each scan considered 1600\n. For parameter pairs with identical test\ndifferent combinations of\nerror, the pair constructing the sparsest solution was kept. Finally, the reported values in\nTab.1 and Tab.2 are averaged over all 100 dataset partitions.\n\u0005 of the MCC in combination with the\nTable 1 shows the optimal parametrization \u0001\n$ ).\nclassi\ufb01cation rate and sparseness of the solution (measured as percentage non-zero\nAdditionally, the corresponding values after the \ufb01rst MCC iteration are given in brackets.\nThe last two columns show the absolute number of iterations and the \ufb01nal number of de-\nactivated examples. For all four datasets the MCC is able to \ufb01nd a sparse solution. In\nparticular for the Heart, Breast-Cancer and Diabetes dataset the solution of the MCC is\nsigni\ufb01cantly sparser than those of the SVM (see Tab.2). Nevertheless, Tab.2 indicates that\nthe classi\ufb01cation rates of the MCC are competitive with those of the SVM.\n\nand \u0001\n\nand \u0005\n\n, resp.\n\n5 Conclusion\n\nThe MCC-approach provides an understanding of SVMs / LPMs in terms of generative\nmodelling using discriminative densities. While usual unsupervised density estimation\nschemes try to minimize some distance criterion (e.g. Kullback-Leibler divergence) be-\n\n\u0001\n\u001d\n\u0014\n\u0014\n\u0002\n\u0003\n\b\n\b\n\b\n\b\n$\n.\n \n\u001f\n\u001d\n\u0006\n\b\n\u001d\n\u0006\n\u001d\n\u0006\n\t\n\n\u0012\n\u0001\n\b\n\fTable 1: Optimal parametrization \u0001\nnumber of iterations of the MCC and number of '\n100 dataset partitions. For the classi\ufb01cation rate and percentage of non-zero\nthe corresponding value after the \ufb01rst MCC iteration is given in brackets.\n\n\u0005 , classi\ufb01cation rate, percentage of non-zero\n$ ,\n\u001d . The results are averaged over all\n\n-coef\ufb01cients\n\nDataset\n\nBreast-Cancer\n\nHeart\nThyroid\nDiabetes\n\n1.38\n2.69\n0.49\n4.52\n\n12.17\n2.066\n\n\u0007\u0015\t\n\u000b\u000e\u0016\n2.624\n\nClassif. rate\n(74.4\n(84.1\n(95.5\n(76.5\n\n74.3\n84.3\n95.5\n76.6\n\n)\n)\n)\n)\n\n\u0002\u0004\u0003\u0006\u0005\b\u0007\n\t\f\u000b\u000e\r\n(13.8\n(21.2\n(46.1\n(5.5\n\n13.6\n20.4\n46.1\n5.3\n\n)\n\n)\n)\n)\n\nIter.\n\n2.23\n3.10\n1.00\n5.86\n\n\u0002\u0010\u000f\u0012\u0011\u0013\t\n2.6\n6.4\n0.0\n40.7\n\nTable 2: Summary of the performance of the KDC, SVM and MCC for the four benchmark\ndatasets. Given are the classi\ufb01cation rates with percentage of non-zero\nNote that our results for the SVM are slightly better to those reported in [8]. One reason\ncould be the coarse parameter selection for the SVM as already mentioned by the author.\n\n$ (in brackets).\n\nDataset\n\nBreast-Cancer\n\nHeart\nThyroid\nDiabetes\n\nKDC\n\n73.1\n84.1\n95.6\n74.2\n\n(100\n(100\n(100\n(100\n\n)\n)\n)\n)\n\nSVM\n\n74.5\n84.4\n95.7\n76.7\n\n(58.5\n(60.9\n(15.8\n(53.6\n\n)\n)\n)\n)\n\nMCC\n\n74.3\n84.3\n95.5\n76.6\n\n(13.6\n(20.4\n(46.1\n( 5.3\n\n)\n)\n)\n)\n\ntween the models and the true densities, MC-estimation aims at learning of densities which\nrepresent the differences of the underlying distributions in an optimal way for classi\ufb01ca-\ntion. Future work will address the investigation of the general multiclass performance and\nthe capability to cope with misslabeled data.\n\nReferences\n\n[1] C. M. Bishop. Neural Networks for Pattern Recognition. Clarendon Press, Oxford, 1995.\n[2] C. Cortes and V. Vapnik. Support-vector networks. Machine Learning, 20(3):273\u2013297, 1995.\n[3] R. O. Duda and P. E. Hart. Pattern Classi\ufb01cation and Scene Analysis. Wiley, New York, 1973.\n[4] T. Graepel, R. Herbrich, B. Scholkopf, A. Smola, P. Bartlett, K. Robert-Muller, K. Obermayer,\n\nand B. Williamson. Classi\ufb01cation on proximity data with lp\u2013machines, 1999.\n\n[5] S.S. Keerthi, S.K. Shevade, C. Bhattacharyya, and K.R.K. Murthy.\n\nImprovements to platt\u2019s\nSMO algorithm for SVM classi\ufb01er design. Technical report, Dept of CSA, IISc, Bangalore,\nIndia, 1999.\n\n[6] P. Meinicke, T. Twellmann, and H. Ritter. Maximum contrast classi\ufb01ers. In Proc. of the Int.\n\nConf. on Arti\ufb01cial Neural Networks, Berlin, 2002. Springer. in press.\n\n[7] J. Platt. Fast training of support vector machines using sequential minimal optimization. In\nB. Sch\u00a8olkopf, C. J. C. Burges, and A. J. Smola, editors, Advances in Kernel Methods \u2014 Support\nVector Learning, pages 185\u2013208, Cambridge, MA, 1999. MIT Press.\n\n[8] G. R\u00a8atsch, T. Onoda, and K.-R. M\u00a8uller. Soft margins for AdaBoost. Technical Report NC-TR-\n1998-021, Department of Computer Science, Royal Holloway, University of London, Egham,\nUK, August 1998. Submitted to Machine Learning.\n\n[9] B. D. Ripley. Pattern Recognition and Neural Networks. Cambridge University Press, Cam-\n\nbridge, 1996.\n\n[10] B. Sch\u00a8olkopf and A. J. Smola. Learning with Kernels. MIT Press, 2002.\n[11] D. W. Scott. Multivariate Density Estimation. Wiley, 1992.\n[12] V. N. Vapnik. The Nature of Statistical Learning Theory. Springer, New York, 1995.\n\n\u0012\n\u0001\n\b\n$\n\u0001\n\b\n\n\u0001\n\u0002\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\b\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\u0014\n\f", "award": [], "sourceid": 2336, "authors": [{"given_name": "Peter", "family_name": "Meinicke", "institution": null}, {"given_name": "Thorsten", "family_name": "Twellmann", "institution": null}, {"given_name": "Helge", "family_name": "Ritter", "institution": null}]}