{"title": "Ex ante coordination and collusion in zero-sum multi-player extensive-form games", "book": "Advances in Neural Information Processing Systems", "page_first": 9638, "page_last": 9648, "abstract": "Recent milestones in equilibrium computation, such as the success of Libratus, show that it is possible to compute strong solutions to two-player zero-sum games in theory and practice. This is not the case for games with more than two players, which remain one of the main open challenges in computational game theory. This paper focuses on zero-sum games where a team of players faces an opponent, as is the case, for example, in Bridge, collusion in poker, and many non-recreational applications such as war, where the colluders do not have time or means of communicating during battle, collusion in bidding, where communication during the auction is illegal, and coordinated swindling in public. The possibility for the team members to communicate before game play\u2014that is, coordinate their strategies ex ante\u2014makes the use of behavioral strategies unsatisfactory. The reasons for this are closely related to the fact that the team can be represented as a single player with imperfect recall. We propose a new game representation, the realization form, that generalizes the sequence form but can also be applied to imperfect-recall games. Then, we use it to derive an auxiliary game that is equivalent to the original one. It provides a sound way to map the problem of finding an optimal ex-ante-correlated strategy for the team to the well-understood Nash equilibrium-finding problem in a (larger) two-player zero-sum perfect-recall game. By reasoning over the auxiliary game, we devise an anytime algorithm, fictitious team-play, that is guaranteed to converge to an optimal coordinated strategy for the team against an optimal opponent, and that is dramatically faster than the prior state-of-the-art algorithm for this problem.", "full_text": "Exantecoordinationandcollusioninzero-summulti-playerextensive-formgamesGabrieleFarina\u2217ComputerScienceDepartmentCarnegieMellonUniversitygfarina@cs.cmu.eduAndreaCelli\u2217DEIBPolitecnicodiMilanoandrea.celli@polimi.itNicolaGattiDEIBPolitecnicodiMilanonicola.gatti@polimi.itTuomasSandholmComputerScienceDepartmentCarnegieMellonUniversitysandholm@cs.cmu.eduAbstractRecentmilestonesinequilibriumcomputation,suchasthesuccessofLibratus,showthatitispossibletocomputestrongsolutionstotwo-playerzero-sumgamesintheoryandpractice.Thisisnotthecaseforgameswithmorethantwoplayers,whichremainoneofthemainopenchallengesincomputationalgametheory.Thispaperfocusesonzero-sumgameswhereateamofplayersfacesanopponent,asisthecase,forexample,inBridge,collusioninpoker,andmanynon-recreationalapplicationssuchaswar,wherethecolludersdonothavetimeormeansofcom-municatingduringbattle,collusioninbidding,wherecommunicationduringtheauctionisillegal,andcoordinatedswindlinginpublic.Thepossibilityfortheteammemberstocommunicatebeforegameplay\u2014thatis,coordinatetheirstrategiesexante\u2014makestheuseofbehavioralstrategiesunsatisfactory.Thereasonsforthisarecloselyrelatedtothefactthattheteamcanberepresentedasasingleplayerwithimperfectrecall.Weproposeanewgamerepresentation,therealizationform,thatgeneralizesthesequenceformbutcanalsobeappliedtoimperfect-recallgames.Then,weuseittoderiveanauxiliarygamethatisequivalenttotheoriginalone.Itprovidesasoundwaytomaptheproblemof\ufb01ndinganoptimalex-ante-coordinatedstrategyfortheteamtothewell-understoodNashequilibrium-\ufb01ndingproblemina(larger)two-playerzero-sumperfect-recallgame.Byreasoningovertheauxiliarygame,wedeviseananytimealgorithm,\ufb01ctitiousteam-play,thatisguaranteedtoconvergetoanoptimalcoordinatedstrategyfortheteamagainstanoptimalopponent,andthatisdramaticallyfasterthanthepriorstate-of-the-artalgorithmforthisproblem.1IntroductionInrecentyears,computationalstudiesonimperfect-informationgameshavelargelyfocusedontwo-playerzero-sumgames.Inthatsetting,AItechniqueshaveachievedremarkableresults,suchasdefeatingtophumanspecialistprofessionalsinheads-upno-limitTexashold\u2019empoker[4,5].Fewerresultsareknownforsettingswithmorethantwoplayers.Yet,manystrategicinteractionsprovideplayerswithincentivestoteamup.Insomecases,playersmayhaveasimilargoalandmaybewillingtocoordinateandsharetheir\ufb01nalreward.Consider,asanillustration,thecaseofapoker\u2217Equalcontribution.32ndConferenceonNeuralInformationProcessingSystems(NeurIPS2018),Montr\u00b4eal,Canada.\fgamewiththreeormoreplayers,whereallbutoneofthemcolludeagainstanidenti\ufb01edtargetplayerandwillsharethewinningsafterthegame.Inothersettings,playersmightbeforcedtocooperatebythenatureoftheinteractionitself.Thisisthecase,forinstance,inthecard-playingphaseofBridge,whereateamoftwoplayers,calledthe\u201cdefenders\u201d,playsagainstathirdplayer,the\u201cdeclarer\u201d.Situationsofateamganginguponaplayerare,ofcourse,ubiquitousinmanynon-recreationalapplicationsaswell,suchaswarwherethecolludersdonothavetimeormeansofcommunicatingduringbattle,collusioninbiddingwherecommunicationduringtheauctionisillegal,coordinatedswindlinginpublic,andsoon.Thebene\ufb01tsfromcoordination/collusiondependonthecommunicationpossibilitiesamongteammembers.Inthispaper,weareinterestedinexantecoordination,wheretheteammembershaveanopportunitytodiscussandagreeontacticsbeforethegamestarts,butwillbeunabletocommunicateduringthegame,exceptthroughtheirpublicly-observedactions.2Theteamfacesanopponentinazero-sumgame(asin,forexample,multi-playerpokerwithcollusionandBridge).Evenwithoutcommunicationduringthegame,theplanningphasegivestheteammembersanad-vantage:forinstance,theteammemberscouldskewtheirstrategiestousecertainactionstosignalabouttheirstate(forexample,thattheyhaveparticularcards).Inotherwords,byhavingagreedoneachmember\u2019splannedreactionunderanypossiblecircumstanceofthegame,informationcanbesilentlypropagatedintheclear,bysimplyobservingpublicinformation.Exantecoordinationcanenabletheteammemberstoobtainsigni\ufb01cantlyhigherutility(uptoafactorlinearinthenumberofthegame-treeleaves)thantheutilitytheywouldobtainbyabstainingfromcoordination[1,6,25].FindinganequilibriumwithexantecoordinationisNP-hardandinapproximable[1,6].Theonlyknownalgorithmisbasedonahybridrepresentationofthegame,whereteammembersplayjointnormal-formactionswhiletheadversaryemployssequence-formstrategies[6].Wewilldevelopdramaticallyfasteralgorithmsinthispaper.Ateamthatexantecoordinatescanbemodeledasasinglemeta-player.Thismeta-playertypicallyhasimperfectrecall,giventhattheteammembersobservedifferentaspectsoftheplay(opponent\u2019smoves,eachothers\u2019moves,andchance\u2019smoves)andcannotcommunicateduringthegame.Then,solvingthegameamountstocomputingaNashequilibrium(NE)innormal-formstrategiesinatwo-playerzero-sumimperfect-recallgame.Thefocusonnormal-formstrategiesiscrucial.Indeed,itisknownthatbehavioralstrategies,thatprovideacompactrepresentationoftheplayers\u2019strategies,cannotbeemployedinimperfect-recallgameswithoutincurringalossofexpressiveness[18].Someimperfect-recallgamesdonotevenhaveanyNEinbehavioralstrategies[27].EvenwhenaNEinbehavioralstrategiesexists,itsvaluecanbeuptoalinearfactor(inthenumberofthegame-treeleaves)worsethanthatofaNEinnormal-formstrategies.Forthesereasons,recentef\ufb01cienttechniquesforapproximatingmaxminbehavioralstrategypro\ufb01lesinimperfect-recallgames[9,7]arenotapplicabletoourdomain.Maincontributionsofthispaper.Our\ufb01rstcontributionisanewgamerepresentation,whichwecalltherealizationform.Inperfect-recallgamesitessentiallycoincideswiththesequenceform,but,unlikethesequenceform,itcanalsobeusedinimperfect-recallgames.Byexploitingtherealizationform,weproduceatwo-playerauxiliarygamethathasperfectrecall,andisequivalenttothenormalformoftheoriginalgame,butsigni\ufb01cantlymoreconcise.Furthermore,weproposeananytimealgorithm,\ufb01ctitiousteam-play,whichisavariationof\ufb01ctitiousplay[2].Itisguaranteedtoconvergetoanoptimalsolutioninthesettingwheretheteammemberscoordinateexante.Experimentsshowthatitisdramaticallyfasterthanthepriorstate-of-the-artalgorithmforthisproblem.2PreliminariesInthissectionweprovideabriefoverviewofextensive-formgames(seealsothetextbookbyShohamandLeyton-Brown[20]).Anextensive-formgame\u0393hasa\ufb01nitesetPofplayersanda\ufb01nitesetofactionsA.Histhesetofallpossiblenodes,describedassequencesofactions(histo-ries).A(h)isthesetofactionsavailableatnodeh.Ifa\u2208A(h)leadstoh0,wewriteha=h0.2Thiskindofcoordinationhassometimesbeenreferredtoasexantecorrelationamongteammembers[6].However,wewillnotusethattermbecausethissettingisquitedifferentthantheusualnotionofcorrelationingametheory.Intheusualcorrelationsetting,theindividualplayershavetobeincentivizedtofollowtherecommendationsofthecorrelationdevice.Incontrast,herethereisnoneedforsuchincentivesbecausethemembersoftheteamcansharethebene\ufb01tsfromcoordination.2\fP(h)\u2208P\u222a{c}istheplayerwhoactsath,wherecdenoteschance.Hiisthesetofdecisionnodeswhereplayeriacts.Zisthesetofterminalnodes.Foreachplayeri\u2208P,thereisapayofffunctionui:Z\u2192R.Anextensive-formgamewithimperfectinformationhasasetIofinformationsets.Decisionnodeswithinthesameinformationsetarenotdistinguishablefortheplayerwhoseturnitistomove.Byde\ufb01nition,foranyI\u2208I,A(I)=A(h),forallh\u2208I.IiistheinformationpartitionofHi.Apurenormal-formplanforplayeriisatuple\u03c3\u2208\u03a3i=\u00d7I\u2208IiA(I)thatspeci\ufb01esanactionforeachinformationsetofthatplayer.\u03c3(I)denotestheactionselectedin\u03c3atinformationsetI.Anormal-formstrategyxiforplayeriisde\ufb01nedasxi:\u03a3i\u2192\u2206|\u03a3i|.WedenotebyXithenormal-formstrategyspaceofplayeri.Abehavioralstrategy\u03c0i\u2208\u03a0iassociateseachI\u2208IiwithaprobabilityvectoroverA(I).\u03c0i(I,a)denotestheprobabilitywithwhichichoosesactionaatI.\u03c0cisthestrategyofavirtualplayer,\u201cchance\u201d,whoplaysnon-strategicallyandisusedtorepresentexogenousstochasticity.Theexpectedpayoffofplayeri,whensheplaysxiandtheopponentsplayx\u2212i,isdenoted,withanoverloadofnotation,byui(xi,x\u2212i).Denoteby\u03c1xii(z)theprobabilitywithwhichplayeriplaystoreachzwhenfollowingstrategyxi(\u03c1\u03c0ii(z)isde\ufb01nedanalogously).Then,\u03c1x(z)=Qi\u2208P\u222a{c}\u03c1xii(z)istheprobabilityofreachingzwhenplayersfollowbehavioralstrategypro\ufb01lex.Wesaythatxi,x0iarerealizationequivalentif,foranyx\u2212iandforanyz\u2208Z,\u03c1x(z)=\u03c1x0(z),wherex=(xi,x\u2212i),x0=(x0i,x\u2212i).Thesamede\ufb01ni-tionholdsforstrategiesindifferentrepresentations(e.g.,behavioralandsequenceform).Similarly,twostrategiesxi,x0iarepayoffequivalentif,\u2200j\u2208Pand\u2200x\u2212i,uj(xi,x\u2212i)=uj(x0i,x\u2212i).Aplayerhasperfectrecallifshehasperfectmemoryofherpastactionsandobservations.Formally,\u2200xi,\u2200I\u2208Ii,\u2200h,h0\u2208I,\u03c1xii(h)=\u03c1xii(h0).\u0393hasperfectrecallifeveryplayerhasperfectrecall.BR(x\u2212i)denotesthebestresponseofplayeriagainstastrategypro\ufb01lex\u2212i.Abestresponseisastrategysuchthatui(BR(x\u2212i),x\u2212i)=maxxi\u2208Xiui(xi,x\u2212i).ANE[17]isastrategypro\ufb01leinwhichnoplayercanimproveherutilitybyunilaterallydeviatingfromherstrategy.Therefore,foreachplayeri,aNEx\u2217=(x\u2217i,x\u2217\u2212i)satis\ufb01esui(x\u2217i,x\u2217\u2212i)=ui(BR(x\u2217\u2212i),x\u2217\u2212i).Thesequenceform[12,23]ofagameisacompactrepresentationapplicableonlytogameswithperfectrecall.Itdecomposesstrategiesintosequencesofactionsandtheirrealizationprobabilities.Asequenceqi\u2208Qiforplayeri,de\ufb01nedbyanodeh,isatuplespecifyingplayeri\u2019sactionsonthepathfromtheroottoh.Asequenceissaidterminalif,togetherwithsomesequencesoftheotherplayers,leadstoaterminalnode.q\u2205denotesthe\ufb01ctitioussequenceleadingtotherootnodeandqaistheextendedsequenceobtainedbyappendingactionatoq.Asequence-formstrategyforplayeriisafunctionri:Qi\u2192[0,1],s.t.ri(q\u2205)=1and,foreachI\u2208IiandsequenceqleadingtoI,\u2212ri(q)+Pa\u2208A(I)ri(qa)=0.3Team-maxminequilibriumwithcoordinationdevice(TMECor)Inthesettingofexantecoordination,teammembershavetheopportunitytodiscusstacticsbeforethegamebegins,butareotherwiseunabletocommunicateduringthegame,exceptviapublicly-observedactions.Apowerful,game-theoreticwaytothinkaboutexantecoordinationisthroughacoordinationdevice.Intheplanningphasebeforethegamestarts,theteammembersidentifyasetofjointpurenormal-formplans.Then,justbeforetheplay,thecoordinationdevicewillrandomlydrawoneofthenormal-formplansfromagivenprobabilitydistribution,andtheteammemberswillallactasspeci\ufb01edintheselectedplan.ANEwhereteammembersplayexantecoordinatednormal-formstrategiesiscalledateam-maxminequilibriumwithcoordinationdevice(TMECor)[6].3Inanapproximateversion,\u0001-TMECor,neithertheteamnortheopponentcangainmorethan\u0001bydeviatingfromtheirstrategy,assumingthattheotherdoesnotdeviate.Bysamplingarecommendationfromajointprobabilitydistributionover\u03a31,\u03a32,thecoordinationdeviceintroducesacorrelationbetweenthestrategiesoftheteammembersthatisotherwiseimpos-sibletocaptureusingbehavioralstrategies.Inotherwords,ingeneralthereexistsnobehavioralstrategyfortheteamplayerthatisrealization-equivalenttothenormal-formstrategyinducedbythecoordinationdevice,asthefollowingexamplefurtherillustrates.3Theyactuallycalleditcorrelation,notcoordination.Asexplainedintheintroduction,wewillusethetermcoordination.However,wewillkeeptheiracronymTMECorinsteadofswitchingtotheacronymTMECoor.3\fExample1.Considerthezero-sumgameinFigure1.Twoteammembers(Players1and2)playagainstanadversaryA.Theteamobtainsacumulativepayoffof2whenthegameendsat1or8,andapayoffof0otherwise.Avalidexantecoordinationdeviceisasfollows:theteammemberstossanunbiasedcoin;ifheadscomesup,Player1willplayactionAandPlayer2willplayactionC;otherwise,Player1willplayactionBandPlayer2willplayactionD.Therealizationinducedontheleavesissuchthat\u03c1(1)=\u03c1(8)=1/2and\u03c1(i)=0fori6\u2208{1,8}.Nobehavioralstrategyfortheteammembersisabletoinducethesamerealization.ThiscoordinationdeviceisenoughtoovercometheimperfectinformationofPlayer2aboutPlayer1\u2019smove,asPlayer2knowswhatactionwillbeplayedbyPlayer1eventhoughPlayer2willnotobserveitduringthegame.A1C2DA3C4DB\u20185C6DA7C8DBrPlayer2Player1Figure1:Exampleofextensive-formgamewithateam.Theuppercaselettersdenotetheactionnames.Thecirclednumbersuniquelyidentifytheterminalnodes.APlayer11E2FA3E4FB\u2018Player15E6FC7E8FDrPlayer2Figure2:Agamewherecoordinatedstrategieshaveaweaksignalingpower.Theuppercaselettersde-notetheactionnames.Thecirclednumbersuniquelyidentifytheterminalnodes.Onemightwonderwhetherthereisvalueinforcingthecoordinationdevicetoonlyinducenormal-formstrategiesforwhicharealization-equivalenttupleofbehavioralstrategies(oneforeachteammember)exists.Indeed,undersucharestriction,theproblemofconstructinganoptimalcoordina-tiondevicewouldamountto\ufb01ndingtheoptimaltupleofbehavioralstrategies(oneforeachteammember)thatmaximizestheteam\u2019sutility.Thissolutionconceptisknownasteam-maxminequilib-rium(TME)[25].TMEoffersconceptualsimplicitythatunfortunatelycomesatahighcost.First,\ufb01ndingthebesttupleofbehavioralstrategiesisanon-linear,non-convexoptimizationproblem.Moreover,restrictingtoTMEsisalsoundesirableintermsof\ufb01nalutilityfortheteam,sinceitmayincurinanarbitrarilylargelosscomparedtoaTMECor[6].Interestingly,aswewillproveinSection4.2,thereisastrongconnectionbetweenTMEandTMECor.Thelattersolutionconceptcanbeseenasthenatural\u201cconvexi\ufb01cation\u201doftheformer,inasensethatwewillmakepreciseinTheorem2.4Realizationform:auniversal,low-dimensionalgamerepresentationInthissection,weintroducetherealizationformofagame,whichenablesonetorepresentthestrategyspaceofaplayerbyanumberofvariablesthatislinearinthegamesize(asopposedtoexponentialasinthenormalform),eveningameswithimperfectrecall.Foreachplayeri,arealization-formstrategyisavectorthatspeci\ufb01estheprobabilitywithwhichiplaystoreachthedifferentterminalnodes.Themappingfromnormal-formstrategiestorealization-formstrategiesallowsustocompresstheactionspacefromXi,whichhasasmanycoordinatesasthenumberofnormal-formplans\u2014usuallyexponentialinthesizeofthetree\u2014toaspacethathasonecoordinateforeachterminalnode.Thismappingismany-to-onebecauseoftheredundanciesinthenormal-formrepresentation.Givenarealization-formstrategy,allthenormal-formstrategiesthatinduceitarepayoffequivalent.Theconstructionoftherealizationformreliesonthefollowingobservation.Observation1.Let\u0393beagameandz\u2208Zbeaterminalnode.Givenanormal-formstrategypro\ufb01lex=(x1,...,xn)\u2208X1\u00d7\u00b7\u00b7\u00b7\u00d7Xn,theprobabilityofreachingzcanbeuniquelydecomposedastheproductofthecontributionsofeachindividualplayer,pluschance\u2019scontribution.Formally,\u03c1x(z)=\u03c1xccQi\u2208P\u03c1xii(z).De\ufb01nition1(Realizationfunction).Let\u0393beagame.Therealizationfunctionofplayeri\u2208Pisthefunctionf\u0393i:Xi\u2192[0,1]|Z|thatmapseverynormal-formstrategyforplayeritothecorrespondingvectorofrealizationsforeachterminalnode:f\u0393i:Xi3x7\u2192(\u03c1xi(z1),...,\u03c1xi(z|Z|)).Weareinterestedintherangeoff\u0393i,calledtherealizationpolytopeofplayeri.4\fDe\ufb01nition2(Realizationpolytopeandstrategies).Playeri\u2019srealizationpolytope\u2126\u0393iingame\u0393istherangeoff\u0393i,thatisthesetofallpossiblerealizationvectorsforplayeri:\u2126\u0393i:=f\u0393i(Xi).Wecallanelement\u03c9i\u2208\u2126\u0393iarealization-formstrategy(or,simply,realization)ofplayeri.Thefunctionthatmapsatupleofrealization-formstrategies,oneforeachplayer,tothepayoffofeachplayer,ismultilinear.ThisisbyconstructionandfollowsfromObservation1.Moreover,therealizationfunctionhasthefollowingstrongproperty(allproofsareprovidedinAppendixB).Lemma1.f\u0393iisalinearfunctionand\u2126\u0393iisaconvexpolytope.Forplayerswithperfectrecall,therealizationformistheprojectionofthesequenceform,wherevariablesrelatedtonon-terminalsequencesaredropped.Inotherwords,whentheperfect-recallpropertyissatis\ufb01ed,itispossibletomovebetweenthesequence-formandtherealization-formrepresentationsbymeansofasimplelineartransformation.Therefore,therealizationpolytopeofperfect-recallgamescanbedescribedwithalinearnumber(inthegamesize)oflinearconstraints.Conversely,ingameswithimperfectrecallthenumberofconstraintsrequiredtodescribetherealiza-tionpolytopemaybeexponential4.Akeyfeatureoftherealizationformisthatitcanbeappliedtobothsettingswithoutanymodi\ufb01cation.Forexample,anoptimalNEinatwo-playerzero-sumgame,withorwithoutperfectrecalland/orinformation,canbecomputedthroughthebilinearsaddle-pointproblemmax\u03c91\u2208\u2126\u03931min\u03c92\u2208\u2126\u03931\u03c9>1U\u03c92,whereUisa(diagonal)|Z|\u00d7|Z|payoffmatrix.Finally,therealizationformofagameisformallyde\ufb01nedasfollows.De\ufb01nition3(Realizationform).Givenanextensive-formgame\u0393,itsrealizationformisatuple(P,Z,U,\u2126\u0393),where\u2126\u0393speci\ufb01esarealizationpolytopeforeachi\u2208P.4.1TwoexamplesofrealizationpolytopesToillustratetherealization-formconstruction,weconsidertwothree-playerzero-sumextensive-formgameswithperfectrecall,whereateamcomposedoftwoplayersplayingagainstthethirdplayer.Asalreadyobserved,sincetheteammemberhavethesameincentives,theteamasawholebehavesasasinglemeta-playerwith(potentially)imperfectrecall.AsweshowinExample2,ex-antecoordinationallowsteammemberstobehaveasasingleplayerwithperfectrecall.Incontrast,inExample3,thesignalingpowerofexantecoordinatedstrategiesisnotenoughtofullyrevealprivateteammembers\u2019information.Example2.ConsiderthegamedepictedinFigure1.XTisthe4-dimensionalsimplexcorre-spondingtothespaceofprobabilitydistributionsoverthesetofpurenormal-formplans\u03a3T={AC,AD,BC,BD}.Givenx\u2208XT,theprobabilitywithwhichTplaystoreachacertainout-comeisthesumofeveryx(\u03c3)suchthatplan\u03c3\u2208\u03a3Tisconsistentwiththeoutcome(i.e.,theoutcomeisreachableifTplays\u03c3).Intheexample,wehave:fT(x)=(x(A,C),x(A,D),x(B,C),x(B,D),x(A,C),x(A,D),x(B,C),x(B,D)),whereoutcomesareorderedfromlefttorightinthetree.Then,therealizationpolytopeisdescribedbyPolytope1.TheseconstraintsshowthatPlayerThasperfectrecallwhenemployingcoordinatedstrategies.Indeed,theconstraintscoincideswiththesequence-formconstraintsobtainedwhensplittingPlayer2\u2019sinformationsetintotwoinformationsets,oneforeachaction{A,B}.Example3.InthegameinFigure2,theteamPlayerThasimperfectrecallevenwhencoordinationisallowed.Inthiscase,thesignalingpowerofexantecoordinatedstrategiesisnotenoughforPlayer1topropagatetheinformationobserved(thatis,A\u2019smove)toPlayer2.Itcanbeveri\ufb01edthattherealizationpolytope\u2126\u0393TischaracterizedbythesetofconstraintsinPolytope2(seeAppendixAformoredetails).Asonemightexpect,thispolytopecontainsPolytope1.4.2Relationshipwithteammax-minequilibriumInthissubsectionwestudytherelationshipwithteammax-minequilibrium,andproveafactofpotentialindependentinterest.Thissubsectionisnotneededforunderstandingtherestofthepaper.Weprovethattherealizationpolytopeofanon-absent-mindedplayeristheconvexhullofthesetofrealizationsthatarereachablestartingfrombehavioralstrategies.ThisgivesaprecisemeaningtoourclaimthattheTMECorconceptistheconvexi\ufb01cationoftheTMEconcept.4Understandinginwhichsubclassesofimperfect-recallgamesthenumberofconstraintsremainspolyno-mialisaninterestingopenproblem.5\f\uf8f1\uf8f4\uf8f4\uf8f2\uf8f4\uf8f4\uf8f3\u03c9(5)+\u03c9(6)+\u03c9(7)+\u03c9(8)=1,\u03c9(2)=\u03c9(6),\u03c9(4)=\u03c9(8),\u03c9(1)=\u03c9(5),\u03c9(3)=\u03c9(7),\u03c9(i)\u22650i\u2208{1,2,3,4,5,6,7,8}.Polytope1:DescriptionoftherealizationpolytopeforthegameofFigure1.\uf8f1\uf8f4\uf8f4\uf8f2\uf8f4\uf8f4\uf8f3\u03c9(5)+\u03c9(6)+\u03c9(7)+\u03c9(8)=1,\u03c9(2)+\u03c9(4)=\u03c9(6)+\u03c9(8),\u03c9(1)+\u03c9(3)=\u03c9(5)+\u03c9(7),\u03c9(i)\u22650i\u2208{1,2,3,4,5,6,7,8}.Polytope2:DescriptionoftherealizationpolytopeforthegameofFigure2.De\ufb01nition4.Let\u0393beagame.Thebehavioral-realizationfunctionofplayeriisthefunction\u02dcf\u0393i:\u03a0i3\u03c07\u2192(\u03c1\u03c0i(z1),...,\u03c1\u03c0i(z|Z|))\u2208[0,1]|Z|.Accordingly,thebehavioral-realizationsetofplayeriistherangeof\u02dcf\u0393i,thatis\u02dc\u2126\u0393i:=\u02dcf\u0393i(\u03a0i).Thissetisgenerallynon-convex.Denotingbyco(\u00b7)theconvexhullofaset,wehavethefollowing:Theorem2.Consideragame\u0393.Ifplayeriisnotabsent-minded,then\u2126\u0393i=co(cid:0)\u02dc\u2126\u0393i(cid:1).5Auxiliarygame:anequivalentgamethatenablestheuseofbehavioralstrategiesIntherestofthispaper,wefocusonthree-playerzero-sumextensive-formgameswithperfectrecall,andwewillmodeltheinteractionofateamcomposedoftwoplayersplayingagainstthethirdplayer.Thetheorydevelopedalsoappliestosettingswithteamswithanarbitrarynumberofplayers.Weprovethatitispossibletoconstructanauxiliarygamewiththefollowingproperties:\u2022itisatwo-playerperfect-recallgamebetweentheadversaryAandateam-playerT;\u2022forbothplayers,thesetofbehavioralstrategiesisas\u201cexpressive\u201dasthesetofthenormal-formstrategiesintheoriginalgame(i.e.,inthecaseoftheteam,thesetofstrategiesthatteammemberscanachievethroughexantecoordination).Toaccomplishthis,weintroducearootnode\u03c6,whosebranchescorrespondtothenormal-formstrategiesofthe\ufb01rstplayeroftheteam.Thisrepresentationenablestheteamtoexpressanyprob-abilitydistributionovertheensuingsubtrees,andleadstoanequivalencebetweenthebehavioralstrategiesinthisnewperfect-recallgame(theauxiliarygame)andthenormal-formstrategiesoftheoriginaltwo-playerimperfect-recallgamebetweentheteamandtheopponent.Theauxiliarygameisaperfect-recallrepresentationoftheoriginalimperfect-recallgamesuchthattheexpressivenessofbehavioral(andsequence-form)strategiesisincreasedtomatchtheexpressivenessofnormal-formstrategiesintheoriginalgame.Considerageneric\u0393withP={1,2,A},where1and2areteammembers.WewillrefertoPlayer1asthepivotplayer.Forany\u03c31\u2208\u03a31,wede\ufb01ne\u0393\u03c31asthetwo-playergamewithP={2,A}thatweobtainfrom\u0393by\ufb01xingthechoicesofPlayer1asfollows:\u2200I\u2208I1and\u2200a\u2208A(I),ifa=\u03c31(I),then\u03c01,\u03c31(I,a)=1;otherwise,\u03c01,\u03c31(I,a)=0.Once\u03c01,\u03c31hasbeen\ufb01xedin\u0393\u03c31,decisionnodesbelongingtoPlayer1canbeconsideredasiftheywerechancenodes.Theauxiliarygameof\u0393,denotedwith\u0393\u2217,isde\ufb01nedasfollows.De\ufb01nition5(AuxiliaryGame).Theauxiliarygame\u0393\u2217isatwo-playergameobtainedfrom\u0393inthefollowingway:\u2022P={T,A};\u2022theroot\u03c6isadecisionnodeofPlayerTwithA(\u03c6)={a\u03c3}\u03c3\u2208\u03a31;\u2022eacha\u03c3isfollowedbyasubtree\u0393\u03c3;\u2022AdoesnotobservetheactionchosenbyTat\u03c6.\u03c6a\u03c3\u00b7\u00b7\u00b7\u00b7\u00b7\u00b7\u00b7\u00b7\u00b7\u0393\u03c3Figure3:Structureoftheauxiliarygame\u0393\u2217.Byconstruction,allthedecisionnodesofanyinformationsetofteamTarepartofthesamesubtree\u0393\u03c3.Intuitively,thisisbecause,intheoriginalgame,teammembersjointlypickanactionfromtheirjointprobabilitydistributionand,therefore,everyteammemberknowswhattheothermemberisgoingtoplay.Theopponenthasthesamenumberofinformationsetsbothin\u0393and\u0393\u2217.Thisisbecauseshedoesnotobservethechoiceat\u03c6and,therefore,herinformationsetsspanacrossallsubtrees\u0393\u03c3.ThebasicstructureoftheauxiliarygametreeisdepictedinFigure3(informationsetsofAareomittedforclarity).Gameswithmorethantwoteammemberscanberepresentedthough6\fa\u0393\u2217whichhasanumberofsubtreesequaltotheCartesianproductofthenormal-formplansofallteammembersexceptone.Thenextlemmaisfundamentaltounderstandtheequivalencebetweenbehavioralstrategiesof\u0393\u2217andnormal-formstrategiesof\u0393.Intuitively,itjusti\ufb01estheintroductionoftherootnode\u03c6,whosebranchescorrespondtothenormal-formstrategiesofthepivotplayer.ThisrepresentationenablestheteamTtoexpressanyconvexcombinationofrealizationsinthe\u0393\u03c3subtrees.Lemma3.Forany\u0393,\u2126\u0393T=co(cid:16)S\u03c3\u2208\u03a31\u2126\u0393\u03c3T(cid:17).ThefollowingtheoremfollowsfromLemma3andcharacterizestherelationshipbetween\u0393and\u0393\u2217.ItshowsthatthereisastrongconnectionbetweenthestrategiesofPlayerTintheauxiliarygameandtheexantecoordinatedstrategiesfortheteammembersintheoriginalgame\u0393.Theorem4.Games\u0393and\u0393\u2217arerealization-formequivalentinthefollowingsense:(i)Team.Givenanydistributionovertheactionsatthegametreeroot\u03c6(i.e.,achoice\u03a313\u03c37\u2192\u03bb\u03c3\u22650suchthatP\u03c3\u03bb\u03c3=1)andanychoiceofrealizations{\u03c9\u03c3\u2208\u2126\u0393\u03c3T}\u03c3\u2208\u03a31,wehavethatP\u03c3\u2208\u03a31\u03bb\u03c3\u03c9\u03c3\u2208\u2126\u0393T.Theconverseisalsotrue:givenany\u03c9\u2208\u2126\u0393T,thereexistsachoiceof{\u03bb\u03c3}\u03c3\u2208\u03a31andrealizations{\u03c9\u03c3\u2208\u2126\u0393\u03c3T}\u03c3\u2208\u03a31suchthat\u03c9=P\u03c3\u2208\u03a31\u03bb\u03c3\u03c9\u03c3.(ii)Adversary.Therealizationpolytopeoftheadversarysatis\ufb01es\u2126\u0393A=\u2126\u0393\u2217A.ThefollowingisthenadirectconsequenceofTheorem4Corollary1.Thesetofpayoffsreachablein\u0393coincideswiththesetofpayoffsreachablein\u0393\u2217.Speci\ufb01cally,anystrategy{\u03bb\u03c3}\u03c3\u2208\u03a31,{\u03c9\u03c3}\u03c3\u2208\u03a31over\u0393\u2217ispayoff-equivalenttotherealization-formstrategy\u03c9=P\u03c3\u2208\u03a31\u03bb\u03c3\u03c9\u03c3in\u0393.Remark1.Since\u0393\u03c3hasperfectrecall,everyrealization\u03c9\u03c3\u2208\u2126\u0393\u03c3canbeinducedbyTviabehavioralstrategies.Theaboveshowsthatforeveryexantecoordinatedstrategyfortheteamin\u0393,thereexistsacor-responding(payoff-equivalent)behavioralstrategyforTin\u0393\u2217,andviceversa.Hence,duetorealization-formequivalencebetween\u0393and\u0393\u2217,\ufb01ndingaTMECorin\u0393(employingexanteco-ordinatednormal-formstrategies),isequivalentto\ufb01ndingaNEin\u0393\u2217(withbehavioralstrategies).6Fictitiousteam-play:ananytimealgorithmforTMECorThissectionintroducesananytimealgorithm,\ufb01ctitiousteam-play,for\ufb01ndingaTMECor.Itfollowsfromtheprevioussectionthatinorderto\ufb01ndaTMECorin\u0393,itsuf\ufb01cesto\ufb01ndatwo-playerNEintheauxiliarygame\u0393\u2217(andviceversa,althoughwedonotusethisseconddirection).Further-more,since\u0393\u2217isatwo-playerperfect-recallzero-sumgame,the\ufb01ctitiousplay(FP)algorithmcanbeappliedwithitstheoreticalguaranteeofconvergingtoaNE.Fictitiousplay[2,19]isanitera-tivealgorithmoriginallydescribedfornormal-formgames.Itkeepstrackofaveragenormal-formstrategies\u00afxi,whichareoutputintheend,andtheyconvergetoaNE.Atiterationt,playericom-putesthebestresponseagainsttheopponent\u2019sempiricaldistributionofplayuptotimet\u22121,thatis,xti=BR(\u00afxt\u22121\u2212i).Thenheraveragestrategyisupdatedas\u00afxti=t\u22121t\u00afxt\u22121i+1txti.Conceptually,our\ufb01ctitiousteam-playalgorithmcoincideswithFPappliedtotheauxiliarygame\u0393\u2217.However,inordertoavoidtheexponentialsizeof\u0393\u2217,our\ufb01ctitiousteam-playalgorithmdoesnotexplicitlyworkontheauxiliarygame.Rather,itencodesthebest-responseproblemsbymeansofmixedintegerlinearprograms(MILPs)ontheoriginalgame\u0393.Themainalgorithm.ThepseudocodeofthemainalgorithmisgivenasAlgorithm1,whereBRA(\u00b7)andBRT(\u00b7)arethesubroutinesforsolvingthebest-responseproblems.Ouralgorithmemploysrealization-formstrategies.Thisallowsforasigni\ufb01cantlymoreintuitivewayofperformingaveraging(Steps7,8,10)thanwhatisdoneinfull-widthextensive-form\ufb01ctitiousplay[11],whichemploysbehavioralstrategies.Ouralgorithmmaintainsanaveragerealization\u00af\u03c9Afortheadversary.Moreover,the|\u03a31|-dimensionalvector\u00af\u03bbkeepstheempiricalfrequenciesofactionsatnode\u03c6inauxiliarygame\u0393\u2217(seeFigure3).Finally,\u2200\u03c3\u2208\u03a31,\u00af\u03c9T,\u03c3\u2208\u2126\u0393\u03c3Tistheaveragerealizationoftheteaminthesubtree\u0393\u03c3.7\fAlgorithm1Fictitiousteam-play1:functionFICTITIOUSTEAMPLAY(\u0393)2:Initialize\u00af\u03c9A3:\u00af\u03bb\u2190(0,...,0),t\u219014:\u00af\u03c9T,\u03c3\u2190(0,...,0)\u2200\u03c3\u2208\u03a315:whilewithincomputationalbudgetdo6:(\u03c3t,\u03c9tT)\u2190BRT(\u00af\u03c9A)7:\u00af\u03bb\u2190(1\u22121t)\u00af\u03bb+1t1\u03c3t8:\u00af\u03c9T,\u03c3t\u2190(1\u22121t)\u00af\u03c9T,\u03c3t+1t\u03c9tT9:\u03c9tA\u2190BRA(\u00af\u03bb,{\u00af\u03c9T,\u03c3}\u03c3)10:\u00af\u03c9A\u2190(1\u22121t)\u00af\u03c9A+1t\u03c9tA11:t\u2190t+112:return(\u00af\u03bb,(\u00af\u03c9T,\u03c3)\u03c3\u2208\u03a31)Aftertiterationsofthealgorithm,onlytpairsofstrategiesaregenerated.Hence,anoptimizedimplementationofthealgorithmcanemployalazydatastructuretokeeptrackofthechangesto\u00af\u03bband\u00af\u03c9T,\u03c3.Initially(Step2),theaveragerealization\u00af\u03c9Aoftheadver-saryissettotherealization-formstrategyequivalenttoauniformbehavioralstrategypro\ufb01le.Ateachiterationthealgorithm\ufb01rstcomputesateam\u2019sbest-responseagainst\u00af\u03c9A.Werequirethatthechosenbestresponseassignprob-ability1tooneoftheavailableactions(say,a\u03c3t)atnode\u03c6.(Apure\u2014thatis,non-randomized\u2014bestresponseal-waysexistsand,therefore,inparticulartherealwaysexistsatleastonebestresponseselectingasingleactionattherootwithprobabilityone.)Then,theaver-agefrequenciesandteam\u2019srealizationsareupdatedonthebasisoftheobserved(\u03c3t,\u03c9tT).Finally,theadversary\u2019sbestresponse\u03c9tAagainsttheupdatedaveragestrategyoftheteamiscomputed,andtheempiricaldistributionofplayoftheadversaryisupdated.Theexantecoordinatedstrategypro\ufb01lefortheteamisimplicitlyrepresentedbythepair(\u00af\u03bb,\u00af\u03c9T,\u03c3).Inparticular,thatpairencodesacoordinationdevicethatoperatesasfollows:\u2022Atthebeginningofthegame,apurenormal-formplan\u02dc\u03c3\u2208\u03a3issampledaccordingtothediscreteprobabilitydistributionencodedby\u00af\u03bb.Player1willplaythegameaccordingtothesampledplan.\u2022Player2willplayaccordingtoanynormal-formstrategyinf\u221212(\u00af\u03c9T,\u02dc\u03c3),thatis,anynormal-formstrategywhoserealizationis\u00af\u03c9T,\u02dc\u03c3.Thecorrectnessofthealgorithmthenisadirectconsequenceofrealization-equivalencebetween\u0393and\u0393\u2217,whichwasshowninTheorem4.Inparticular,thestrategyoftheteamconvergestoapro\ufb01lethatispartofanormal-formNEintheoriginalgame\u0393.Best-responsesubroutines.Theproblemof\ufb01ndingtheadversary\u2019sbestresponsetoapairofstrate-giesoftheteam,namelyBRA(\u00af\u03bb,{\u00af\u03c9T,\u03c3}\u03c3),canbeef\ufb01cientlytackledbyworkingon\u0393(secondpointofTheorem4).Incontrast,theproblemofcomputingBRT(\u00af\u03c9A)isNP-hard[24],andinapprox-imable[6].CelliandGatti[6]proposeaMILPformulationtosolvetheteambest-responseproblem.InAppendixEweproposeanalternativeMILPformulationinwhichthenumberofbinaryvariablesispolynomialin\u0393andproportionaltothenumberofsequencesofthepivotplayer.Inouralgorithm,weemployameta-oraclethatusessimultaneously,asparallelprocesses,bothsubroutines,andstopsthemassoonasoneofthetwohasfoundasolutionor,inthecaseatime-limitisreached,itstopsbothsubroutinesanditreturnsthebestsolution(intermsofteam\u2019sutility).ThiscircumventstheneedtoproveoptimalityintheMILP,whichoftentakesmostoftheMILP-solvingtime,andopensthedoorstoheuristicMILP-solvingtechniques.Oneofthekeyfeaturesofthemeta-oracleisthatitsperformancesarenotimpactedbythesizeof\u0393\u2217,whichisneverexplicitlyemployedinthebest-responsescomputation.7ExperimentsWeconductedexperimentsonthree-playerKuhnpokergamesandthree-playerLeduchold\u2019empokergames.Thesearestandardgamesinthecomputationalgametheoryliterature,anddescriptionofthemcanbefoundinAppendixF.Ourinstancesareparametricinthenumberofranksinthedeck.TheinstancesadoptedarelistedinTables1and2,whereKrandLrdenote,respectively,aKuhninstancewithrranksandaLeducinstancewithrranks(i.e.,3rtotalcards).Table1alsodisplaystheinstances\u2019dimensionsintermsofthenumberofinformationsetsperplayerandthenumberofsequences(i.e.,numberofinformationset\u2013actionpairs)perplayer,aswellasthepayoffdispersion\u2206u\u2014thatis,thedifferencebetweenthemaximumandminimumattainableteamutility.Fictitiousteam-play.Weinstantiated\ufb01ctitiousteam-playwiththemeta-oraclepreviouslydis-cussed,whichreturnsthebestsolutionfoundbytheMILPoracleswithinthetimelimit.Weleteachbest-responseformulationrunontheGurobi8.0MILPsolver,withatimelimitof15secondsand5000maximumiterations.Ouralgorithmisananytimealgorithm,soitdoesnotrequireatargetaccuracy\u0001for\u0001-TMECortobespeci\ufb01edinadvance.Table1showstheany-8\fGameTreesize\u2206uFictitiousteam-playHCGInf.Seq.10%5%2%1.5%1%0.5%K3251360s0s0s1s1s1s0sK4331761s1s4s4s30s1m12s9sK5412161s2s44s1m4m15s8m57s1m58sK6492561s12s43s5m15s8m30s23m32s25m26sK7572964s17s2m15s5m46s6m31s23m49s2h50mL34572292115s1m14m05s30m40s1h34m30s>24hoomL4801401211s1m31s11m8s51m5s6h51m>24hoomTable1:Comparisonbetweentheruntimesof\ufb01ctitiousteam-play(forvariouslevelsofaccuracy)andthehybridcolumngeneration(HCG)algorithm.(oom:outofmemory.)GameTeamUtilityAdv1Adv2Adv3K30.00000.00000.0003K40.04050.0259-0.0446K50.04340.0156-0.0282K60.05140.0271-0.0253K70.05920.0285-0.0259L30.23320.20890.1475L40.19910.1419-0.0223Table2:Valuesoftheaver-agestrategypro\ufb01lefordif-ferentchoicesofadversary.timeperformance,thatis,thetimeittooktoreachan\u03b1\u2206u-TMECorfordifferentaccuracies\u03b1\u2208{10%,5%,2%,1.5%,1%,0.5%}.ResultsinTable1assumethattheteamconsistsofthe\ufb01rstandthirdmoverinthegame;theopponentisthesecondmover.Table2showsthevalueoftheaveragestrategycomputedby\ufb01ctitiousteam-playfordifferentchoicesoftheopponentplayer.Thisvaluecorrespondstotheexpectedutilityoftheteamfortheaveragestrategypro\ufb01le(\u00af\u03bb,\u00af\u03c9T,\u03c3)atit-eration1000.InAppendixF.3weshowtheminimumcumulativeutilitythattheteamisguaranteedtoachieve,thatis\u2212BRA(\u00af\u03bb,{\u00af\u03c9T,\u03c3}\u03c3).Hybridcolumngenerationbenchmark.Wecomparedagainstthehybridcolumngeneration(HCG)algorithm[6],whichistheonlyprioralgorithmforthisproblem.Tomakethecompari-sonfair,weinstantiateHCGwiththesamemeta-oraclediscussedintheprevioussection.WeagainuseGurobi8.0MILPsolvertosolvethebestresponseproblemfortheteam.However,inthecaseofHCG,notimelimitcanbesetonGurobiwithoutinvalidatingthetheoreticalconvergenceguaranteeofthealgorithm.Thisisadrawback,asitpreventsHCGfromrunninginananytimefashion,de-spitecolumngenerationotherwisebeingananytimealgorithm.IntheLeducpokerinstances,HCGexceededthememorybudget(40GB).Ourexperimentsshowthat\ufb01ctitiousteam-playscalestosigni\ufb01cantlylargergamesthanHCG.Inter-estingly,inalmostallthegames,thevalueoftheteamwasnon-negative:bycolluding,theteamwasabletoachievevictory.Moreover,inAppendixF.4,weshowthataTMECorprovidestotheteamasubstantialpayoffincreaseoverthesettingwhereteammembersplayinbehavioralstrategies.8ConclusionsandfutureresearchThestudyofalgorithmsformulti-playergamesischallenging.Inthispaper,weproposedanalgo-rithmforsettingsinwhichateamofplayersfacesanadversaryandtheteammemberscanexploitonlyexantecoordination,discussingandagreeingontacticsbeforethegamestarts.Our\ufb01rstcon-tributionwastherealizationform,anovelrepresentationthatallowsustorepresentthestrategiesofthenormalformmoreconcisely.Therealizationformisalsoapplicabletoimperfect-recallgames.Weusedittoderiveatwo-playerperfect-recallauxiliarygamethatisequivalenttotheoriginalgame,andprovidesatheoreticallysoundwaytomaptheproblemof\ufb01ndinganoptimalex-ante-coordinatedstrategyfortheteamtoaclassicalwell-understoodNashequilibrium-\ufb01ndingprobleminatwo-playerzero-sumperfect-recallgame.Oursecondcontributionwasthedesignofthe\ufb01c-titiousteam-playalgorithm,whichemploysanovelbest-responsemeta-oracle.Theanytimealgo-rithmisguaranteedtoconvergetoanequilibrium.Ourexperimentsshowedthat\ufb01ctitiousteam-playisdramaticallyfasterthantheprioralgorithmsforthisproblem.Inthefuture,itwouldbeinterestingtoadaptotherpopularequilibriumcomputationtechniquesfromthetwo-playersetting(suchasCFR)foroursetting,byreasoningovertheauxiliarygame.Thestudyofalgorithmsforteamgamescouldshedfurtherlightonhowtodealwithimperfect-recallgames,thatarereceivingincreasingattentioninthecommunityduetotheapplicationofimperfect-recallabstractionstothecomputationofstrategiesforlargeextensive-formgames[26,14,10,3,8,13].Acknowledgments.ThismaterialisbasedonworksupportedbytheNationalScienceFoundationundergrantsIIS-1718457,IIS-1617590,andCCF-1733556,andtheAROunderawardW911NF-17-1-0082.9\fReferences[1]N.Basilico,A.Celli,G.DeNittis,andN.Gatti.Team-maxminequilibrium:ef\ufb01ciencyboundsandalgorithms.InAAAIConferenceonArti\ufb01cialIntelligence(AAAI),pages356\u2013362,2017.[2]G.W.Brown.Iterativesolutionofgamesby\ufb01ctitiousplay.Activityanalysisofproductionandallocation,13(1):374\u2013376,1951.[3]N.Brown,S.Ganzfried,andT.Sandholm.Hierarchicalabstraction,distributedequilibriumcomputation,andpost-processing,withapplicationtoachampionno-limitTexasHold\u2019emagent.InInternationalConferenceonAutonomousAgentsandMulti-AgentSystems(AAMAS),2015.[4]N.BrownandT.Sandholm.Safeandnestedsubgamesolvingforimperfect-informationgames.InAdvancesinNeuralInformationProcessingSystems(NIPS),pages689\u2013699,2017.[5]N.BrownandT.Sandholm.SuperhumanAIforheads-upno-limitpoker:Libratusbeatstopprofessionals.Science,pageeaao1733,2017.[6]A.CelliandN.Gatti.Computationalresultsforextensive-formadversarialteamgames.InAAAIConferenceonArti\ufb01cialIntelligence(AAAI),2018.[7]J.\u02c7Cerm\u00b4ak,B.Bo\u02c7sansk`y,K.Hor\u00b4ak,V.Lis`y,andM.P\u02c7echou\u02c7cek.Approximatingmaxminstrategiesinimperfectrecallgamesusinga-lossrecallproperty.InternationalJournalofAp-proximateReasoning,93:290\u2013326,2018.[8]J.\u02c7Cerm\u00b4ak,B.Bo\u02c7sansky,andV.Lis\u00b4y.Analgorithmforconstructingandsolvingimperfectrecallabstractionsoflargeextensive-formgames.InProceedingsoftheInternationalJointConferenceonArti\ufb01cialIntelligence(IJCAI),pages936\u2013942,2017.[9]J.\u02c7Cerm\u00b4ak,B.Bo\u02c7sansk`y,andM.P\u02c7echou\u02c7cek.Combiningincrementalstrategygenerationandbranchandboundsearchforcomputingmaxminstrategiesinimperfectrecallgames.InInter-nationalConferenceonAutonomousAgentsandMulti-AgentSystems(AAMAS),pages902\u2013910,2017.[10]S.GanzfriedandT.Sandholm.Potential-awareimperfect-recallabstractionwithearthmover\u2019sdistanceinimperfect-informationgames.InAAAIConferenceonArti\ufb01cialIntelligence(AAAI),2014.[11]J.Heinrich,M.Lanctot,andD.Silver.Fictitiousself-playinextensive-formgames.InInter-nationalConferenceonMachineLearning(ICML),pages805\u2013813,2015.[12]D.Koller,N.Megiddo,andB.VonStengel.Ef\ufb01cientcomputationofequilibriaforextensivetwo-persongames.Gamesandeconomicbehavior,14(2):247\u2013259,1996.[13]C.KroerandT.Sandholm.Imperfect-recallabstractionswithboundsingames.InProceedingsoftheACMConferenceonEconomicsandComputation(EC),2016.[14]M.Lanctot,R.Gibson,N.Burch,M.Zinkevich,andM.Bowling.No-regretlearninginextensive-formgameswithimperfectrecall.InInternationalConferenceonMachineLearning(ICML),2012.[15]M.Maschler,S.Zamir,E.Solan,andM.Borns.GameTheory.CambridgeUniversityPress,2013.[16]H.B.McMahan,G.J.Gordon,andA.Blum.Planninginthepresenceofcostfunctionscontrolledbyanadversary.InInternationalConferenceonMachineLearning(ICML),pages536\u2013543,2003.[17]J.Nash.Equilibriumpointsinn-persongames.ProceedingsoftheNationalAcademyofSciences,36:48\u201349,1950.[18]M.PiccioneandA.Rubinstein.Ontheinterpretationofdecisionproblemswithimperfectrecall.GamesandEconomicBehavior,20(1):3\u201324,1997.10\f[19]J.Robinson.Aniterativemethodofsolvingagame.Annalsofmathematics,pages296\u2013301,1951.[20]Y.ShohamandK.Leyton-Brown.Multiagentsystems:Algorithmic,game-theoretic,andlog-icalfoundations.CambridgeUniversityPress,2008.[21]F.Southey,M.Bowling,B.Larson,C.Piccione,N.Burch,D.Billings,andC.Rayner.Bayes\u2019bluff:Opponentmodellinginpoker.InProceedingsofthe21stAnnualConferenceonUncer-taintyinArti\ufb01cialIntelligence(UAI),July2005.[22]M.TawarmalaniandN.V.Sahinidis.Apolyhedralbranch-and-cutapproachtoglobalopti-mization.MathematicalProgramming,103:225\u2013249,2005.[23]B.VonStengel.Ef\ufb01cientcomputationofbehaviorstrategies.GamesandEconomicBehavior,14(2):220\u2013246,1996.[24]B.vonStengelandF.Forges.Extensive-formcorrelatedequilibrium:De\ufb01nitionandcompu-tationalcomplexity.MathematicsofOperationsResearch,33(4):1002\u20131022,2008.[25]B.vonStengelandD.Koller.Team-maxminequilibria.GamesandEconomicBehavior,21(1-2):309\u2013321,1997.[26]K.Waugh,M.Zinkevich,M.Johanson,M.Kan,D.Schnizlein,andM.Bowling.Apracti-caluseofimperfectrecall.InSymposiumonAbstraction,ReformulationandApproximation(SARA),2009.[27]P.C.Wichardt.ExistenceofNashequilibriain\ufb01niteextensiveformgameswithimperfectrecall:Acounterexample.GamesandEconomicBehavior,63(1):366\u2013369,2008.11\f", "award": [], "sourceid": 5940, "authors": [{"given_name": "Gabriele", "family_name": "Farina", "institution": "Carnegie Mellon University"}, {"given_name": "Andrea", "family_name": "Celli", "institution": "Politecnico di Milano"}, {"given_name": "Nicola", "family_name": "Gatti", "institution": "Politecnico di Milano"}, {"given_name": "Tuomas", "family_name": "Sandholm", "institution": "Carnegie Mellon University"}]}