语音识别外文翻译外文文献英文文献.docx

文档编号：26388510
上传时间：2023-06-18
格式：DOCX
页数：13
大小：27.47KB

《语音识别外文翻译外文文献英文文献.docx》由会员分享，可在线阅读，更多相关《语音识别外文翻译外文文献英文文献.docx（13页珍藏版）》请在冰豆网上搜索。

语音识别外文翻译外文文献英文文献.docx

语音识别外文翻译外文文献英文文献

SpeechRecognition

VictorZue,RonCole,&WayneWard

MITLaboratoryforComputerScience,Cambridge,Massachusetts,USAOregonGraduateInstituteofScience&Technology,Portland,Oregon,USA

CarnegieMellonUniversity,Pittsburgh,Pennsylvania,USA

1DefiningtheProblem

Speechrecognitionistheprocessofconvertinganacousticsignal,capturedbyamicrophoneoratelephone,toasetofwords.Therecognizedwordscanbethefinalresults,asforapplicationssuchascommands&control,dataentry,anddocumentpreparation.Theycanalsoserveastheinputtofurtherlinguisticprocessinginordertoachievespeechunderstanding,asubjectcoveredinsection.

Speechrecognitionsystemscanbecharacterizedbymanyparameters,someofthemoreimportantofwhichareshowninFigure.Anisolated-wordspeechrecognitionsystemrequiresthatthespeakerpausebrieflybetweenwords,whereasacontinuousspeechrecognitionsystemdoesnot.Spontaneous,orextemporaneouslygenerated,speechcontainsdisfluencies,andismuchmoredifficulttorecognizethanspeechreadfromscript.Somesystemsrequirespeakerenrollment---ausermustprovidesamplesofhisorherspeechbeforeusingthem,whereasothersystemsaresaidtobespeaker-independent,inthatnoenrollmentisnecessary.Someoftheotherparametersdependonthespecifictask.Recognitionisgenerallymoredifficultwhenvocabulariesarelargeorhavemanysimilar-soundingwords.Whenspeechisproducedinasequenceofwords,languagemodelsorartificialgrammarsareusedtorestrictthecombinationofwords.

Thesimplestlanguagemodelcanbespecifiedasafinite-statenetwork,wherethepermissiblewordsfollowingeachwordaregivenexplicitly.Moregenerallanguagemodelsapproximatingnaturallanguagearespecifiedintermsofacontext-sensitivegrammar.

Onepopularmeasureofthedifficultyofthetask,combiningthevocabularysizeandthe1languagemodel,isperplexity,looselydefinedasthegeometricmeanofthenumberofwordsthatcanfollowawordafterthelanguagemodelhasbeenapplied（seesectionforadiscussionoflanguagemodelingingeneralandperplexityinparticular）.Finally,therearesomeexternalparametersthatcanaffectspeechrecognitionsystemperformance,includingthecharacteristicsoftheenvironmentalnoiseandthetypeandtheplacementofthemicrophone.

Speechrecognitionisadifficultproblem,largelybecauseofthemanysourcesofvariabilityassociatedwiththesignal.First,theacousticrealizationsofphonemes,thesmallestsoundunitsofwhichwordsarecomposed,arehighlydependentonthecontextinwhichtheyappear.Thesephoneticvariabilitiesareexemplifiedbytheacousticdifferencesofthephoneme，Atwordboundaries,contextualvariationscanbequitedramatic---makinggasshortagesoundlikegashshortageinAmericanEnglish,anddevoandaresoundlikedevandareinItalian.

Second,acousticvariabilitiescanresultfromchangesintheenvironmentaswellasinthepositionandcharacteristicsofthetransducer.Third,within-speakervariabilitiescanresultfromchangesinthespeaker'sphysicalandemotionalstate,speakingrate,orvoicequality.Finally,differencesinsociolinguisticbackground,dialect,andvocaltractsizeandshapecancontributetoacross-speakervariabilities.

Figureshowsthemajorcomponentsofatypicalspeechrecognitionsystem.Thedigitizedspeechsignalisfirsttransformedintoasetofusefulmeasurementsorfeaturesatafixedrate,2typicallyonceevery10--20msec（seesectionsand11.3forsignalrepresentationanddigitalsignalprocessing,respectively）.Thesemeasurementsarethenusedtosearchforthemostlikelywordcandidate,makinguseofconstraintsimposedbytheacoustic,lexical,andlanguagemodels.Throughoutthisprocess,trainingdataareusedtodeterminethevaluesofthemodelparameters.

Speechrecognitionsystemsattempttomodelthesourcesofvariabilitydescribedaboveinseveralways.Atthelevelofsignalrepresentation,researchershavedevelopedrepresentationsthatemphasizeperceptuallyimportantspeaker-independentfeaturesofthesignal,andde-emphasizespeaker-dependentcharacteristics.Attheacousticphoneticlevel,speakervariabilityistypicallymodeledusingstatisticaltechniquesappliedtolargeamountsofdata.Speakeradaptationalgorithmshavealsobeendevelopedthatadaptspeaker-independentacousticmodelstothoseofthecurrentspeakerduringsystemuse,（seesection）.Effectsoflinguisticcontextattheacousticphoneticlevelaretypicallyhandledbytrainingseparatemodelsforphonemesindifferentcontexts;thisiscalledcontextdependentacousticmodeling.

Wordlevelvariabilitycanbehandledbyallowingalternatepronunciationsofwordsinrepresentationsknownaspronunciationnetworks.Commonalternatepronunciationsofwords,aswellaseffectsofdialectandaccentarehandledbyallowingsearchalgorithmstofindalternatepathsofphonemesthroughthesenetworks.Statisticallanguagemodels,basedonestimatesofthefrequencyofoccurrenceofwordsequences,areoftenusedtoguidethesearchthroughthemostprobablesequenceofwords.

ThedominantrecognitionparadigminthepastfifteenyearsisknownashiddenMarkovmodels（HMM）.AnHMMisadoublystochasticmodel,inwhichthegenerationoftheunderlyingphonemestringandtheframe-by-frame,surfaceacousticrealizationsarebothrepresentedprobabilisticallyasMarkovprocesses,asdiscussedinsections,and11.2.Neuralnetworkshavealsobeenusedtoestimatetheframebasedscores;thesescoresarethenintegratedintoHMM-basedsystemarchitectures,inwhathascometobeknownashybridsystems,asdescribedinsection11.5.

Aninterestingfeatureofframe-basedHMMsystemsisthatspeechsegmentsareidentifiedduringthesearchprocess,ratherthanexplicitly.Analternateapproachistofirstidentifyspeechsegments,thenclassifythesegmentsandusethesegmentscorestorecognizewords.Thisapproachhasproducedcompetitiverecognitionperformanceinseveraltasks.

2StateoftheArt

Commentsaboutthestate-of-the-artneedtobemadeinthecontextofspecificapplicationswhichreflecttheconstraintsonthetask.Moreover,differenttechnologiesaresometimesappropriatefordifferenttasks.Forexample,whenthevocabularyissmall,theentirewordcanbemodeledasasingleunit.Suchanapproachisnotpracticalforlargevocabularies,wherewordmodelsmustbebuiltupfromsubwordunits.

Thepastdecadehaswitnessedsignificantprogressinspeechrecognitiontechnology.Worderrorratescontinuetodropbyafactorof2everytwoyears.Substantialprogresshasbeenmadeinthebasictechnology,leadingtotheloweringofbarrierstospeakerindependence,continuousspeech,andlargevocabularies.Thereareseveralfactorsthathavecontributedtothisrapidprogress.First,thereisthecomingofageoftheHMM.HMMispowerfulinthat,withtheavailabilityoftrainingdata,theparametersofthemodelcanbetrainedautomaticallytogiveoptimalperformance.

Second,muchefforthasgoneintothedevelopmentoflargespeechcorporaforsystemdevelopment,training,andtesting.Someofthesecorporaaredesignedforacousticphoneticresearch,whileothersarehighlytaskspecific.Nowadays,itisnotuncommontohavetensofthousandsofsentencesavailableforsystemtrainingandtesting.Thesecorporapermitresearcherstoquantifytheacousticcuesimportantforphoneticcontrastsandtodetermineparametersoftherecognizersinastatisticallymeaningfulway.Whilemanyofthesecorpora（e.g.,TIMIT,RM,ATIS,andWSJ;seesection12.3）wereoriginallycollectedunderthesponsorshipoftheU.S.DefenseAdvancedResearchProjectsAgency（ARPA）tospurhumanlanguagetechnologydevelopmentamongitscontractors,theyhaveneverthelessgainedworld-wideacceptance（e.g.,inCanada,France,Germany,Japan,andtheU.K.）asstandardsonwhichtoevaluatespeechrecognition.

Third,progresshasbeenbroughtaboutbytheestablishmentofstandardsforperformanceevaluation.Onlyadecadeago,researcherstrainedandtestedtheirsystemsusinglocallycollecteddata,andhadnotbeenverycarefulindelineatingtrainingandtestingsets.Asaresult,itwasverydifficulttocompareperformanceacrosssystems,andasystem'sperformancetypicallydegradedwhenitwaspresentedwithpreviouslyunseendata.Therecentavailabilityofalargebodyofdatainthepublicdomain,coupledwiththespecificationofevaluationstandards,hasresultedinuniformdocumentationoftestresults,thuscontributingtogreaterreliabilityinmonitoringprogress（corpusdevelopmentactivitiesandevaluationmethodologiesaresummarizedinchapters12and13respectively）.

Finally,advancesincomputertechnologyhavealsoindirectlyinfluencedourprogress.Theavailabilityoffastcomputerswithinexpensivemassstoragecapabilitieshasenabledresearcherstorunmanylargescaleexperimentsinashortamountoftime.Thismeansthattheelapsedtimebetweenanideaanditsimplementationandevaluationisgreatlyreduced.Infact,speechrecognitionsystemswithreasonableperformancecannowruninrealtimeusinghigh-endworkstationswithoutadditionalhardware---afeatunimaginableonlyafewyearsago.

Oneofthemostpopular,andpotentiallymostusefultaskswithlowperplexity（PP=11）istherecognitionofdigits.ForAmericanEnglish,speaker-independentrecognitionofdigitstringsspokencontinuouslyandrestrictedtotelephonebandwidthcanachieveanerrorrateof0.3%whenthestringlengthisknown.

Oneofthebestknownmoderate-perplexitytasksisthe1,000-wordso-calledResource5Management（RM）task,inwhichinquiriescanbema