E2E Spoken Entity Extraction for Virtual Agents

التفاصيل البيبلوغرافية
العنوان: E2E Spoken Entity Extraction for Virtual Agents
المؤلفون: Singla, Karan, Kim, Yeon-Jun
بيانات النشر: arXiv, 2023.
سنة النشر: 2023
مصطلحات موضوعية: FOS: Computer and information sciences, Sound (cs.SD), Computer Science - Computation and Language, Audio and Speech Processing (eess.AS), FOS: Electrical engineering, electronic engineering, information engineering, Computation and Language (cs.CL), Computer Science - Sound, Electrical Engineering and Systems Science - Audio and Speech Processing
الوصف: This paper rethink some aspects of speech processing using speech encoders, specifically about extracting entities directly from speech, without intermediate textual representation. In human-computer conversations, extracting entities such as names, street addresses and email addresses from speech is a challenging task. In this paper, we study the impact of fine-tuning pre-trained speech encoders on extracting spoken entities in human-readable form directly from speech without the need for text transcription. We illustrate that such a direct approach optimizes the encoder to transcribe only the entity relevant portions of speech ignoring the superfluous portions such as carrier phrases, or spell name entities. In the context of dialog from an enterprise virtual agent, we demonstrate that the 1-step approach outperforms the typical 2-step approach which first generates lexical transcriptions followed by text-based entity extraction for identifying spoken entities.
DOI: 10.48550/arxiv.2302.10186
الوصول الحر: https://explore.openaire.eu/search/publication?articleId=doi_dedup___::649c401081fd5855da2e063f115feea5Test
حقوق: OPEN
رقم الانضمام: edsair.doi.dedup.....649c401081fd5855da2e063f115feea5
قاعدة البيانات: OpenAIRE