<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title>Alexa Prize</title>
    <link>https://www.amazon.science/alexa-prize</link>
    <description>Alexa Prize</description>
    <language>en-US</language>
    <lastBuildDate>Tue, 03 Oct 2023 13:00:34 GMT</lastBuildDate>
    <atom:link href="https://www.amazon.science/alexa-prize.rss" type="application/rss+xml" rel="self" />
    <item>
      <title>Alexa Prize TaskBot Challenge 2 winner announced</title>
      <link>https://www.amazon.science/alexa-prize/taskbot-challenge/2022</link>
      <description>Team TWIZ from NOVA School of Science and Technology awarded $500,000 prize for first-place overall performance.</description>
      <pubDate>Tue, 03 Oct 2023 13:00:34 GMT</pubDate>
      <guid>https://www.amazon.science/alexa-prize/taskbot-challenge/2022</guid>
    </item>
    <item>
      <title>Advancing conversational task assistance: the second Alexa Prize TaskBot challenge</title>
      <link>https://www.amazon.science/alexa-prize/proceedings/alexa-lets-work-together-introducing-the-second-alexa-prize-taskbot-challenge</link>
      <description>Since its inception in 2016, the Alexa Prize program has enabled hundreds of university students and faculty to explore and compete in the development of conversational agents through the SocialBot Grand Challenge, whose goal is to build agents capable of conversing coherently and engagingly with humans on popular topics. As conversational agents attempt to assist users with increasingly complex tasks, new conversational AI techniques and evaluation platforms are needed. The Alexa Prize TaskBot Challenge, now in its second year, introduced the requirements of interactively assisting humans with real-world tasks, while making use of both voice and visual modalities. This challenge requires the TaskBots to identify and understand the user&amp;#8217;s need, identify and integrate task and domain knowledge into the interaction, and develop new ways of engaging the user without distracting them from the task at hand, among other challenges. This paper provides an overview of the second TaskBot challenge, in which both new and returning teams participated. We describe the infrastructure support and the new models provided to the teams with the CoBot Toolkit. We then summarize the approaches the participating teams took to address research challenges, including changes and improvements from the previous year. Finally, we analyze the performance of the competing TaskBots and discuss some of the lessons learned.</description>
      <pubDate>Tue, 03 Oct 2023 12:16:29 GMT</pubDate>
      <guid>https://www.amazon.science/alexa-prize/proceedings/alexa-lets-work-together-introducing-the-second-alexa-prize-taskbot-challenge</guid>
    </item>
    <item>
      <title>TACO 2.0: A task-oriented dialogue system with mixed initiatives and multi-modal interaction</title>
      <link>https://www.amazon.science/alexa-prize/proceedings/taco-2-0-a-task-oriented-dialogue-system-with-mixed-initiatives-and-multi-modal-interaction</link>
      <description>In the inaugural Alexa Prize TaskBot Challenge, we introduced TACO 1.0, a task-oriented digital assistant, designed to guide users through multi-step tasks in cooking and home improvement. Equipped with a suite of components including language understanding, dialogue management, and response generation, bolstered by a search engine, TACO 1.0 set a robust foundation in user-centered assistance. Building on this, we present TACO 2.0, aspiring to deliver a more collaborative and engaging dialogue experience. Towards this end, we refine our mechanisms to better accommodate the dynamic nature of real-world conversations, supporting mixed initiatives from both users and agents. In terms of user initiative, we develop an upgraded hierarchical intent recognition module and a more powerful question- answering system to accurately comprehend and respond to user needs. To cater to agent initiative, we incorporate a chit-chat functionality, allowing for multi-turn casual conversations. Furthermore, we delve into a series of strategies for multi-modal interaction to continuously improve user engagement.</description>
      <pubDate>Tue, 03 Oct 2023 12:15:42 GMT</pubDate>
      <guid>https://www.amazon.science/alexa-prize/proceedings/taco-2-0-a-task-oriented-dialogue-system-with-mixed-initiatives-and-multi-modal-interaction</guid>
    </item>
    <item>
      <title>DiWBot: A cooking and DIY conversation guidance system</title>
      <link>https://www.amazon.science/alexa-prize/proceedings/diwbot-a-cooking-and-diy-conversation-guidance-system</link>
      <description>As conversational agents become sophisticated, they are becoming more prevalent in the daily lives of people. They have the ability to help users accomplish daily tasks and help in their day-to-day lives. In this report we summarize our findings and development of our taskbot: Do it With Bot (DiWBot) built to help people in cooking and DIY tasks for the 2023 Alexa Prize TaskBot Challenge. We present the engineering techniques employed to build our bot and an analysis of the conversations between our taskbot and its human users. This analysis reveals what conversation behaviors and taskbot features are preferred by the human users and which behaviors impact the ratings in a negative manner. We present this report to aid in the development of similar taskbots.</description>
      <pubDate>Tue, 03 Oct 2023 12:15:06 GMT</pubDate>
      <guid>https://www.amazon.science/alexa-prize/proceedings/diwbot-a-cooking-and-diy-conversation-guidance-system</guid>
    </item>
    <item>
      <title>EvoquerBot: A multimedia chatbot leveraging synthetic data for cross-domain assistance</title>
      <link>https://www.amazon.science/alexa-prize/proceedings/evoquerbot-a-multimedia-chatbot-leveraging-synthetic-data-for-cross-domain-assistance</link>
      <description>EvoquerBot, developed for the TaskBot challenge, is a multimedia chatbot designed to assist users in completing cooking and DIY tasks within a single session. The bot leverages a coordinated orchestration of submodules for intent classification, task recommendation, task description, and step navigation. This paper addresses the challenges of short development and model training time, data quality in both NLP and multimedia sectors, multimedia response handling, and tailoring the conversation flow to domain-specific user experiences. To overcome these, we propose agile classifier development, data augmentation, multimedia response design, and domain-specific dialogue state machines. The conversation flow is governed by an efficient intent classifier and a recursion-based state machine, further enhanced with features such as Cooking Image Augmentation and DIY Substep Decomposition. The effectiveness of our system is validated by the superior relevance of task recommendations, demonstrating its ability to enhance user experience.</description>
      <pubDate>Tue, 03 Oct 2023 12:14:29 GMT</pubDate>
      <guid>https://www.amazon.science/alexa-prize/proceedings/evoquerbot-a-multimedia-chatbot-leveraging-synthetic-data-for-cross-domain-assistance</guid>
    </item>
    <item>
      <title>Sage: A multimodal knowledge graph-based conversational agent for complex task guidance</title>
      <link>https://www.amazon.science/alexa-prize/proceedings/sage-a-multimodal-knowledge-graph-based-conversational-agent-for-complex-task-guidance</link>
      <description>This paper presents Sage, a task-oriented multimodal conversational agent devel- oped for the Alexa Prize TaskBot Challenge 2. Focusing on cooking and DIY tasks, Sage integrates task-oriented dialogues with engaging general chats for a human-like interaction model. Its innovative hierarchical dialogue state management, based on hierarchical state machines, enables a flexible conversation flow managing both cross-task and inner-task intents. To offer comprehensive task- related insights, Sage employs a multimodal task knowledge graph, integrating diverse online data with advanced image generation and large language model techniques. Moreover, Sage pioneers an open-domain intent grounding approach with a T5-based model for high-level intent classification and an LLM-based model for open-domain demand understanding. These strategies allow Sage to handle complex user requests, fostering dynamic, relevant conversations. At the end of the semifinals, Sage achieved an average rating of 3.57/5.0.</description>
      <pubDate>Tue, 03 Oct 2023 12:13:32 GMT</pubDate>
      <guid>https://www.amazon.science/alexa-prize/proceedings/sage-a-multimodal-knowledge-graph-based-conversational-agent-for-complex-task-guidance</guid>
    </item>
    <item>
      <title>TWIZ-v2: The wizard of multimodal conversational-stimulus</title>
      <link>https://www.amazon.science/alexa-prize/proceedings/twiz-the-wizard-of-multimodal-conversational-stimulus</link>
      <description>In this report, we describe the vision, challenges, and scientific contributions of the Task Wizard team, TWIZ, in the Alexa Prize TaskBot Challenge 2022. Our vision, is to build TWIZ bot as an helpful, multimodal, knowledgeable, and engaging assistant that can guide users towards the successful completion of complex manual tasks. To achieve this, we focus our efforts on three main research questions: (1) Humanly-shaped conversations, by providing information in a knowledgeable way; (2) Multimodal stimulus, making use of various modalities including voice, images, and videos; and (3) Zero-shot conversational flows, to improve the robustness of the interaction to unseen scenarios. TWIZ is an assistant capable of supporting a wide range of tasks, with several innovative features such as creative cooking, video navigation through voice, and the robust TWIZ-LLM a model trained for dialoguing about complex manual tasks. Given ratings and feedback provided by users, we observed that TWIZ bot is an effective and robust system, capable of guiding users through tasks while providing several multimodal stimuli.</description>
      <pubDate>Tue, 03 Oct 2023 12:12:15 GMT</pubDate>
      <guid>https://www.amazon.science/alexa-prize/proceedings/twiz-the-wizard-of-multimodal-conversational-stimulus</guid>
    </item>
    <item>
      <title>ISABEL: An inclusive and collaborative task-oriented dialogue system</title>
      <link>https://www.amazon.science/alexa-prize/proceedings/isabel-an-inclusive-and-collaborative-task-oriented-dialogue-system</link>
      <description>In the rapidly evolving landscape of multimodal interactive technologies, a critical gap persists in their utility and reach across diverse user populations. These technologies, while advanced, exhibit a narrow application of modalities, consequently marginalizing certain groups. Additionally, their superficial representations of context pose a challenge in accommodating users with varied preferences, cultural backgrounds, or dialects, thereby limiting their collaborative efficacy. The impoverished representations of context fail to accommodate creative, human-like, and flexible communication styles. Compounded by static generative capabilities that dampen user retention rates, these issues are further amplified in the face of safety concerns, especially in the age of large language models. This work explores these multifaceted limitations and strives to highlight avenues for enhancing accessibility, inclusivity, engagement, safety, and flexibility. We propose an IncluSive And collaBorativE ALexa skill (ISABEL). Our novel system is the first to combine diverse theories from machine learning, cognitive science, and linguistics. Pairing these theories with community outreach and co-design, we are able to build: 1. new, sophisticated representations of context that support equitable and human-like under- standing of diverse user populations; 2. a first-of-its-kind, multimodal interface &amp;#8211; co-designed with the Deaf and Hard of Hearing (DHH) community &amp;#8211; using touch and visual communication to support workflows for users with diverse capabilities; 3. and, novel, neurosymbolic strategies to incorporate new generative AI technologies and enable safe, efficient, and engaging response generation. As summarized in Figure 1, these contributions interact and culminate to achieve three primary goals in our design of ISABEL: inclusivity, human-likeness, and safety.</description>
      <pubDate>Tue, 03 Oct 2023 12:10:45 GMT</pubDate>
      <guid>https://www.amazon.science/alexa-prize/proceedings/isabel-an-inclusive-and-collaborative-task-oriented-dialogue-system</guid>
    </item>
    <item>
      <title>BoilerBot: A reliable task-oriented chatbot enhanced with large language models</title>
      <link>https://www.amazon.science/alexa-prize/proceedings/boilerbot-a-reliable-task-oriented-chatbot-enhanced-with-large-language-models</link>
      <description>This paper outlines the design and deployment of BoilerBot: a task-oriented multi- modal conversational agent developed for the Alexa Prize TaskBot 2 competition. BoilerBot features flexible response generation, leveraging large language models (LLMs) to enable adaptable user experiences. We discuss our novel contributions towards advancing state-of-the-art task-oriented conversational agents, highlighting user-facing challenges and proposing fault-tolerant, iterative solutions for carefully guided workflows that enable Alexa users to maximize BoilerBot&amp;#8217;s functionality.</description>
      <pubDate>Tue, 03 Oct 2023 12:09:45 GMT</pubDate>
      <guid>https://www.amazon.science/alexa-prize/proceedings/boilerbot-a-reliable-task-oriented-chatbot-enhanced-with-large-language-models</guid>
    </item>
    <item>
      <title>GRILLBot-v2: Generative models for multi-modal task-oriented assistance</title>
      <link>https://www.amazon.science/alexa-prize/proceedings/grillbot-v2-generative-models-for-multi-modal-task-oriented-assistance</link>
      <description>We present our Alexa TaskBot Challenge virtual assistant, GRILLBot-v2, which advances multi-modal task-oriented conversations by leveraging generative lan- guage models and automatic extraction and augmentation of interactive task data. GRILLBot-v2 is a conversational system based on open-source software and pub- licly available data. The task of manually crafting engaging conversational content is expensive and time-consuming. Conversely, using solely generative models has the danger of hallucinations or forgetting conversational history. Therefore, we propose advances in task-oriented automated data extraction and augmentation to create multi-modal corpora, including tasks, domain knowledge, and associated videos. Specifically, we include rich content such as custom jokes, system-initiative questions, and multi-modal content. Furthermore, this rich data allows us to ground our generative models to create complex and interactive conversations. For example, we make meaningful improvements to task QA by grounding answer generation on conversation history and our knowledge and task corpora, achieving a factoid accuracy of 0.94 on our test set. We also use structured information within tasks (e.g., category tags for recipes) to create a mixed-initiative approach to guide users to more relevant search results, increasing the average conversation rating by over &amp;#8764; 15% (guided search) during the semi-finals. The success of GRILLBot-v2 during the competition motivates using synthetic content from generative models for tasks by grounding them on relevant knowledge and task data.</description>
      <pubDate>Tue, 03 Oct 2023 12:08:53 GMT</pubDate>
      <guid>https://www.amazon.science/alexa-prize/proceedings/grillbot-v2-generative-models-for-multi-modal-task-oriented-assistance</guid>
    </item>
  </channel>
</rss>
