<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">lingngu</journal-id><journal-title-group><journal-title xml:lang="ru">Вестник НГУ. Серия: Лингвистика и межкультурная коммуникация</journal-title><trans-title-group xml:lang="en"><trans-title>NSU Vestnik. Series: Linguistics and Intercultural Communication</trans-title></trans-title-group></journal-title-group><issn pub-type="ppub">1818-7935</issn><publisher><publisher-name>Новосибирский государственный университет</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.25205/1818-7935-2020-18-3-16-34</article-id><article-id custom-type="elpub" pub-id-type="custom">lingngu-136</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>ПРИКЛАДНАЯ ЛИНГВИСТИКА</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="en"><subject>APPLIED LINGUISTICS</subject></subj-group></article-categories><title-group><article-title>Сравнение моделей векторного представления текстов в задаче создания чат-бота</article-title><trans-title-group xml:lang="en"><trans-title>Text Vectorization Methods for Retrieval-Based Chatbot</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-4450-2566</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Жеребцова</surname><given-names>Ю. А.</given-names></name><name name-style="western" xml:lang="en"><surname>Zherebtsova</surname><given-names>Y. A.</given-names></name></name-alternatives><email xlink:type="simple">julia.zherebtsova@gmail.com</email><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-4523-5167</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Чижик</surname><given-names>А. В.</given-names></name><name name-style="western" xml:lang="en"><surname>Chizhik</surname><given-names>A. V.</given-names></name></name-alternatives><email xlink:type="simple">afrancuzova@mail.ru</email><xref ref-type="aff" rid="aff-2"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Национальный исследовательский университет ИТМО</institution><country>Россия</country></aff><aff xml:lang="en"><institution>ITMO University</institution><country>Russian Federation</country></aff></aff-alternatives><aff-alternatives id="aff-2"><aff xml:lang="ru"><institution>Национальный исследовательский университет ИТМО; Санкт-Петербургский государственный университет</institution><country>Россия</country></aff><aff xml:lang="en"><institution>ITMO University; Saint Petersburg State University</institution><country>Russian Federation</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2020</year></pub-date><pub-date pub-type="epub"><day>10</day><month>11</month><year>2020</year></pub-date><volume>18</volume><issue>3</issue><fpage>16</fpage><lpage>34</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Жеребцова Ю.А., Чижик А.В., 2020</copyright-statement><copyright-year>2020</copyright-year><copyright-holder xml:lang="ru">Жеребцова Ю.А., Чижик А.В.</copyright-holder><copyright-holder xml:lang="en">Zherebtsova Y.A., Chizhik A.V.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://lingngu.elpub.ru/jour/article/view/136">https://lingngu.elpub.ru/jour/article/view/136</self-uri><abstract><p>На сегодняшний день одним из стремительно развивающихся направлений научных исследований является создание разговорного интеллекта, способного поддерживать полноценный человеко-машинный диалог на произвольное количество тем. Благодаря большому количеству индустриальных разработок, нуждающихся во взаимодействии гаджетов и человека, интерес к этой проблеме возрос в последние годы. В данной работе представлен краткий обзор архитектур современных разговорных агентов (чат-ботов) по выдаче ответа пользователю, выделены основные преимущества и недостатки каждого подхода. Отдельно приведен краткий обзор и сравнительный анализ актуальных на сегодняшний день методов векторизации текстовых данных в задачах создания современных разговорных агентов. Представлены результаты эксперимента по созданию русскоязычного чат-бота ранжирующего типа: проанализированы особенности открытых источников данных с диалогами на русском языке, описан алгоритм обработки собранных данных для реализации бота, ранжирования ответов и выбора ответной реплики, опубликован итоговый набор данных и программный код. Также были проанализированы проблемы чат-ботов ранжирующего типа (на примере создания бота, поддерживающего беседу по узкопрофильной теме о пленочной фотографии). Кроме того, были проанализированы особенности открытых источников данных с диалогами на русском языке, доступных на сегодняшний день, собран и проанализирован необходимый набор данных для обучения чат-бота, продемонстрирована его работа, а также количественная оценка качества ответов пользователю. Авторы раскрывают проблематику оценки качества работы чат-ботов, в частности обсуждаются вопросы выбора метрик. Также демонстрируются примеры диалогов чат-бота, реализованного на моделях векторизации, давших хорошие показатели при автоматической оценке.</p></abstract><trans-abstract xml:lang="en"><p>Nowadays, a field of dialogue systems and conversational agents is one of the rapidly growing research areas in artificial intelligence applications. Business and industry are showing increasing interest in implementing intelligent conversational agents into their products. There are numerous applications of chatbots in industry, banking, healthcare, and education; and it keeps on growing year-by-year. Many recent studies has tended to focus on possibility of creating intelligent bots helping users not only to accomplish specific tasks (by identifying their intents from text or voice conversations using artificial intelligence), but to capture the user’s identity, attributes, engagement data, and any feedback the user provides - to better handle a wide variety of conversational topics imitating human-like behavior. In this paper, we review the recent progress in developing intelligent conversational agents (or chatbots), its current architecture (rule-based, retrieval based and generative-based models) as well as discuss the main advantages and disadvantages of the approaches. Additionally, we conduct a comparative analysis of state-of-the-art text data vectorization methods (i. e. word/sentence embeddings) which we apply in implementation of a retrieval-based chatbot as an experiment. The results of the experiment are presented as a quality of the chatbot responses selection using various R10@k measures. We also focus on the features of open data sources providing dialogues in Russian. Natural language processing (NLP) techniques for the collected dialogue data are described. Both the final dataset and program code are published. In this paper, the authors also discuss the issues of assessing the quality of chatbots response selection, in particular, emphasizing the importance of choosing the proper evaluation method. We also demonstrate examples of chatbot dialogues implemented using text vectorization models (TF-IDF-weighted Word2Vec embeddings and LASER sentence embeddings) which revealed best performance. Our future work research is also briefly described in this paper.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>обработка естественного языка</kwd><kwd>компьютерная лингвистика</kwd><kwd>машинное обучение</kwd><kwd>диалоговые системы</kwd><kwd>интеллектуальные чат-боты</kwd><kwd>эмбеддинги слов</kwd><kwd>разговорный интеллект</kwd><kwd>ранжирующие чат-боты</kwd><kwd>порождающие модели</kwd><kwd>векторные представления текста</kwd></kwd-group><kwd-group xml:lang="en"><kwd>natural language processing</kwd><kwd>natural language understanding</kwd><kwd>dialogue systems</kwd><kwd>conversational AI</kwd><kwd>intelligent chatbot</kwd><kwd>retrieval-based chatbot</kwd><kwd>word embeddings</kwd><kwd>text vectorization</kwd><kwd>generative models</kwd></kwd-group></article-meta></front><back><ref-list><title>References</title></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
