# 1. Introduction

Language AI for social impact - a Playbook on how to use language technology for community engagement.

| <mark style="color:blue;">**How can CLEAR Global help?**</mark>                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                  |
| -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| <p>CLEAR Global’s mission is to help people get vital information and be heard, whatever language they speak. We help our partner organizations to listen to the communities they work with and communicate with them effectively. <br><br>Our tech-focused work helps organizations to find use cases where language technology could help to get users more actively engaged and scale up communications efforts. We develop language AI solutions such as chatbots, machine translation, and speech solutions for low-resource languages. These are languages that don’t have enough data to create such language solutions. <br><br>CLEAR Global’s user experience (UX) team can help with user research and UX design, and advise on human-centered design to tech interventions. <br><br>Our Language Services team can translate messages and documents into local languages, help with audio translations and pictures, train staff and volunteers, and advise on two-way communication. We also work with partners to field test and revise materials so they are easier to understand and have more impact. This work is backed up by research and language mapping and by assessing the communication needs of target populations.<br><br></p><p>For more information visit our <a href="https://clearglobal.org/">website</a> or contact us at <a href="mailto:info@clearglobal.org"><info@clearglobal.org></a>.</p> |

{% hint style="info" %}
**This playbook is currently still under development by CLEAR Global in collaboration with NLP experts and relevant organizations working to provide feedback and validate the playbook.**
{% endhint %}

\ <mark style="color:blue;">**Chapter 1 Overview:**</mark> In this chapter, we introduce CLEAR Global, an organization that aims to make communication across languages more effective. CLEAR Global focuses on developing language AI solutions. These include chatbots, machine translation, and speech solutions. Our main focus is on low-resource languages.

This playbook is a guide for program and technology partners and aims to help them understand and make use of language technology. It looks at various key objectives:&#x20;

* learning about language technology
* understanding relevant terms&#x20;
* choosing use cases that will have an impact&#x20;
* making communication better and helping people to work together well&#x20;
* managing data effectively using language technology
* making use of chatbots and machine translation.

The playbook focuses on practical application through real-world examples and aims to get partners to learn together. In the “Acknowledgments” section, we thank the various people who have helped with the “4 Billion Conversations” project. We highlight the work of Natural Language Processing (NLP) researchers and partner organizations. Overall, the playbook serves as a helpful and broad resource. It should help people to use language technology to improve communication and get people in their community more actively engaged.\
\
\
**Welcome to the Playbook for Language Technology!**

This playbook aims to help you understand, develop, and use language technology effectively. You may have little or no knowledge of this technology and maybe just starting to learn about it. Or you may have already used language technology in your work. Whatever the case, this playbook will be your go-to resource. It will help you understand, make use of, and maximize the potential of language technology.&#x20;

{% hint style="info" %}
**This section of the playbook requires little or no technical expertise and has been designed to be plain and concise with a focus on introducing language technology and how it can be integrated into programs, for organizations looking to use language AI for community engagement.**&#x20;
{% endhint %}

This playbook is a key resource for program and technology partners. It will help you to understand how to use and develop language technology and give you plenty of practical guidance. The focus is on low-resource and minority languages. The digital world is expanding, and language technology can help bridge gaps in communication and give people better access to information. But the major languages are dominant and this is a problem. We have written this playbook to help organizations build and make use of language technology solutions for low-resource languages. We also want to help organizations find use cases that will have an impact and work with their communities, partners, and supporters to set up successful language technology projects.


# 1.1 How to use the partner playbook

In each section of this playbook, you’ll find new information, steps you can take, and examples. These will guide you toward your goal of finding and building use cases with language technology that will have a real impact. Here are some ways you can use this playbook:

#### **a)** Step-by-step learning

Start by reading the initial chapters to get an overview of language technology. Then you can dive into more specific topics that interest you the most. Use the sidebar on the left to select the various chapters and sections. Note that the final chapters are aimed at more technical readers. The earlier ones are intended for those who are developing programs.&#x20;

#### **b)** Practical examples

The playbook is full of real-world examples and case studies. They show how language technology can be used successfully to build chatbots and machine translation solutions. Take the time to study these examples and think about how you could apply similar ideas in your context.

#### **c)** Learning together

Feel free to engage with other partners, experts, and our team. Discuss the content, share thoughts and ideas, give feedback, and ask for help where needed. Working with others makes the learning experience richer and makes it easier to overcome challenges.

**d) Using the Table of Contents**

The Table of Contents for the playbook allows you to switch between sections of the playbook in any order. This makes it easy for you to get the information that is most relevant to your experience. You can move around the different sections using the section headers on the left-hand side of the playbook. These sections also have sub-sections which are extensions of each main section. If you click on the main section, you will see the sub-sections as a drop-down menu. Each main section has a title and a number. These help to organize the playbook and guide users to the information they need.


# 1.2 Chapter overviews

Use Table 1 to guide you through this partner playbook. <br>

| Chapter                                                | Content                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| ------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Chapter 1 – Introduction                               | We introduce CLEAR Global, and guide you on how to best use this partner playbook.                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| Chapter 2 – Overview of language technology            | We introduce the basic ideas, concepts, and tools used in this field. This chapter also includes a section with a list of the most common terminology so you don’t get lost in jargons and acronyms like NLP, ASR, STT and so on.                                                                                                                                                                                                                                                                                                                               |
| Chapter 3 – Opportunities for partners                 | We look at real-life examples that show how language technology can solve practical problems. The focus here is on reaching marginalized communities using languages that don't have a lot of resources                                                                                                                                                                                                                                                                                                                                                         |
| Chapter 4 – Finding use cases that will have an impact | We explain how to find and understand situations where language technology can make an impact. We'll help you decide if an idea is realistic, and look at the results you can expect.                                                                                                                                                                                                                                                                                                                                                                           |
| Chapter 5 –Communication and working together          | <p>We look at understanding local community issues, working with partners or local communities to solve problems, and sharing the impact of the projects with these communities.<br>In <a href="https://docs.google.com/document/d/1APbhZkFf-bASY8B__r1FmEtIwU9WNoi7/edit#heading=h.dnmnw53uz44m">Chapter 5</a>, we'll also show you ways to collect, organize, and manage the data you need. This is especially relevant for languages that may not have a lot of information available. Good data is very important for language technology to work well.</p> |
| Chapter 6 – Implementing language technology           | <p>This chapter is aimed at a more technical audience to help them design and develop language technology solutions.<br><br>We present a high-level workflow to navigate the language technology landscape. We show them how to decide whether their language is suitably prepared and evaluate their solution in real life.</p>                                                                                                                                                                                                                                |
| Chapter 7 – Guidelines for development and deployment  | We present practical knowledge on making use of solutions, with a focus on chatbots and machine translation. This chapter, aimed at a technical audience, dives into building data, training models, deployment, and continuous evaluation for impact.                                                                                                                                                                                                                                                                                                          |

*Table 1: Chapter summary for the partner playbook*

<br>


# 1.3 Acknowledgements

This playbook was created as part of the 4 Billion Conversations project. The project is being run by CLEAR Global and is funded by Deutsche Gesellschaft für Internationale Zusammenarbeit (GIZ).

We would like to say thank you to the many people who helped with this document, including:

* Natural Language Processing (NLP) researchers Dr. David Adelani, Dr. Duygu Ataman, and Dr. Mathias Müller for sharing their insights and giving feedback, and
* the many partners that shared their time, knowledge and feedback with us, including Malaica, Pattan (Pakistan), Families Fit for Children (Uganda), and Reach a Hand Uganda.

<br>


# 2. Overview of Language Technology

<mark style="color:blue;">**Chapter 2 Overview:**</mark> In this chapter, we give you an overview of language technology. Language technology focuses on making systems do useful tasks with human language, whether spoken or written. This chapter highlights the huge impact language technology can have on communication, automation, and interaction.

We look at the benefits of using language technology in communication. We also look at its role in automating translation and interpreting services. It enables real-time communication with communities and makes it easier to collect and analyze data, and evaluate programs. The chapter focuses on the potential for language technology to bridge language gaps, increase inclusion, and get more people to participate in programs.

The section on potential uses makes it clear how important language technology can be in disaster response, education, healthcare, and environmental conservation. We look at the way language technology can help in these areas:

* supporting humanitarian aid&#x20;
* improving access to educational materials
* making healthcare communications more effective and
* helping the environment by passing on messages about climate change.

We present the key terminology and concepts used in language technology, explaining terms like Artificial Intelligence (AI), Automatic Speech Recognition (ASR), Text-to-Speech (TTS), Machine Translation (MT), and chatbot. Understanding these key terms will help readers to make better use of the playbook.\
\
To sum up, Chapter 2 is like a foundation. It gives you the basic knowledge you need to understand the importance of language technology. It can help you deal with communication challenges, improve inclusion, and make programs more effective in many different sectors.<br>

{% hint style="info" %}
**This section of the playbook requires little or no technical expertise and has been designed to be plain and concise with a focus on introducing language technology and how it can be integrated into programs, for organizations looking to use language AI for community engagement.**&#x20;
{% endhint %}


# 2.1 Definition and uses of language technology

There are various ways to define language technology. Here is a simple, easy description: language technology is about “getting computers to do useful things with human language, whether in spoken or written form”.

Language technology lets us automate tasks in verbal or written interaction. Examples are interactions between people, translating from one language to another, transcribing spoken words into text, or text into spoken word. Language technology can also analyze language, classify content, and make sense of the message. We call this part listening.

New developments in language technology include some interactive products. There are now machines that respond to voice commands, such as voice assistants. Speakers of major languages can communicate across language barriers using automated translation. They can use search engines to find information. Many of us have used online bots. Generative artificial intelligence (AI) has recently made huge progress.

\
Many people use these amazing technologies every day and it’s quite normal for them. But there are not many products that speakers of under-resourced languages can use. Language technology in these languages may not exist, or it may not be widely available. This means that many people are left behind. Learn more about language technology[ here](https://www.dfki.de/~hansu/LT.pdf)<br>


# 2.2 How language technology helps with communication

Language technology has completely changed the way we communicate with communities. Our communication is now more program-focused and efficient. There are many ways we can use language technology when communicating with communities. It can make the programs much more effective.

Firstly, language technology means we can automate translation and interpreting services. In multilingual settings, effective communication can be difficult. Language technology tools can help organizations to bridge gaps in language and/or literacy. For example, machine translation and speech recognition. We can make sure people understand our messages and we can reach communities with several languages. This means more people are included and can take part in programs.

Language technology also allows real-time communication with communities. By using instant messaging platforms and chatbots, organizations can respond to queries or concerns from community members immediately. People feel listened to and valued. They are more likely to trust the organization and get involved in the program.

Language technology also helps us to collect and analyze data. Organizations can use natural language processing algorithms to find things out from all the text and voice data they collect from surveys or feedback forms. This means they can identify the needs and preferences of a community. They can then make changes to the program in line with these.

A further plus point is that language technology makes it easier to evaluate programs. Automated sentiment analysis tools allow organizations to get an accurate idea of public opinion about their programs, general opinions, and current events. Organizations can also use these tools to analyze social media posts or online reviews. This allows them to find out how effective their programs are. They can make any improvements needed and get an understanding of how communities feel about the program. This information is helpful for future program design.

To sum up, language technology can be very helpful as it provides information and makes two-way communication with communities easier. From translation services to real-time communication and the ability to analyze data, these new developments can help us to scale up community communication. We continue to welcome such technological innovations in our society today. But we must also use these tools to improve understanding between diverse communities so that our programs can be more effective.

<br>

<br>


# 2.3 Areas where language technology can be used

### Technology should both listen and provide information, whatever the language                         &#x20;

A good use of language technology helps users to communicate with the system in their preferred language (context). We could also use gestures or sign languages to communicate with computers or sign language speakers. The goal is to make it easy for us to use all sorts of technology and access information from all over the world. Technology should help us at work and in our daily lives.\
\
\
Language technology helps people and systems communicate with each other

Communication between people who speak different languages can be difficult, especially if they are speakers of under-resourced languages. One of the main goals of language technology is to translate different human languages smoothly and automatically. This is not easy, but language experts have now developed software systems that make the work of human translators easier. Automatic translations aren’t always perfect, but they can be helpful when people need information in a foreign language. For example, in crises when we need to simplify large amounts of text and get people to understand it.

Disaster response and humanitarian aid                                                                                                                             Language technology can support humanitarian aid in emergencies. It can quickly translate important information for both the people affected and the people who are there to help. So everyone gets the important messages they need, even if they speak different languages. When we want to help communities that don’t usually get much attention, we can use language technology for low-resource languages to help translate important documents and instructions. We can then be sure they get the aid they need in the best way possible.<br>

Education

In education, language technology can improve access to educational materials and information that wasn’t previously available in all languages. You can use language technology to create personalized learning plans and make learning easier. The Languages in Education report by UNESCO tells us that 40% of children around the world learn in a language that they don’t speak at home (UNESCO). They have to learn the language of teaching over time. Language technology can help them to learn a new language.

\
Healthcare

In healthcare settings, effective communication and connecting with people are very important. Nowadays, healthcare institutions take a more patient-centered approach. They see connecting with people as a key priority. It helps to improve communication with the community and make services more accessible. A key part of this is talking to people in their own language and showing an understanding of their bodies and health beliefs.&#x20;

Language technology can be very useful in the field of healthcare. It can help create chatbots that give vital information on health-related topics. It can translate content for people who speak different languages. And it can make it easier for doctors and nurses to talk to patients if they don't speak the same language. Automated Q\&A or virtual assistants can give answers to frequently asked questions and reduce the number of phone calls needed. Everyone can understand their health better and get the care they need.<br>

Helping the environment

Language development projects can help to provide information about the environment in different languages. Language technology can spread messages about climate change and climate resilience in the languages of the communities affected by climate change. Language technology can also help to translate messages and educational information about taking care of nature. This means that people all around the world can learn how to protect our planet.&#x20;

Language technology plays a huge role in making sure that those who are most affected by climate change have a strong voice when they call for action from those in power. Governments can then put in place preventive, reparative, and climate-friendly policies and measures. More people will then know about environmental problems and how to solve them, and can help to find solutions.

<br>


# 2.4 Key terminology and concepts

This section looks at the most common terms and concepts used in language technology. These terms will help readers to make good use of the playbook. They will have a clear understanding of the terminology, abbreviations, and concepts used in language technology.\
\
**Language Technology (LT):** Information technologies that focus on human language, in both spoken and written forms. Human Language Technology includes various technologies that are used to process, understand, and generate language. Imagine the helpful tools on your phone or computer that understand and generate words, like language translation apps or voice assistants. These technologies help us to communicate and interact with devices using language. When you type a message on your phone and it suggests the next word, or you talk to a virtual assistant like Siri or Alexa, this is language technology in action. It's like having a smart friend who understands and responds to what you say. It makes technology more accessible and user-friendly.

\
**Artificial Intelligence (AI):** Artificial Intelligence (AI) refers to computer systems that can do tasks that normally require human intelligence. AI is like giving computers the ability to do things that usually only humans can do. Imagine having a smart friend who can learn, make decisions, and solve problems but doesn’t need to be programmed for each task.

**Automatic Speech Recognition (ASR) or Speech-to-Text (STT):** ASR converts spoken language into text. It's used in voice assistants and transcription services, for example. The terms ASR and STT are both used and they are almost the same. But STT can be a semi-manual process, while ASR is fully automated. ASR is like a smart listener that turns spoken words into written text on your device. This process is usually automated while involves a bit of human touch, a bit of fine-tuning to make sure every word is correct.\
\
**Text-to-Speech (TTS):** TTS technology converts written text into spoken language. It's used in screen readers and virtual assistants, and to generate audio content. You can find TTS on accessibility tools like screen readers. It brings written content to life through spoken words. Imagine your phone reading out loud the messages you get, or an audiobook telling your favorite story. This is made possible by TTS technology.<br>

**Machine Translation (MT):** MT is the technology that uses algorithms and models to automatically translate text or speech from one language to another. It's the tech behind apps and systems that make your friend's text in Swahili appear in English. It lets you connect with people around the world with no language barriers.

**Speech Translation (STS):** STS technology translates spoken language from one language to another in spoken or written form. It uses ASR, MT, and TTS and allows you to have a seamless conversation. If you're talking, it turns your words into written text (ASR) and translates them into the other language (MT). It can then even read the translated message out loud (TTS). You can have a smooth and natural chat with someone who speaks a language you don't. Your communication will feel effortless.

**Language Model (LM):** A language model is a statistical or machine learning model that looks at the structure and patterns of written language. Large Language Models (LLMs) are a type of LM that has been in the focus because they can generate meaningful text that is relevant to a context. These models are often trained on massive amounts of text data. They can be fine-tuned for specific tasks such as completing texts, generating language, and even translation.&#x20;

**Chatbot:** An application that uses Artificial Intelligence (AI) to hold conversations that feel like human conversations. The interactions are usually text-based. Imagine having a friendly virtual assistant on your computer or phone who is trained to give you certain information. A chatbot is like a digital buddy who uses artificial intelligence to chat with you. It’s like chatting with a real person but via text messages.<br>

**Conversational AI:** A branch of artificial intelligence that aims to create natural and engaging conversations between machines and humans. This technology makes interactions with machines feel more like talking to a person. It makes it easier and more fun to get things done. Whether you need answers, want to set a reminder, or just have a casual chat, Conversational AI adds a human touch to your interactions with technology.

**Dialogue system:** The technical framework that allows chatbots to have back-and-forth conversations with users. A dialogue system is the technology behind a chatbot. It makes them capable of having meaningful back-and-forth conversations with you. Imagine interacting with a virtual assistant that not only understands what you say but also responds in a way that feels like a real conversation.

**Knowledge technologies:** Technologies that form a bridge between language and concepts or tasks in the real world. They connect an understanding of language to knowledge and information in a specific field. Knowledge technologies are like digital interpreters. They connect language with real-world concepts and tasks. This makes information easier to access and more useful. Imagine you're talking to a super-smart assistant who not only understands what you're saying but also knows a lot about specific topics.<br>

**Natural Language Processing (NLP):** A field of artificial intelligence that gives computers the ability to understand, interpret, and generate human language. It includes various language technology tasks, such as sentiment analysis and text summarization. Natural Language Processing is like giving super-smart brains to computers so they can understand, talk, and write like humans. Imagine your computer not just reading your words but understanding the feelings behind them. Or writing a summary of a long article for you.

\
**Natural Language Understanding (NLU):** The process of giving machines the ability to understand the meaning and context of human language. This makes it possible to have more complicated interactions. NLU is what allows computers or systems to understand a voice request. For example, when you ask them to find the right music, even without using specific technical terms. It's like giving machines the ability to understand the deeper meaning and essence of our words. Interactions with technology can be much smoother and more human-like.

\
**Annotated Data:** Data that has been manually labeled with specific information. This might be named entities, transcription, sentiment labels, or part-of-speech tags. Annotated data is information that has been carefully tagged or labeled by humans so that computers can understand it more easily. It's like providing special notes or labels on a map to help someone (in this case, a computer) find their way and understand different aspects of a large amount of information.

\
**Graphical Processing Units (GPUs):** Specialized electronic circuits that speed up graphics rendering. They are widely used in language technology for parallel processing tasks, like developing models. Similarly, a Tensor Processing Unit (TPU) is a specialized kind of hardware that speeds up machine learning workloads, giving better performance for deep learning tasks.


# 3. Partner Opportunities

<mark style="color:blue;">**Chapter 3 Overview:**</mark> *In this chapter, the playbook explores the numerous opportunities that language technology offers to organizations aiming to enhance efficiency and broaden their reach. By integrating language technology into their operations, organizations can utilize applications such as search engines and chatbots to provide information and amplify the voices of those they serve in accessible ways.*

*The chapter delves into several real-world examples of language technology applications developed by CLEAR Global, showcasing their impact across various sectors. These examples include TILES, a voice-enabled AI kiosk for climate resilience in rural areas; Kompas, an AI solution for crisis response in Ukraine; MT Rwanda, a machine translation initiative improving learning experiences; Shehu, a chatbot addressing COVID-19 queries in Nigeria; and a multilingual chatbot facilitating business registration in Kenya.*

*Additionally, the playbook highlights non-CLEAR Global examples like Me Bote na Conversa, a chatbot translating corporate terms in Brazil; and MAIA, a chatbot combatting gender-based harassment in Brazil. It also introduces UNICEF's U-Report Information Chatbot, fostering youth engagement.*

*The chapter emphasizes the importance of bridging the technical gap within organizations, providing principles and best practices for impactful solutions. It underscores the significance of a user-centered approach, well-defined use cases, and continuous user involvement.*

*Furthermore, the playbook guides organizations in engaging with language technology providers. It suggests considering factors such as expertise, collaborative attitude, resource availability, values alignment, past collaborations, communication skills, geographic considerations, size and scale, and network reach when selecting a partner. The playbook recommends appointing a dedicated team to ensure effective communication and collaboration between technical and non-technical teams, especially in the development of language technology solutions.*

{% hint style="info" %}
**This section of the playbook requires little or no technical expertise and has been designed to be plain and concise with a focus on introducing language technology and how it can be integrated into programs, for organizations looking to use language AI for community engagement.**&#x20;
{% endhint %}

Language technology presents opportunities for organizations to enhance their efficiency and reach. By integrating language technology into their processes, organizations can use applications such as internet search engines, spoken language dialog systems (e.g., chatbots), and much more. With language technology, organizations can help many more of the people they serve get information or have their voices heard, in ways that are simple and accessible.


# 3.1 Enabling Organizations with Language Technology

We look at several real-world examples of language technology applications that CLEAR Global has developed. We showcase their impact across various sectors.&#x20;

For example:&#x20;

* TILES, a voice-based AI kiosk to make people in rural areas more resilient to climate change&#x20;
* Kompas, an AI solution for crisis response in Ukraine&#x20;
* MT Rwanda, a machine translation project to make learning easier&#x20;
* Shehu, a chatbot that answers questions on COVID-19 in Nigeria and
* a multilingual chatbot that helps people to register businesses in Kenya.<br>

The playbook also highlights some non-CLEAR Global examples.

For example:&#x20;

* Me Bote na Conversa, a chatbot that translates business terms in Brazil and
* MAIA, a chatbot that combats gender-based harassment, also in Brazil.&#x20;

We also look at UNICEF's U-Report Information Chatbot, which aims to get more youth engagement.

\
We discuss the importance of bridging the technical gap within organizations. We also provide principles and best practices for solutions that will have an impact.&#x20;

We focus on:

* a user-centered approach
* well-defined use cases and
* strong involvement of users.

The playbook also helps organizations to engage with providers of language technology.&#x20;

When you choose a partner, there are lots of factors to think about:\ <br>

* Do they have the right expertise?
* Can you work well with them?
* Do they have the resources?
* Do you have the same values?&#x20;
* Who else have they worked with?
* Do they have good communication skills?
* Where are they based?
* How big are they?
* Do they have a wide network?

We suggest that you build a team for this task. This should make sure that communication between technical and non-technical teams is effective, and that they work together well. This is very important when developing language technology solutions.

There are many use cases for language technology across various sectors. Here are some examples to show how we’ve used language technology in our work in the past.

#### CLEAR Global Examples:

**Using language technology for climate resilience – TILES**

CLEAR Global worked with partner organizations to develop a voice-based multilingual AI kiosk solution for areas where there is low literacy and low-connectivity. It’s called TILES (Touch Interface for Language Enabled Service). Users can use voice and gestures to ask questions via a screen and get answers in their chosen language.

In many areas, for example, rural areas, the most marginalized people don’t have the same access to key information that many of us do. They have to ask aid workers, village elders, friends, neighbors, and relatives to answer their questions. They are dependent on other people being there to help. Sometimes the information is no longer accurate after being passed from person to person. They may also need information that no one in the community can help them with.

The problem is even worse if the person doesn’t speak the majority language. Many people don’t have access to the internet. Data coverage may be weak (3G does not give a person full access to the internet), data plans are expensive, and they may not have access to a device. Information in the right language may be unavailable or incomplete. \
\
CLEAR Global developed the Tiles device together with **Gram Vaani** in India. They wanted to help farmers to switch to sustainable farming methods and to make them more resilient to climate change. With this technology-powered solution, farmers could ask questions in their own language using their voice. The device used pre-recorded audio or visual content to answer the farmers’ questions. It was used in areas facing the effects of climate change.<br>

**Using language technology to respond to crises – Kompas**&#x20;

The latest UNHCR figures say that around 7.9 million refugees have left Ukraine for Europe since February 2022. Over 6.5 million people are displaced within Ukraine. Too much information was spread and this made it difficult for refugees to find facts they could rely on. A lot of information was posted by people who wanted to help. It was hard to know if the news was true, or if it was up to date. Trying to decide which information was correct took up a lot of time and was stressful. There was also a risk that people would follow incorrect or old information about the current situation.\
\
CLEAR Global wanted to help so they developed Kompas. This is a multilingual artificial intelligence tool that allows people to search for information. Its sources are carefully chosen, checked, and up to date. It uses channels and websites that are already popular with the affected people. It gives people affected by the war in Ukraine easy access to safe and usable information that they may need and want, in their chosen language.&#x20;

Kompas worked in parallel with the efforts of United for Ukraine and other major information providers in Ukraine and provided a system for accessing information. Users can search for relevant information by asking open-ended questions. Service providers also collected the questions and got feedback on how useful they were. They could then adapt their services and add any information that was missing.<br>

**Improving learner experiences in the education sector through language technology – MT Rwanda**

Another example of how language technology is used is the MT Rwanda Machine Translation project. This was set up by CLEAR Global and GIZ Digital Solutions for Sustainable Development (DSSD) in Rwanda. The project aimed to make the preconditions better for the use of machine translation in the public sector and the digital ecosystem. CLEAR Global worked with local Kigali-based technical partner Digital Umuganda to create a digital learning platform and look at how machine translation could make the user and learner experience better.\
\
The platform offers tailored learning and training of skills that are in demand. There are over 300 courses in various subjects. However, research showed that people in Rwanda were not using the platform because there was not much content in the local language.&#x20;

\
The use case for machine translation in the education sector included:\
\
1\) localized instructions in Kinyarwanda (Non-MT), so that users could find their way around the platform in Kinyarwanda

2\) complete subject translation meant that all the courses could be translated into Kinyarwanda and

3\) in-line translation allowed learners to highlight words and phrases within the courses and get a translation into Kinyarwanda.\
\
This meant users could get more context and a better understanding of professional terminology. They could also improve their English skills. Getting the community involved was also important. A group of Kinyarwanda speakers created, translated, and checked the data sets needed to train and build the translation model.

#### **Ensuring Inclusive Accessibility in Humanitarian Response - Shehu**

The CLEAR Global team worked with Mercy Corps to create a chatbot in Borno State, Nigeria. The aim was to answer questions about the coronavirus in English, Hausa, and Kanuri. The chatbot, called Shehu, provided accurate, up-to-date information. It also dealt with issues that people were worried about. This helped to stop people from getting incorrect information.\
\
Shehu answered COVID-19 questions in the user’s chosen language. This helped to build strong relationships, trust, and awareness. Because it was multilingual, it gave more people access to key information, for example, vaccine details and directions to vaccination sites.

\
**Financial inclusion and multilingual chatbot – business registration process made simple**

Another example where language technology helped to overcome a socio-economic challenge was the GIZ Fair Forward – Artificial Intelligence for All project in Kenya. This was run in partnership with THiNK and Made by People. Dealing with government services can be difficult and most information is in English. This project kicked off with a human-centered design research phase. The aim of this was to better understand the barriers in communication and language that Kenyan citizens were facing.&#x20;

The design process showed that common but important tasks like registering a business online are long and difficult. This is because the legal terminology is complex and usually only in English. Most business owners had similar questions and issues and faced very similar challenges with the registration process. This research led to the design of an AI-powered conversational chatbot that used Swahili, English, and Sheng. The chatbot could answer the most frequently asked questions, guide users through the registration process, and listen to the issues they face. This helped business owners to work through the complex registration process. It also gave phone line staff more time to focus on the more difficult cases.&#x20;

Some of the benefits:

* easy access to information on business registration&#x20;
* simplified, non-technical content that is easy to understand&#x20;
* accurate content&#x20;
* answers to user questions are regularly updated and&#x20;
* fast responses and feedback.

This was an innovative solution to some of the challenges identified. It also encouraged more people to set up businesses. For example, people who had access to technology but found it difficult to register a business in Kenya.

#### Non-CLEAR Global examples:

**An Inclusive chatbot that translates business terms into English on WhatsApp – Me Bote Na Conversa**

In Brazil, they often use English words in business language. But a large percentage of the population doesn’t speak good English. This makes professionals feel insecure and unprepared. Meta wanted to help and worked with Indique uma Preta and MOOC to develop a chatbot named Me Bote na Conversa. This chatbot translates English terms into Portuguese and explains them, allowing much broader communication.

Me Bote na Conversa aims to take away the language barrier that is common in business settings. This is very important as only 5% of Brazilians speak good English\*, even though English terms are used a lot in everyday business activities. Words like "budget," "approach," and "briefing" are now very common but not everyone understands them.\
\
The chatbot works via WhatsApp and is easy to access and free. Indique uma Preta helped to develop the database, which includes over 350 useful business terms. Users can also suggest further terms to add.<br>

**Using language technology to combat gender-based harassment and violence – Chatbot MAIA**

According to statistics provided by the United Nations (UN), a staggering 95% of online harassment targets women. In response to this concerning trend, Weni has introduced Maia ("My Artificial Intelligence Friend" in English), a virtual assistant designed to aid young girls. Its primary focus is to help these individuals recognize signs of an abusive relationship and offer guidance on the appropriate action to take in such situations. The launch of Maia forms an integral part of the broader #NamoroLegal campaign, an initiative spearheaded by the São Paulo Public Ministry (MPSP) and supported by Microsoft Brazil.

#### **Empowering and connecting young people around the world to be agents of change - UNICEF U-Report**

The chatbot works via WhatsApp, and is easy to access and free. Indique uma Preta helped to develop the database, which includes over 350 useful business terms. Users can also sug UNICEF U-Report

U-Report Information Chatbot is a social platform that can bring about big changes. It was set up by UNICEF, and you can access it through SMS, Facebook, and Twitter, for example. This innovative platform gives young people the chance to voice their opinions and get involved. They can become catalysts for positive change within their societies. The mechanism works by gathering points of view and data from young persons on a wide range of subjects that are important to them. These subjects include job opportunities, dealing with prejudice, and child marriage.

Through U-Report, these young people can take part in polls, raise awareness about key issues, and stand up for the rights of children. Most importantly, the information obtained and the findings are then passed back to the communities they came from. The findings are also sent to policymakers, who have the authority to make decisions that directly impact the lives of young people.

The U-Report for Humanitarian Action initiative is a joint effort of the Office of Innovation, Program Division, Communication for Development, and Office of Emergency Programs. In February 2020, they developed a U-Report Information Chatbot. They wanted to improve communication on the risks of COVID-19 and get more people involved. The chatbot uses languages like English, Spanish, French, Bahasa, Arabic, Vietnamese and Thai. \
\
U-Report is like a mouthpiece. It gives young people a stronger voice and channels their worries and hopes into real change. This allows young people to make a helpful contribution. It also helps to connect communities and policymakers, so they can work together to create a better future for the youth.\ <br>


# 3.2 Bridging the Technical Gap

Some organizations have a technical team who have already developed and launched digital tools. Other organizations are just starting and have very little or no technical capacity. Whatever your situation, there are some principles and best practices that your organization should follow. This will help you to build solutions that have a positive impact.&#x20;

If a tech solution is to be useful and valuable, you’ll need to consider the following:

* To start the process, you need to define the problem and understand the target audience’s needs. You also need to find out if they have access to technology. And you need to know what languages they speak.&#x20;
* It’s important to build a solid use case. To develop a use case, the organization needs to know who they want to reach. They need to know where the communication gaps are, and what impact they are hoping for.&#x20;
* Technical solutions should start with the problem and aim to solve it, not the other way around.&#x20;
* The solution should become part of programs that are already running and should deal with current communication problems.&#x20;
* A well-defined use case and problem statement will make the project more successful. It will also have a bigger impact than if the solution is driven by tech.&#x20;
* It’s important to involve users, the target audience, and those working on the design process. They should be involved at all stages. This will make sure that the solution is user-centered. It should then also reflect the needs and hopes of the target audience and the aims of the organization.&#x20;

To learn more about how to develop a strong use case and identify program requirements, please see Section [4. Finding use cases that will have an impact](https://docs.google.com/document/d/1APbhZkFf-bASY8B__r1FmEtIwU9WNoi7/edit#heading=h.s6p46cjp0cxz)


# 3.3 Dealing with language technology providers

Once a use case is developed, it will be easier to identify and select an appropriate technology provider. The technical development work can then be carried out by partners who specialize in the field and have a good understanding of the processes around building or deploying digital solutions, including language technology.

Areas that are useful to explore, define, and consider when choosing a technical partner:

* **Relevant expertise**: Assess the partner’s expertise in language technology. Look for partners with experience and knowledge that directly relate to the objectives of your project.
* **Collaborative attitude**: Choose partners who are willing to engage in a collaborative process. A partner who is open to sharing feedback, participating in discussions, and contributing insights will enhance the delivery process.
* **Resource availability**: Consider the partner’s availability and resources to actively participate in the process. Ensure they have the time and the right personnel assigned to work effectively.
* **Similar values and goals**: Select partners whose organizational values and goals align with yours. This alignment will create a more cohesive partnership and increase the likelihood of a successful collaboration.
* **Past collaborations**: If possible, review the partner’s history of collaborations. This can provide insights into their approach and commitment.
* **Communication skills**: Effective communication is crucial for a successful tech delivery process. Choose partners who are articulate, responsive, and capable of clear and constructive communication.
* **Geographic considerations:** Depending on the nature of your project and process, consider whether geographic location matters. Local partners might offer more direct engagement opportunities.
* **Size and scale**: Consider the size and scale of the partner’s organization. Larger organizations might offer a broader range of perspectives and resources, while smaller organizations could provide more focused, personalized feedback.
* **Network and reach**: Partners with a wide network or reach can help amplify the impact of the project by also involving a larger community of stakeholders, such as linguists or communities of speakers of marginalized languages.

Many types of partners can be considered, including non-profit organizations or social enterprises that have tech for social good as their main focus. There are also opportunities to reach out to commercial for-profit organizations that may do in-kind work, although it is still advised to formalize the collaboration to establish responsibilities, processes, communication, and accountability, as well as continuity.

Building teams for language tech projects\
\
For a language technology project to be successful, we suggest that the owner of the use case forms a team. This could even be a small volunteer group. You should have at least one focal point in this team who represents the needs of the organization and the users. Doing this will make sure that everyone understands their needs and stays in contact with the target audience. It should also ensure that the project is completed on time and as agreed.

When you are developing language technology solutions, it is very important to work together closely. Building a community to work together means you can develop your expertise together. It also makes sure that the solution design fits to the specific domain, languages spoken, access to technology, and data factors. This is especially true in small volunteer teams. Working together in this way will make your project more effective. It also means people have support and can learn and grow within the team.


# 4. Identifying Impactful Use Cases

<mark style="color:blue;">**Chapter 4 Overview:**</mark>

Chapter overview:&#x20;

In this chapter, we look at the way organizations select language technology and how they can use it effectively. We compare the process to choosing the right tool for the job. We also explain the importance of making sure the use of technology fits with the organization’s goals, the workability of the idea, and the overall benefits.

This chapter also helps organizations to work through the process of finding use cases for language technology that will have a strong impact. The focus is on the needs of users, workability, and continuous improvement. We offer practical steps so that organizations can make sure the technology they introduce fits in with their mission and goals.

When we bring language technology into our everyday lives, it’s important to make smart choices about where and how to use it. Think of it as picking the perfect tool for the job. Our success in using language technology will depend on how well it fits with what we want to achieve, how workable it is to use, and what benefits it brings to the organization.

Use Table 2 from this partner playbook to get a summary of the sections in Chapter 4.

<br>

| Section                                             | Content                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| --------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 4.1 –  Setting criteria to help choose the use case | We show organizations how to set clear objectives so that they can use language technology effectively. They need to identify the challenges that language technology can overcome. They also need to understand their technical capabilities and evaluate the impact the technology might have. And they need to make sure it meets the needs of stakeholders and can be scaled up.                                                                                                                                                                                                               |
| 4.2 – Carrying out a needs assessment               | <p>It is very important to do a needs assessment using the principles of human-centered design (HCD).</p><p>To do this, you need to:</p><ul><li>identify stakeholders and end users </li><li>collect existing information </li><li>build empathy through interviews or surveys </li><li>create user personas </li><li>map user journeys </li><li>define and prioritize problems </li><li>validate findings and </li><li>share results. </li></ul><p>This process forms the basis for solutions that will have an impact. This is because they reflect an in-depth understanding of user needs.</p> |
| 4.3 – Evaluating what can be done and what works    | To check whether the project is workable, the playbook suggests small-scale proof of concepts (POC) to test language technology solutions. You will need to define key performance indicators (KPIs) to measure success, and carry out a cost-benefit analysis. You also need to identify risks and develop strategies to deal with them. The chapter suggests a step-by-step approach so you can keep checking your progress and making improvements. It’s important that you can adapt if there are changes in needs or in technology.                                                           |

Table 2: Section summary for Chapter 4

<br>

{% hint style="info" %}
**This section of the playbook requires little or no technical expertise and has been designed to be plain and concise with a focus on introducing language technology and how it can be integrated into programs, for organizations looking to use language AI for community engagement.**&#x20;
{% endhint %}

When it comes to bringing language technology into our everyday lives, making smart choices about where and how to use it is important. Think of it as picking the perfect tool for a job. The success of using language technology depends on how well it fits with what we want to achieve, how feasible it is to use, and how beneficial it can be to the organization.


# 4.1 Setting criteria to help choose the use case

If your organization wants to start using language technology in your work, there are several criteria you can use to choose a use case for language technology. The first step is to outline the goals you hope to achieve by using language technology.\
\
Setting goals

You can do this by identifying the challenges that language technology can help you with. For example:

* improving customer service&#x20;
* boosting productivity&#x20;
* improving communication or&#x20;
* making a service more accessible to marginalized groups and communities.

Technical capacity

The next step is to understand your technical capacity. Is there a gap between what you want to do and your technical skills? If you identify this technical gap, it will make it easier when you are working with technology partners.\
\
Potential impact

It’s important to look at the impact language technology could have. This will help you decide if a use case is worth following up.

You can ask questions like:

* “Will this improve the daily lives of its users?”&#x20;
* “Will more people have access to our service?”&#x20;
* “Will this make internal or external teams more efficient?” or&#x20;
* “Will it cut the cost of achieving our goals?” <br>

Meeting the needs of stakeholders and scalability

Another way of choosing use cases with an impact is by simply talking to stakeholders. You need to understand their pain points and find out what they expect when it comes to communication. This will also help you to understand whether the use case can be scaled up. Can the use case be expanded? Can it be used in other areas or projects within the organization?

<br>


# 4.2 Conducting A Needs Assessment

A needs assessment is a crucial first step. It will make sure that the solutions you develop are a good fit for the needs and wishes of the end users. To do a needs assessment, you need to collect and analyze information so you can understand the challenges, needs, and desires of the audience. You can achieve this by following the principles of human-centered design (HCD).&#x20;

In the paragraphs below, you will find a breakdown showing you how to carry out a needs assessment using HCD. You can read more about HCD and design thinking [here](https://www.ideo.org/tools). <br>

Identify stakeholders and end users

First you need to identify the individuals and groups that will be affected by the solution. This means both primary end-users and secondary stakeholders. You need to understand their characteristics, behaviors, and points of view. Then you can create user personas that reflect the user group that you are designing the solution for.<br>

Collect existing information

Start by collecting existing data, reports, and research that is relevant to the problem. This will help you to understand the context and any existing projects related to the challenge you are dealing with.

Understand and build empathy&#x20;

Do interviews and surveys, or set up focus group discussions to listen to your target audience. Ask open-ended questions so people can share their experiences, needs, frustrations, and hopes on this topic. It’s also helpful to look at the daily life of individual communities. It can help you to understand where and how people face communication challenges and how they manage them.&#x20;

&#x20;\
Build user personas

Use the information you have collected to build user personas. A user persona is a like a character in a story. It is a description of a person in one of your target groups. It describes that group by stating common challenges, goals, wishes, and characteristics. For example, the language they speak and their reading level. This is key information to get an in-depth understanding of your target audience.&#x20;

\
Map user journeys

Create user journey maps so you can picture the complete experience of the end users. This helps to identify frustrations and opportunities. It will show you where the solution can make an impact.

\
Define problems and prioritize them

Once you have a better understanding of the challenges, frustrations, goals, and needs, you’ll need to produce a clear and distinct problem statement. This will say what you want to focus on. Not all challenges can be solved in one go, so set clear priorities.

\
Checking and feedback

Now you need to check that the problem and needs you have identified are correct. You can do this by sharing them with stakeholders and end users. Collect feedback to correct them if needed. This checking process makes sure that your findings from the analysis agree with the community’s point of view.

\
Sharing findings

Share the findings from the needs assessment with all relevant stakeholders. This ensures transparency and makes sure that everyone has the same understanding. This is very important if you want to work together effectively.

The needs assessment phase in HCD is the preparation stage. The next stages are forming ideas, making prototypes, and testing. By getting an in-depth understanding of the end users and their needs, HCD practitioners can design solutions that fit well with their audience. They can also deal with any difficulties. Key elements of this process are empathy, close working relationships, and a strong desire to create solutions that are meaningful and have a real impact.

<br>


# 4.3 Evaluating What Can Be Done and What Works

You need to decide if it is possible and practical to use language technology for the purpose you've chosen. How much will it improve things? To find out whether a use is workable and will have the impact you hope for, you need to test it out without taking any risks. And you need to get data that will show that you need the use case. One way to do this is by developing proof of concepts (POC). You can then measure the performance of the POC, carry out a cost-benefit analysis, and identify and deal with any risks involved with the specific use cases. We’ll explain these now.<br>

Conduct small-scale proof of concepts to see if the language technology solutions are workable for the use case you’ve chosen.

A proof of concept is like a trial run. It involves trying out small versions of your language technology solutions to see if they work as you expect. You can catch any problems or challenges early on, so you can make improvements before moving forward to a larger version.

\
Define key performance indicators (KPIs) to measure the success of your technology.\
Performance indicators are like measuring sticks. They help you see how well your language technology solution is doing. By choosing certain measurements, you can track the progress and success of the use case. For example, how accurately it translates or how quickly it responds. User satisfaction is also important. It tells you if people find the technology helpful and easy to use.

\
Look at the cost of language technology and compare it with the expected benefits.

This means looking at how much money and resources you will need to bring in the technology for the use case you’ve chosen. You will also need to look at the benefits or positive results you expect to get. This analysis helps you to compare the costs and benefits and make smart decisions.

\
Identify the risks of the use case and develop strategies to deal with them.

Risks are things that could go wrong or cause problems when you use language technology. It's important to think about these ahead of time and come up with ways to handle them if they happen. These strategies are like backup plans that help you deal with the risks and minimize their impact. One potential risk might be problems with accuracy and quality. You can deal with this by asking human experts to regularly test and check the accuracy of language technology solutions. You can also set up quality assurance processes to catch and correct errors.

\
When you introduce language technology for low-resource languages, these concepts are very important. They help to make sure that your plans are practical and that your technology works well. They also make sure that you achieve what you want, that it’s worth the cost, and that you’re ready for any surprises along the way.\
\
Use a step-by-step approach so you can keep assessing and improving language technology solutions.

As you start to introduce use cases and collect data on their impact, you’ll need to review your priorities regularly and identify areas for improvement. The needs of the organization or the end user may change. New opportunities may come up, and new technology may emerge. All these things can affect your decision-making about language technology. If you review your plans regularly, you can adapt your strategy when needed and make sure it is effective over time. You can also deliver added value to the target user groups and the wider organization.<br>


# 5 Communication and working together

<mark style="color:blue;">**Chapter 5 Overview:**</mark>\
In this chapter, we share some useful tips on effective communication and good working relationships with communities and partners. We show how they can make your language technology projects more successful. These tips are based on our own experience.

\
The chapter also suggests that you share best practices and lessons learned. We advise you to be fully transparent in the partnership so you can make improvements at any stage.

\
In summary, this chapter focuses on the different types of communication needed for successful language technology projects. We outline strategies for effective communication within language communities. We also show you how to build strong working partnerships. The focus is on transparency, trust, and learning together.

\
Use Table 3 from this partner playbook to get a summary of the sections in Chapter 5.

<br>

| Section                                            | Content                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| -------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| 5.1 –  Communicating with communities              | <p>To develop language technology solutions, you often need a very large amount of language data. In this chapter, we explain the importance of working closely with language communities to create this data. Strategies used may include hiring translators, paying local people for shorter translations, and asking for volunteers to do translations. It is very important that you are clear in your communications about translation. You need to recognize that it is a nuanced process that requires an understanding of the local culture.</p><p><br></p><p>To communicate well, you will need to:</p><ul><li>be clear about your expectations</li><li>give translators detailed instructions </li><li>ensure open dialogue to sort out any issues and </li></ul><p>make it possible for translators to work together.</p> |
| 5.2 – Communicating and working well with partners | <p>This part looks at the special approach to communication that you need when working with partners. The focus is on long-term thinking and the importance of having clear roles and responsibilities so the partnership will last. We talk about building trust, a gradual process in which each partner has their own unique strengths. </p><p><br></p><p>Effective communication means being clear about the expertise that each partner can offer and creating multi-level communication channels. We underline the need for varied channels based on type of information, urgency, and privacy needs. You can make it easier to interact with the user organization through local partner ownership and multichannel engagement. This ensures a user-centered approach. </p>                                                   |

*Table 3: Section summary for Chapter 5*

<br>

*5.2 Communication and Collaboration with Partners:*

*This part delves into the distinctive communication approach needed when collaborating with partners. It stresses long-term thinking, emphasizing defining roles and responsibilities for enduring partnerships. Building trust is highlighted as a gradual process acknowledging the unique strengths of each partner. Effective communication involves clearly articulating expertise contributions and establishing multi-level communication channels. It underlines the need for varied channels based on information types, urgency, and privacy considerations. User-organization interaction is facilitated through local partner ownership and multichannel engagement, ensuring a user-centered approach. The chapter advocates for sharing best practices, and lessons learned, and maintaining transparency throughout the partnership to foster an environment of continuous improvement.*

*In summary, this chapter emphasizes the nuanced communication required for successful language technology projects. It outlines strategies for effective communication within language communities and offers insights into building strong, collaborative partnerships with a focus on transparency, trust, and mutual learning.*

{% hint style="info" %}
**This section of the playbook requires little or no technical expertise and has been designed to be plain and concise with a focus on introducing language technology and how it can be integrated into programs, for organizations looking to use language AI for community engagement.**&#x20;
{% endhint %}

In this section, we are sharing useful tips gathered from experience on how effective communication and collaboration with communities and partners contribute to successful language technology projects.


# 5.1 Communicating with Communities

To develop language tech solutions, you need a very large amount of language data to build machine translation solutions. This is especially true for low-resource languages as there are not many datasets available in these languages. People speak these languages in areas across Asia, South America and Africa. Often there is no good-quality data available in these languages. They aren’t seen as the language of the internet or the power language. <br>

When building language technology solutions to improve communications in these low-resource languages, it’s important to work with the local communities. They can help to create the data for such language technology solutions. They can also make sure the data is accurate and includes cultural nuances, which play a key role in communication in these languages. This means, for example, asking for help from Swahili communities when setting up projects to build STT or ASR solutions in Swahili. This will ensure that end users of the language technology solution will have a good level of understanding.<br>

Building datasets means setting up teams within the community to help deliver solutions that can support the use of language technology. You can do things like hire translators to convert a lot of text or pay locals to translate shorter pieces. You can also ask volunteers or members of the community to do translations in their free time.

To follow are some key points about the approach to take when creating datasets. This will help you to deliver language technology that aids communication:

* It’s very important to explain how data helps to build better language technology solutions if you want to transfer language from humans to models effectively and accurately. Language technology plays a key role in bridging the gap between different cultures, languages, and societies. But it’s important to understand that using language technology is not just a mechanical process of converting words from one language to another. It also needs to convey meaning and cultural nuance.                                                                                                                                <br>
* If you want to communicate effectively about introducing language technology, it’s important to tell people what you expect and give them clear guidelines. Organizations should give translators or volunteers detailed instructions. These should explain the purpose, target audience, and desired tone of their translations. This helps them to understand the context and to adapt their work if needed, so they can provide better data.

\
To sum up, for effective communication about translation and the data collection process, volunteers, translators and organizations need to set clear expectations. There also needs to be open dialogue between the volunteers themselves. If we keep these channels of communication open, we can ensure accurate translations with the correct meanings. We can then use these to break down language barriers.

* Communication between the volunteers themselves is also vital. This will make sure terminology and style are consistent across various projects. Native speakers understand that there are different ways to refer to the same word in languages like Swahili, Hausa and Somali. Platforms that allow volunteers and translators to work together mean they can share knowledge, ask questions, and ask one another for advice. They are very important in ensuring a smooth and effective process for data collection when building language technology solutions.<br>
* Translators or volunteers also need to communicate with organizations if they come across any challenges or things that are not clear. If there are any terms or phrases that have multiple meanings or a special cultural meaning, they should ask for clarification. Open dialogue between volunteers and organizations means that any issues can be dealt with quickly. The resulting translations will then be more accurate.

<br>

<br>


# 5.2 Communicating and working well with partners

Handling communication with partners while working together on a language tech project requires a different approach to communication with language communities. This is because the aim and intention are different. You may need a more formal approach to communicating with partners, with more focus on goals, especially when working together for strategic reasons. We have learned from experience that this kind of communication needs to be driven by the following:

* Long-term thinking

Your aim may be to build a long-term partnership as a long-term strategic investment. You will need to clearly define roles and responsibilities with a long-term goal. With this type of communication, you need to set the roles and responsibilities of each party from the start. This can be very helpful and may prevent any conflicts in roles. When an overlap occurs, this could become an area where you work together and learn together.&#x20;

* Building trust

It is very important for a successful partnership that the partner organizations trust one another. Building trust is a process. You need to take into account the difference between and uniqueness of each partner. If you trust one another, you can communicate freely. You can exchange ideas, debate, and discuss the progress of the project and the milestones.&#x20;

* Being clear about the different kinds of expertise needed in the project

Each organization brings its expertise, and this is the basis for a partnership. For example, CLEAR Global has many years of expertise in software development, NLP, human-centered design, and partnerships. Another local organization may have expertise in local dynamics and needs. Or it may have a better understanding of the language barriers. By working together, the two organizations complement each other.&#x20;

To handle this kind of communication so that you work together well and get the most benefit, you need a well-guided approach to communication and the right channel. We have also learned from experience that, when communicating with partners, we need the following high-level steps:&#x20;

<br>

1. Setting up multi-level communication channels

This means setting up various possible ways of communicating. This is to meet different needs for information sharing, working together, decision-making, and sharing feedback. We need to understand that different types of information and interaction need different levels of detail, urgency, and privacy considerations. Experience has shown that it is important to identify communication needs, levels, and suitable channels.&#x20;

Setting up proper labels and channels of communication makes it easier to share information and creates better ways of making decisions. For instance, to communicate with partners, you might need to use official channels that allow you to make references but are also faster and more efficient. When communicating with others working on the project, you might require fewer official channels but also fewer details. You may also need an option for urgent communications. You might also want to create communication guidelines. And you may want to ensure consistency and adaptability to give you the flexibility that you may need in a virtual work environment.&#x20;

2. &#x20;Aiding interaction with the user organization

To aid interaction with the user organization, we have learned from experience that it is important to take ownership of the process and final product with a local partner. For this to work well, partners need to be involved in multiple communication channels, with several channels of communication and knowledge sharing.&#x20;

The communication process should be completely user-centered so that user needs are met. You also need to involve all stakeholders in finding a solution, at every stage. We have also learned that local partners need to build user communities for any tech product and shape communication around the product. It is useful to be transparent in sharing feedback and providing communication leadership.

3. &#x20;Sharing best practices and lessons learned

Some of the best communication practices are as follows:

* make the process two-way&#x20;
* be clear to avoid misunderstandings&#x20;
* manage expectations at all times&#x20;
* use multiple channels&#x20;
* set up a feedback loop&#x20;
* be transparent&#x20;
* be timely and
* be aware of cultural factors.

\
These are key practices for smooth communication and to allow you to work well in a partnership. It is also important to share lessons learned in the course of the project. These lessons could be learned from hits and misses, challenges, and opportunities. It is vital to collect these lessons at every milestone, keep a record of them in a document, and share them for future partnerships. <br>


# 6. Language Technology Implementation

Now that you have developed a general understanding of language technology and identified your use case, it's time to start building! This chapter will help you locate where your use-case stands in the language technology landscape and how to navigate towards a working solution. We have developed a comprehensive workflow that will help us get a high-level view and then dive into it bit by bit.&#x20;

{% hint style="info" %}
This chapter assumes technical familiarity with Natural Language Processing tools, data and model development.&#x20;
{% endhint %}

This chapter is organized as follows:&#x20;

* [**Section 6.1**](/6.-language-technology-implementation/6.1-navigating-the-language-technology-landscape) delves into our language technology development workflow to gain a high-level comprehension of the key aspects involved in language technology implementation.&#x20;
* [**Section 6.2**](/6.-language-technology-implementation/6.2-creating-a-language-specific-peculiarities-lsp-document) centers on the necessity of a language-specific analysis for creating language tools and outlines the process of creating a language-specific peculiarities (LSP) document.&#x20;
* [**Section 6.3**](/6.-language-technology-implementation/6.3-open-source-data-and-models) curates a compilation of notable open data resources and provides examples of conducting a data search.&#x20;
* [**Section 6.4**](/6.-language-technology-implementation/6.4-assessing-data-and-model-maturity) introduces methodologies for conducting initial assessments of models and data.&#x20;
* [**Section 6.5**](/6.-language-technology-implementation/6.5-key-metrics-for-evaluating-language-solutions) elaborates on the essential metrics vital for evaluating the efficacy and success of a language solution.


# 6.1 Navigating the Language Technology Landscape

<figure><img src="/files/fOh3sQvMYCGGglnYvoVh" alt=""><figcaption></figcaption></figure>

This flowchart describes the key steps and decision points an initiative faces while developing a solution in a marginalized language using NLP.

The starting point of the diagram assumes that you have already come up with a rough idea involving an NLP solution. For example, it could be an already existing solution for a high-resourced language. As NLP languages are **data-driven**, it is highly likely that there are no well-developed solutions for low-resource languages. The diagram also depicts the decisions to follow in this regard.

{% hint style="info" %}
***Data-driven** means that the intelligence that is created with these tools is collected from large volumes of information, or simply data. For example, in the case of machine translation, the engine “models” translation from one language to another by looking at a collection of human-translated documents and sentences. Similarly, a sentiment analyzer learns how to label if a tweet says good or bad about a company from thousands of tweets labeled by humans as carrying a good or bad sentiment.*

*This dependency on data is what makes these technologies accessible to some languages but not to others. Hence the widely used definitions of high-resource and low-resource languages. The available resources for a language directly influence the possibility of developing an application for that language. As the greatest resource of textual data is the internet, which is dominated by a few languages, these technologies tend to focus on only a handful of dominant languages, e.g., English, Spanish, Chinese, Arabic, etc.*

*The diagram below by*[ *Microsoft Research Labs India*](https://www.microsoft.com/en-us/research/publication/ellora-enabling-low-resource-languages-with-technology/) *illustrates the hierarchy created by this “power law” among languages.*
{% endhint %}

![Classification of languages according to the availability of language technology, tools and resources](/files/z7TwZMf6y67kflaLzP1u)

Let’s now navigate through our landscape diagram step by step

![](/files/WmBgTj8ZPYBlghoQg1dV)

### **What kind of model do I need?**

The first point to identify is what kind of AI task or tasks the solution is using. Some of the most common NLP tasks are as follows:

Text-based:

* **Machine Translation:** Converting text from one language to another while preserving its meaning.
* **Information Retrieval:** Finding relevant documents or information from a collection based on user queries.
* **Information Extraction:** Automatically extracting structured information from unstructured text.
* **Sentiment Analysis:** Determining the emotional tone or sentiment expressed in a piece of text.
* **Question Answering:** Automatically answering questions posed in natural language.
* **Text Summarization:** Generating concise summaries from longer texts while retaining key information.
* **Named-Entity Recognition:** Identifying and classifying named entities (such as names of people, places, and organizations) in text.

Speech-based:

* **Automatic Speech Recognition:** Converting spoken language into written text.
* **Text-to-Speech Synthesis:** Generating natural-sounding speech from written text.

Image-based:

* **Image captioning:** Generating descriptive captions for images.
* **Image classification:** Assigning labels to images based on their content.
* **Image generation:** Creating new images based on certain input criteria.
* **Image sentiment analysis:** Determining the emotional tone or sentiment expressed in images.
* **Visual question answering:** Answering questions about the content of images.
* **Scene recognition:** Identifying the type or context of a scene in an image.

{% hint style="info" %}
Recently, most NLP tasks feed on the use of large language models (LLM). A particular task can be achieved through the use of the correct prompts. So you can either look for a model specific for your task or directly for a large language model that supports your language.
{% endhint %}

When you're designing a solution or product, keep in mind that these kinds of NLP tasks compose only a component of the solution. For example, consider a Customer Support Chatbot solution. In this scenario, a **Question Answering** NLP component is employed to handle customer inquiries. However, building the complete solution involves various additional components. These include data collection and preprocessing, named-entity recognition, intent recognition, response generation, a knowledge base, user interface design, back-end integration, sentiment analysis, an escalation mechanism, and mechanisms for continuous learning and improvement.

### **Where should I look for open-source models?**

There are many open-source models for language technology available online, such as on platforms like GitHub, Hugging Face, and Kaggle. These platforms host repositories containing pre-trained models and resources that cover a wide array of NLP tasks. Developers can access, fine-tune, and adapt these models to their specific needs, significantly expediting the development process and benefiting from the collective knowledge of the open-source community. We’ll explore this in more detail in the [**Section 6.3 Open source data and models**](/6.-language-technology-implementation/6.3-open-source-data-and-models).

### **Which model should I choose?**

It’s very easy to get lost in the vast amount of available models in the wild. The choice of the model depends on your use case, performance requirements, and ethical considerations. You should compare different models based on their capabilities, limitations, and suitability for your problem. You should also consider the trade-offs between the complexity, accuracy, and efficiency of the models. We’ll explore this in more detail in the next section.

Let’s say you found various models that could serve your purpose; the next step should be to check if they serve your particular purpose.

![](/files/S4YguuAS0BqtermmCst1)

### **Does this model do what it promises?**

To verify if a model does what it promises, you can get a hint from the published benchmarks of the model. However, the best practice is to test it on your domain with a representative sample of your data. You should also evaluate the model’s quality, robustness, and fairness using appropriate metrics and benchmarks. You should also check if the model’s license and terms of use are compatible with your intended application. We'll explore this in more detail in [6.4 Assessing data and model maturity](/6.-language-technology-implementation/6.4-assessing-data-and-model-maturity).

### **How do I check if a model serves my purpose?**

To check if a model serves your purpose, you need to define clear and measurable objectives for your solution. You should also collect feedback from your users and stakeholders to assess the impact and value of your solution. You should also monitor and update your model regularly to ensure its reliability and relevance.

Once you have identified a working model for your solution, you should make it available to meet your demand. Making your model accessible and usable by your application is a crucial step. Machine learning models are usually served through what's called an **application programming interface (API)**.

{% hint style="info" %}
**What is an API?**

An Application Programming Interface (API) is a set of rules and protocols that allow different software applications to communicate with each other. It defines the methods and data structures that developers can use to interact with a specific software component, such as a machine learning model. APIs enable applications to request certain tasks or information from another system and receive appropriate responses.

**Serving Machine Learning Models through APIs**

Machine learning models, including those used in language technology, are often served through APIs. These APIs expose the functionality of the model to other applications, allowing developers to integrate the model's capabilities without needing to understand its intricate internal workings.

For instance, if you have a machine translation model, you can create an API that inputs text in one language and returns the translated text in another language. This simplifies the integration process, as developers can interact with the model using standard HTTP requests rather than delving into the complexities of the model architecture.
{% endhint %}

<figure><img src="/files/QC3mcNbqPTGVBjOOy3jK" alt=""><figcaption></figcaption></figure>

### **Can I use this model with a ready-made API?**

Some models and data are supported by ready-made APIs that allow you to access them easily and quickly. For example, you can use Google Cloud APIs for language technology tasks such as translation, speech recognition, natural language understanding, etc. However, not all models and data are available through APIs, and you may need to build your own API or use other methods to access them.

### **How do I deploy and scale?**

To deploy and scale your solution, you need to consider the infrastructure, platform, and tools that you will use. You should also consider the security, privacy, and compliance issues that may arise from your solution. You should also plan for maintenance, updates, and improvements to your solution over time. In [Chapter 7](/7-development-and-deployment-guidelines), we give a detailed overview of deploying an MT backend and a RASA-based chatbot.

In the event that you haven’t found a suitable model for your purpose, you should already be thinking about how to obtain high-quality data for your domain. This data can be used to train a model or *fine-tune* an already existing one.

{% hint style="info" %}
**Fine-tuning** involves taking a foundational model and adjusting it to better suit your specific task or domain. If a model for your exact language or task isn't readily available, you might consider fine-tuning from a foundational model or even a similar language.

For instance, imagine you're working on a translation model for Kinyarwanda, a Bantu language. While there might not be a pre-trained model for Kinyarwanda specifically, you could fine-tune a Swahili-English translation model to help with Kinyarwanda-English translation. Swahili and Kinyarwanda share similar linguistic characteristics as Bantu languages, making the fine-tuned model a valuable starting point.
{% endhint %}

Ensuring you have enough high-quality data for training your model is crucial. The availability of data might vary based on your language and specific task. Sometimes, pre-existing datasets might be inadequate, leading to the need for data augmentation or active collection efforts. We will look at various open data resources in [Section 6.3](/6.-language-technology-implementation/6.3-open-source-data-and-models).

<figure><img src="/files/eIMiqBGghdApZqYztcXi" alt=""><figcaption></figcaption></figure>

### **Where should I look for open data?**

There are many sources of open data for language technology, such as HuggingFace, Kaggle, Common Crawl, Common Voice. You can also find some examples of prominent open data resources in [Section 6.3](/6.-language-technology-implementation/6.3-open-source-data-and-models) of this chapter.

### **What type of data do I need?**

The type of data you need depends on your use case and the task at hand. You should look for data that is relevant, representative, reliable, and diverse for your problem. You should also consider the size, format, quality, and licensing of the data.

Once you have collected enough good data, you can start experimenting with training your model.

### **What resources do I need to build a model?**

To build a model, you need to have the skills, tools, and time to perform tasks such as data collection, preprocessing, modeling, training, testing, evaluation, optimization, etc. You may also need to collaborate with other experts or stakeholders to ensure the quality and usability of your model. You should also consider the cost and availability of computational resources such as [GPUs or TPUs](/2.-overview-of-language-technology/2.4-key-terminology-and-concepts) that may be required for building a model.

In the event that you can’t find the right dataset or it’s not good enough to obtain a high-quality dataset, you have the option to take matters into your own hands and collect your own dataset.

Even in cases where you have access to sufficient data, it’s always good practice to design your project in such a way that it accumulates data over time. This ensures the improvement of your system over time.

![](/files/1TRfCe2OXshOsN8foqg1)

### **What open data initiatives are there?**

Open data initiatives like Common Voice and Tatoeba provide valuable platforms where individuals voluntarily contribute by recording and translating sentences in various languages. These initiatives tap into the collective efforts of people around the world, resulting in diverse and substantial datasets that can be used for training and fine-tuning models. By harnessing the power of community-driven data collection, you can ensure the availability of relevant and authentic language resources tailored to your specific needs. We’ll discuss this deeper in [**6.3 Open source data and models**](/6.-language-technology-implementation/6.3-open-source-data-and-models).

### **How much data do I need to collect?**

The amount of data you gather depends on how complex your problem is and how much data you’re building on. A rule of thumb is to collect a dataset as representative as possible of your real world application.

Here's a simple way to think about it: Imagine you're creating a system that understands spoken words, like the kind of system that listens to what people say on the phone. If you want it to work well in the real world, you need data that's similar to the real phone conversations people will have. This data should reflect the different kinds of people who might use your system, like different ages, genders, dialects and levels of education. That way, your system can understand and respond to everyone.

Think about the words people will use, too. If your system is going to be used in a specific area, like banking, it's important to have recordings of conversations about banking. The words, phrases, and terms used in banking will be different from, say, a system meant for ordering food.

### **How do you mobilize a language data community?**

Mobilizing the community means actively engaging a group of people who share a common interest or expertise, often related to language or data. For language technology, it involves reaching out to individuals who can contribute valuable data or insights. This can include native speakers, domain experts, or volunteers. Strategies involve organizing workshops, online challenges, or forums where contributors can share recordings, translations, or annotations. Mobilizing the community helps amass diverse, high-quality data that improves the accuracy and effectiveness of language technology solutions.

Sometimes you have potential data sources that can be processed to obtain the necessary data. For instance, if your goal is to build a translation engine specialized in news articles, you can scrape news articles from websites and translate them into your target language, creating parallel data (refer to Masakhane’s work on [news topic classification](https://arxiv.org/abs/2304.09972) and [news translation](https://aclanthology.org/2022.naacl-main.223/)). Similarly, if you're developing a telephony application for banking, you can gather banking-related text to fine-tune your language model. In these cases, you would need to collect, clean, and restructure the data to prepare it for input into your model.

![](/files/wYqmxxfQKWLwet53eQzu)

### **Where should I look when scraping data?**

When scraping data from websites, adhere to ethical guidelines and the website's terms of use. Prioritize sources that provide structured and relevant data while respecting privacy and legality.

### **What should I care about in terms of licensing?**

Licensing is crucial when using external data or models. Ensure the data and models you use are legally accessible and that you comply with licensing terms. Additionally, consider licensing your own work to enable future sharing and collaboration. We’ll discuss further on open data licensing in [**6.3 Open source data and models**](/6.-language-technology-implementation/6.3-open-source-data-and-models).


# 6.2 Creating a Language-Specific Peculiarities (LSP) Document

Before embarking on the creation of language tools, it is highly recommended to consider the development of a **Language-Specific Peculiarities (LSP) document**. An LSP document is designed to serve as a comprehensive reference guide that outlines the unique linguistic features of a given language. This document is particularly valuable for projects involving Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) systems, providing essential insights into pronunciation, orthography, and other language-specific nuances.

An LSP document offers a structured overview of various linguistic aspects crucial for phonemic and orthographic transcription. It covers elements like phonemic variation, syllabic structure, suprasegmentals such as stress and intonation, and foreign language influences that shape the language's phonological characteristics.

#### **Orthography and Input Conventions**

Understanding the orthographic conventions of a language is essential for accurate transcription. An LSP document delves into orthography by detailing character sets, romanization schemes, punctuation rules, and spelling variations. It also addresses text normalization for elements like numerals, percentages, dates, addresses, acronyms, and abbreviations, ensuring consistency in transcription.

#### **Other Language Items and Function Words**

Beyond phonetic and orthographic considerations, an LSP document delves into other linguistic components. This includes comprehensive tables for digits, alphabet pronunciations, month names, weekday names, and common date and time expressions. Additionally, function word lists for pronouns, articles, conjunctions, prepositions, and filler words enrich the document's linguistic insights.

#### **Creating an LSP Document**

Creating a robust LSP document involves linguistic expertise. Collaborating with linguists who possess a deep understanding of the language's phonetic, phonological, and orthographic characteristics is vital. The linguist's role is to meticulously analyze the language's features and formulate concise guidelines. This can be achieved through a thorough study of linguistic resources, phonetic and orthographic patterns, and interactions with native speakers. The LSP document serves as a reference for developers and researchers, aiding in the accurate development of language technology applications that align closely with the language's natural nuances.

In essence, an LSP document acts as a linguistic compass, guiding the development of language technology solutions that reflect the authenticity and intricacies of the language. Through collaboration with linguists and rigorous study of phonological and orthographic patterns, this document becomes an invaluable tool, enhancing the accuracy and effectiveness of ASR and TTS projects.

In 2022, Appen collaborated with CLEAR Global to create an LSP document for Sheng, a Swahili and English-based language used by the youth in Nairobi, Kenya. You can find this resource together with a template LSP document to guide your LSP creation process through this page.

{% embed url="<https://gamayun.translatorswb.org/data/>" %}


# 6.3 Open source data and models

In this section, we delve into the world of open data sets and models, pivotal resources that have catalyzed the advancement of language technology. Open data sets refer to collections of data that are freely available to the public, fostering collaboration and innovation. We suggest taking an open data strategy when building language technology solutions. This ensures that a wider community can participate in and take advantage of the data you create.

We will now explore where to find such data, important considerations like licensing, and the significance of these resources in driving progress.

#### **Discovering the NLP landscape with HuggingFace**

A notable platform in the realm of open data sets and models is Hugging Face. It offers an extensive collection of pre-trained models and datasets. This platform not only facilitates model search but also offers interactive demonstrations of these models, allowing you to witness their capabilities firsthand. Hugging Face is a hub for initiatives by individuals, organizations, and communities, offering a wealth of language technology resources.

#### **OPUS**

When in search of targeted data sets for specific tasks, platforms like OPUS emerge as invaluable resources, particularly for machine translation data. The OPUS project, an initiative from Helsinki University NLP group, serves as a veritable goldmine of translated texts sourced from across the web. This initiative aims to transform and align freely available online content, augmenting it with linguistic annotations. The outcome is a meticulously curated parallel corpus that's publicly accessible. The OPUS project operates on the principles of open source, offering the corpus as an open content package. The compilation of this corpus relies on a variety of open-source tools, and the entire pre-processing procedure is automated to serve in formats that are readily usable with machine translation libraries.

While both Hugging Face and OPUS fall within the realm of open data, they diverge in their scope and offerings. OPUS stands as a specialized resource focused on parallel data, making it a potent asset for machine translation tasks. This platform operates in a moderated environment, wherein its content is carefully curated, ensuring the quality and relevance of the provided data. Unlike Hugging Face, OPUS is designed to primarily host parallel corpora, and it restricts the ability for users to upload their own datasets. This tailored approach to parallel data signifies OPUS's commitment to facilitating high-quality language resources for the language technology community.

#### **Speech data**

Within the realm of language technology, speech data forms the cornerstone for fueling technologies like automatic speech recognition (ASR) and text-to-speech (TTS). ASR enables machines to transcribe spoken language into text, while TTS transforms text into natural-sounding speech. To drive the development of these technologies, open speech datasets are indispensable. Among the valuable resources is [OpenSLR](https://www.openslr.org/resources.php), a repository of curated open-source speech and audio datasets. This platform provides datasets that serve as the bedrock for training and evaluating some of the most important foundational ASR and TTS models.

#### **Common Voice: A Paradigm Shift for Inclusive Speech Data**

In the landscape of speech data, the Common Voice initiative stands out as a transformative force, reshaping the trajectory of voice-enabled technology. Common Voice hosts a publicly accessible voice dataset that owes its existence to the collaborative contributions of volunteers worldwide. Rather than being owned by corporations, this platform thrives on community-driven efforts, offering an extensive dataset for training machine learning models used in ASR and TTS applications.

Engaging with the Common Voice movement is straightforward yet impactful. Participants can register to create accounts, enabling them to actively contribute to expanding the dataset while keeping track of their progress. The process involves crafting sentences, recording them, and validating the contributions of others. This collective approach not only enriches the dataset but also empowers contributors with a sense of ownership in driving language technology forward. The datasets are downloadable with CC-0 (Public domain) licenses through their portal.

An illustrative triumph of Common Voice lies in languages once deemed low-resource. Languages like Kinyarwanda and Catalan, historically lacking comprehensive resources, have experienced a resurgence driven by community involvement. Through local campaigns (see Digital Umuganda (LINK) and Aina (<https://www.projecteaina.cat/>), these languages have flourished with diverse contributions and taken their place at the top in terms of contributed hours.

As of 11 August 2023, Common Voice supports data contribution in 123 languages, with many more under development. If your language is not on this list yet, you can ask for it to be included. The list of tasks to complete in order for a language to start contributing data is:

* Localization of the site through Mozilla’s platform Pontoon
* Collection of sentences with CC-0 license (Public domain). More information on CC-0 [here.](https://common-voice.github.io/community-playbook/sub_pages/text.html)

#### **Licensing and Sharing of data and models**

Licensing stands as a cornerstone in the realm of open data and models, dictating the rules of engagement for usage and sharing. Diverse licenses exist, each carrying its own set of permissions and restrictions. Understanding these licenses is paramount to responsibly leveraging open resources in your language technology projects. In the open data space, here are some commonly used [Creative Commons (CC) licenses](https://creativecommons.org/about/cclicenses/) that you might encounter: **CC BY (Attribution)**: This license permits you to use, modify, and share the data or model, even for commercial purposes, as long as you provide appropriate attribution to the original creator.

1. **CC BY-SA (Attribution-ShareAlike)**: Similar to CC BY, this license allows usage, modification, and sharing, but any derivative works must be licensed under the same terms. This ensures that subsequent work remains open as well.
2. **CC BY-NC (Attribution-NonCommercial)**: This license permits usage, modification, and sharing for non-commercial purposes while requiring attribution to the original creator.
3. **CC BY-ND (Attribution-NoDerivatives)**: Under this license, you can use and share the work, but modifications or derivatives are not allowed. Proper attribution is necessary.
4. **CC BY-NC-SA (Attribution-NonCommercial-ShareAlike)**: Like CC BY-SA, this license allows modification and sharing, but only for non-commercial purposes. Any derivative works must be licensed under the same terms.
5. **CC BY-NC-ND (Attribution-NonCommercial-NoDerivatives)**: This is the most restrictive CC license. You can download and share the work, but you can’t modify it or use it commercially.

Adhering to the terms of these licenses ensures ethical and legal use of the resources you encounter. While utilizing open data and models is vital, the spirit of collaboration is equally significant. Contributing back to the community by sharing your own fine-tuned models and datasets, even if they aren't perfect, reinforces the communal nature of language technology advancement. This practice enables others to build upon your work, collaboratively enhancing the capabilities of language models, fostering innovation, and nurturing a dynamic ecosystem of collective growth.


# 6.4 Assessing data and model maturity

Before you start creating or using a language technology solution, it's important to check how good the NLP components you're using are. This includes the pre-trained models or the data you use to train them. This initial check helps ensure that your solution is strong and works well in real situations. Whether you're working on machine translation (MT), automatic speech recognition (ASR), or using large language models (LLMs) for natural language generation, it's crucial to evaluate how good your data and models are before getting started.


# 6.4.1 Assessing NLP Data Maturity

Data maturity refers to the quality, quantity, and relevance of the training data used to build an NLP model. The principle of "garbage in, garbage out" holds particularly true in NLP. This concept emphasizes that the quality of your model's output is directly tied to the quality of the data you use for training. While machine learning models can tolerate errors in training data, the overall quality of the data significantly impacts the resulting model's performance.

Errors in training data can manifest as mislabelings, inaccuracies, or misrepresentations of language. Even subtle issues like biases within the data can have profound effects on the model's behavior. It's crucial to recognize that, especially with large datasets, assuming that the data is clean and unbiased is a risky assumption. Ensuring data quality is essential to building accurate and reliable NLP models.

To elaborate on the "garbage in, garbage out" theory, think of it as using a recipe to cook a meal. If you start with poor-quality ingredients or misinterpret the recipe, the final dish won't turn out as expected. Similarly, when training an NLP model, if the training data is of low quality or contains biases, the model's predictions and responses will likely be inaccurate or biased as well. This highlights the importance of thoroughly evaluating your training data to identify and rectify any issues before building your NLP solution.

The three main considerations in assessing data maturity are: **data quality, data quantity and data relevance.**

**Data quality** refers to the overall reliability, accuracy, completeness, and suitability of the data used for a specific purpose.

To ensure data quality, it's essential to delve into the dataset's details, which are often provided in the form of a dataset card or a research paper outlining the collection process. In the case of datasets resulting from scientific research, they typically come with thorough evaluations. Prior to commencing development, it's wise to conduct an audit of the dataset, especially for applications involving low-resource languages that may only be present in large multilingual datasets. Adopting a cautious approach to data quality is crucial in such scenarios.

In addition to relying on documentation, directly examining the data itself is a recommended practice. For a more structured and systematic review, a valuable approach can be found in the paper titled "[Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets](https://arxiv.org/abs/2103.12028)." In this paper, the authors outline a method where they randomly select 100 sentences from the dataset and categorize them based on attributes such as correctness, offensiveness, correct/incorrect language usage, and quality of translation. This approach provides a practical means to assess and validate the quality of the dataset in a more comprehensive manner.

{% hint style="success" %}
**Metric 1 - Sample Correctness Ratio (SCR) for measuring data quality**

The "Sample Correctness Ratio (SCR)" metric can be used to evaluate the quality of labeled datasets in NLP. It calculates the percentage of labeled samples that adhere to predefined correctness criteria. To calculate this metric, randomly select a representative sample size, such as 100 or 1,000 samples, depending on your capacity, from the dataset. Then, determine the number of correctly labeled samples depending on a particular criteria (correct translation, label, etc.) within that sample and divide it by the total sample size, multiplying by 100.

$$C = \frac{\text{Number of Correctly Labeled Samples}}{\text{Total Labeled Samples}} \times 100%$$
{% endhint %}

**Data quantity** refers to the size of the dataset. One cannot expect to train robust speech recognition models with only a few hours of data; however, it might be possible to fine-tune a model for specific acoustic conditions. In the case of machine translation, creating a foundational model from scratch typically requires millions or even billions of parallel sentences. Nevertheless, if an existing model is available and the goal is to specialize it for a particular domain, a few thousand high-quality translations can suffice. Unfortunately, there are no definitive rules for determining exactly how much data is needed for a specific purpose. It's a judgment that an NLP expert can make by consulting the state-of-the-art and through their experience, but can only confirm through experimentation.

**Data relevance** ensures that the training data reflects the scenarios and contexts the model will encounter. Irrelevant or outdated data might result in a model that struggles to handle real-world inputs. In essence, your training data should mirror the environments in which the model will be employed, providing it with the necessary exposure to tackle real-world complexities. For example, when you're constructing a machine translation system intended to cater to the health domain, it's imperative that your training dataset be a true reflection of the language nuances, vocabulary, and terminologies specific to that domain. For instance, if your health-related translation system lacks exposure to medical jargon, it might misinterpret or fail to properly translate complex terms. Or let's say you're developing an ASR solution to transcribe customer service calls for a financial institution. Your ASR system could falter if your training data predominantly contains clear studio recordings and doesn't incorporate the variety of background noises, accents, and speech patterns characteristic of real telephone interactions.

{% hint style="success" %}
**Metric 2 -** **Burstiness score for measuring** **the domain-specificity of a corpus**

Santini et al. explore various metrics to measure the domain-specificity level of a corpus in their paper “[Can We Quantify Domainhood? Exploring Measures to Assess Domain-Specificity in Web Corpora](https://link.springer.com/chapter/10.1007/978-3-319-99133-7_17)”. They conclude that burstiness is the most suitable measure to single out domain-specific words from a specialized corpus and to allow for the quantification of the domainhood of a corpus.

**Calculating Burstiness Score:**

1. **Corpus and Term List:** Start with a corpus of text and a list of terms that are representative of your domain.
2. **Count Term Occurrences:** For each term in your list, count how many times it appears in the entire corpus. Let's call this value E.
3. **Divide the Corpus into Subsets:** Divide your corpus into different subsets of time periods or documents. These subsets could represent different sections of your corpus or specific documents.
4. **Count Term Occurrences in Subsets:** For each term in your list and for each subset, count how many times the term appears in that subset. Let's call this value E\_t.
5. **Calculate Total Time Periods:** Determine the total number of time periods or documents in the entire corpus. Let's call this value T.
6. **Calculate Burstiness Score:** For each term in your list and for each subset, calculate the burstiness score using the formula provided in the Wikipedia definition: $$Burst(e, t) = (E\_t / E) - (1 / T)$$\
   Where:
   * E\_t is the total number of occurrences of the term in the subset t.
   * E is the total number of occurrences of the term in the entire corpus.
   * T is the total number of time periods or documents in the corpus.
7. **Interpret Burstiness Score:** A positive burstiness score indicates that the term is occurring more often in the subset T compared to its occurrences in the entire corpus, suggesting burstiness. A negative score implies the opposite.
8. **Repeat for All Terms and Subsets:** Repeat steps 4 to 7 for each term in your list and for each subset of the corpus.

By following these steps, you'll be able to calculate the burstiness score for each term in your given list across different subsets of the corpus. This will help you identify terms that are important in specific documents or time periods but are unevenly distributed across the entire corpus.
{% endhint %}


# 6.4.2 Assessing NLP Model Maturity:

Model maturity encompasses the model's proficiency in executing specific tasks. It is advisable to do a careful study to assess if the model does what it promises and proves to be useful in your solution.

Prior to embarking on model construction, it is advisable to study baseline performance metrics through pre-existing solutions or well-established benchmarks. Establishing these baselines helps manage expectations concerning the model's capabilities, allowing for a clearer assessment of its improvements.

#### **Performance metrics**

Selecting the appropriate evaluation metrics is crucial to gauging the model's competence both in comparison with other models and also standalone. Some traditionally applied metrics in MT such as [BLEU](https://en.wikipedia.org/wiki/BLEU), chrF++ and translation error rate (TER) employ a lexical similarity measure to compare MT output to reference translations but are known to correlate poorly with human ratings.  Similarly, automatic speech recognition (ASR) utilizes the [word error rate (WER)](https://en.wikipedia.org/wiki/Word_error_rate). [The WMT study published in 2022](https://aclanthology.org/2022.wmt-1.2.pdf) shows that trained metrics based on large language models close this gap by offering a much better and robust measure.

{% hint style="success" %}
**Metric 3 - Performance metrics for pre-assessing model maturity**

The principally crucial aspect of evaluating model maturity is the creation of a representative test set. This set of data samples, reflective of the real-world application you are building for, serves as the benchmark against which your model's performance is measured. Ensuring that the test set accurately encapsulates the linguistic and contextual diversity of your use case is fundamental for obtaining meaningful evaluation results. When using automatic metrics, remember that they provide a quantitative assessment of model performance. However, these metrics may not capture the full linguistic and contextual nuances of language or your particular application’s needs. It’s advisable to complement automatic metrics with qualitative evaluations and human judgments for a comprehensive understanding of your model’s capabilities.

**COMET metric for machine translation**

[The COMET metric](https://github.com/Unbabel/COMET) is a versatile neural MT evaluation metric that assesses the quality of machine-generated translations across various languages. It has gained recognition as a valuable benchmark metric for machine translation. COMET scores typically range from 0 to 1, with higher scores indicating better translation quality. As other automatic evaluation metrics, the raw score itself only gives a guideline in interpreting a model’s quality but it is useful for ranking different MT systems.&#x20;

COMET supports 102 languages out-of-the-box and evaluating in other language pairs other than these would result in unreliable conclusions. For languages out of this list, it’s more recommendable to use a lexical-based metric like chrF.

**chrF metric for machine translation**

[chrF](https://aclanthology.org/W15-3049.pdf) is a lexical similarity based metric that uses character n-grams instead of word n-grams (like BLEU) to compare the MT output with the reference. It is known to correlate better with human evaluation especially in non-latin and high-morphology languages.&#x20;

To apply lexical-based metrics like chrF++, BLEU and TER, you can either use Python library [SacreBLEU](https://github.com/mjpost/sacrebleu) or the web-based platform [MutNMT](https://ntradumatica.uab.cat/evaluate/).&#x20;

**Human evaluation metrics for machine translation**

Human evaluation metrics for machine translation encompass various methods to assess and compare machine translation system performance. Challenges in human evaluation include subjectivity, time, cost, and the presence of multiple standards. These metrics include [Multidimensional Quality Metrics (MQM)](https://themqm.org/), Scalar Quality Metric (SQM), TrueSkill for ranking, Adequacy and Fluency judgment, Relative ranking, Constituent ranking, Yes or No Constituent judgment, and Direct assessment. Each method serves distinct purposes, such as identifying translation errors (MQM), providing segment-level ratings (SQM), ranking systems (TrueSkill), judging adequacy and fluency (Adequacy and Fluency judgment), relative system ranking (Relative ranking), evaluating syntactic constituents (Constituent ranking), assessing acceptability (Yes or No Constituent judgment), and direct rating (Direct assessment) in monolingual, bilingual, or reference-based contexts. Different evaluation methods are suited to different purposes, and selecting the appropriate method is crucial for obtaining meaningful and relevant results.

**Word Error Rate (WER) for Automatic Speech Recognition**

The word error rate (WER) metric assesses the accuracy of the transcribed text compared to the reference text. Understanding WER scores helps evaluate the ASR model’s performance:

* **0-10%:** Exceptional performance, indicating highly accurate transcriptions.
* **10-20%:** Good performance, with low errors, requiring light post-editing.
* **20-30%:** Moderate errors, requiring a high level of post-editing for accurate results.
* **30%+:** Substantial errors, demanding significant post-editing efforts for comprehensible transcriptions.
  {% endhint %}

#### **Downstream Testing: Evaluating Model Maturity in Real-World Contexts**

Assessing the maturity of NLP models goes beyond traditional evaluation metrics like COMET, chrF and WER. Downstream testing plays a crucial role in understanding how well a model performs in actual use cases, providing insights that automated measurements might overlook.

One effective approach in downstream testing is to evaluate the model's performance within the context of the specific tasks it's designed to support. For instance, consider a scenario where Machine Translation (MT) is integrated into customer interactions. While automated metrics might indicate mediocre performance, actual improvements in customer satisfaction within the integrated system could demonstrate the model's value.

Conversely, a model might achieve impressive metrics but fail to address the nuances of the task, making it less useful in real-world applications. As a result, project managers are advised to prioritize the practical impact of the model over purely relying on automated evaluations.

In conclusion, downstream testing involves conducting human evaluations within the context of the intended tasks to determine the model's true maturity. By focusing on the model's real-world utility and its ability to drive meaningful outcomes, this approach offers a more comprehensive understanding of its effectiveness and readiness for deployment.

In conclusion, assessing data and model maturity is a fundamental step in developing effective NLP solutions. It involves evaluating the quality, quantity, and relevance of training data, as well as understanding the model’s capabilities through performance metrics and downstream testing. By conducting thorough assessments, developers can make informed decisions about the suitability of the data and the model for their intended application, leading to more robust and successful NLP solutions.


# 6.5 Key Metrics for Evaluating Language Solutions

When assessing the impact, quality, and user satisfaction of language solutions, a well-structured approach to data collection and analysis is paramount. Designing the solution with a data collection perspective from the beginning enables quantitative analysis, real-time monitoring through dashboards, and the ability to iterate for improvements. Consider the following metrics to comprehensively evaluate your language solution:

1. **User Engagement and Satisfaction:**
   * **Unique Users:** Measure the number of distinct individuals who have interacted with the solution.
   * **Repeat Users:** Evaluate the proportion of users who engage with the solution multiple times.
   * **Conversations:** Count the total number of interactions or conversations initiated by users.
   * **Interactions per User:** Assess the average number of interactions per user, indicating the depth of engagement.
2. **Impact Measurement and User Knowledge:**
   * **Quiz or Assessment:** Integrate quizzes or assessments to gauge users’ knowledge before and after using the solution.
   * **Pilot Testing:** Conduct controlled pilot tests with specific user groups before an official release to assess impact and effectiveness.
3. **User Behavior Insights and Localization:**
   * **User Growth:** Monitor the increase in the number of users over time.
   * **User Preferences:** Understand the preferred channels (SMS, WhatsApp, etc.) and interaction times.
   * **Geographical Distribution:** Analyze where your users are located geographically.
4. **Quality of Output (e.g., Chatbots, MT):**
   * **Accuracy of Responses:** Calculate the percentage of questions accurately answered by the chatbot or machine translation.
   * **Quality Ratings:** Allow users to rate the quality of responses on a scale of 1 to 5, indicating accuracy and naturalness.
   * **Human Evaluation:** Benchmark machine-generated outputs using human-labeled data.
5. **Content Availability, Relevance, and Adaptability:**
   * **Topics with Insufficient Content:** Measure the percentage of topics or queries that lack sufficient content.
   * **Content Tracking:** Continuously monitor and update content to address gaps and improve relevance.
   * **Off-the-Scope Topics:** Analyze user-initiated topics that fall outside the predefined scope and assess whether your solution can adapt to address these topics.
6. **User Feedback and Satisfaction:**
   * **User Ratings:** Determine the percentage of users who rate the solution as “helpful” or provide positive feedback.
   * **Survey Responses:** Gather user feedback through surveys to understand satisfaction levels.
7. **User Demographics and Knowledge Enhancement:**
   * **Demographic Information:** Collect data on user characteristics like gender, age, and location.
   * **User Knowledge:** Evaluate whether users gain knowledge or understanding after interacting with the solution.
8. **User Behavior Insights and Engagement:**
   * **Conversation Duration:** Measure the average time users spend in conversations.
   * **Language Insights:** Analyze conversations across different languages for insights.

When equipped with a comprehensive set of metrics, you’ll have the tools needed to quantitatively assess your language solution’s performance, impact, and user satisfaction. This solution-agnostic approach ensures effective evaluation of a wide range of language solutions and empowers data-driven decision-making to continuously improve and enhance their effectiveness. Subsequent chapters will delve into specific evaluation methods tailored to chatbots and machine translation solutions.


# 7 Development and Deployment Guidelines

Once you have developed your datasets and models, it’s time to deploy them into your solution. This chapter is dedicated for your technical team members to learn ways to develop and deploy language technology solutions. We’ll start by giving a general introduction in model deployment and then dive deeper into chatbot and machine translation development.

{% hint style="info" %}
This chapter assumes technical knowledge on machine learning concepts and server deployment.&#x20;
{% endhint %}


# 7.1 Serving models through an API

Deploying open-source machine translation models involves the implementation of an **application programming interface (API)** to make your model accessible to users and applications. An API acts as a bridge between your application and the machine translation model, allowing seamless communication and integration. For example, in the context of deploying machine translation models, an API allows you to send a source text to the model and receive the translated output in return. It provides a standardized way for applications to communicate with the model.

Traditional server deployment involves setting up a dedicated server to host the model and serve model requests. You'll need to configure the server environment, manage resources, and handle scaling as the number of requests increases. Platforms like Flask, Django, or FastAPI can be used to create the API endpoints and handle the translation process. In [7.2.2 Deploying your own scalable Machine Translation API](/7-development-and-deployment-guidelines/7.2-machine-translation/7.2.2-deploying-your-own-scalable-machine-translation-api), we will present an example of a machine translation API built on the FastAPI framework.

Even though training machine learning models is generally feasible only with a GPU at hand, it is possible to serve ML models without one. Even most large scale models can be loaded, given sufficient memory. Although, having a GPU at hand would decrease inference time substantially. On the one hand, depending on the real-time constraints of your application, this is something to take into consideration. On the other hand, It's important to note that GPU servers come at a higher cost due to their enhanced processing capabilities. However, even in the absence of dedicated GPUs, the deployment of ML models remains viable, making it an accessible option for various scenarios.

### Frameworks for serving models

For streamlined deployment of PyTorch models in production environments, [**TorchServe**](https://pytorch.org/serve/\)) provides a comprehensive solution. TorchServe facilitates efficient model serving with minimal latency, making it suitable for high-performance inference. The platform offers default handlers for common applications such as object detection and text classification, eliminating the need for custom code deployment. With advanced features like multi-model serving, model versioning for A/B testing, monitoring metrics, and RESTful endpoints, TorchServe seamlessly transitions models from research to production. It supports various machine learning environments, including Amazon SageMaker, Kubernetes, Amazon EKS, and Amazon EC2. Detailed information and documentation for using TorchServe can be found in its [official documentation](https://github.com/pytorch/serve/blob/master/docs/README.md).

**Tensorflow Extended (TFX)**, on the other hand, is a comprehensive platform designed for deploying end-to-end production ML pipelines of Tensorflow-based models. As you're ready to move your models from research to production, TFX aids in creating and managing a production pipeline. A TFX pipeline consists of components that implement an ML pipeline tailored for scalable, high-performance machine learning tasks. These components can be built using TFX libraries, which can also be employed individually. More information about TFX and its capabilities can be accessed through its [official documentation](https://www.tensorflow.org/tfx/tutorials).

[**HuggingFace Inference Endpoints**](https://huggingface.co/docs/inference-endpoints/index) provides a secure and convenient solution for deploying Hugging Face Transformers, Sentence-Transformers, and Diffusion models from the Hub onto dedicated and auto-scaling infrastructure managed by Hugging Face. This service enables you to deploy models without the need to rent and administer a server, as you would be paying Hugging Face for their managed infrastructure. The creation of a Hugging Face Endpoint is based on a Hugging Face Model Repository, which generates image artifacts from the chosen model or a custom-provided container image. These image artifacts are detached from the Hugging Face Hub source repositories to ensure heightened security and reliability.

Hugging Face Inference Endpoints supports all tasks associated with Hugging Face Transformers, Sentence-Transformers, and Diffusion, as well as custom tasks not currently supported by Hugging Face Transformers, such as speaker diarization and diffusion. Additionally, the service allows for the use of a custom container image managed externally via services like [Docker Hub](https://hub.docker.com/), [AWS ](https://aws.amazon.com/ecr/)Elastic Container Registry, [Azure Container Registry](https://azure.microsoft.com/en-us/products/container-registry), or Google [Artifact Registry](https://cloud.google.com/artifact-registry). This deployment approach proves particularly advantageous if you seek to sidestep the complexities of managing your own server while leveraging the capabilities of state-of-the-art language models for your applications.

### **Serverless Deployment**

Serverless deployment offers a more hands-off approach, allowing you to focus on the code rather than infrastructure management. Platforms like [AWS Lambda](https://aws.amazon.com/lambda/), [Google Cloud Functions](https://cloud.google.com/functions?hl=en), and [Azure Functions](https://azure.microsoft.com/en-us/products/functions) enable you to deploy your API without provisioning servers. These platforms automatically scale based on demand, which can be particularly advantageous for handling varying workloads.

The potential benefits of serverless computing become clear when considering the 'classic' workflow for setting up a new server in the cloud. With traditional methods, creating a new virtual machine (VM), configuring the machine, setting up NGINX, and managing auto-scaling rules consume significant time and energy. Additionally, you're billed for every second of server uptime, regardless of usage patterns. In contrast, serverless deployment platforms alleviate these concerns, allowing you to focus on writing and deploying your application's core code. This approach is particularly cost-effective for apps with variable usage patterns, as you only pay for the resources you actually use.

As you consider serverless deployment, it's crucial to be aware of certain drawbacks. When serverless functions are inactive, there can be a brief lag when they're reactivated, affecting responsiveness. Limited resources and configurability may hinder tasks requiring extensive CPU or memory usage. Debugging can be challenging due to less control over the runtime environment, especially when dealing with complex architectures. Moreover, vendor lock-in is a concern as cloud providers often offer specific frameworks for serverless deployment, potentially complicating migration to different platforms.


# 7.2 Machine translation

Machine Translation (MT) is defined as the automatic conversion of text in one language to another language. It has evolved through the years from rule-based to statistical approaches, which modeled the probabilities of mappings between sub-phrases between translations. These probabilities are learned in a statistical fashion from parallel texts where sentence-aligned translations are available in the languages involved (referred to as source and target languages). The diagram below illustrates the modeling of translating the word “sure” from English to Spanish using translations made in the UN Parliament.

![Extracting statistics of translations from parallel data](/files/hrpUUXQJT8ryYe2yVOEd)

In recent years, the field of Machine Translation (MT) has undergone a significant paradigm shift, marked by the emergence of Neural Machine Translation (NMT) systems. Unlike SMT, NMT does not explicitly model these sentence-level probabilities. Instead, NMT models estimate probabilities of generating each target token (word or subword) given the source sentence and previously generated tokens. These probabilities, denoted as p(y\_i|x, y\_\<i), are learned during training from parallel corpora where source and target sentences are aligned. Notably, NMT models store complex mathematical functions' parameters rather than explicit probabilities, transforming input sentences (sequences of discrete symbols) into continuous embeddings, which are then used for mathematical manipulation. This transformation enables NMT to work with the continuous vector representations of language, a departure from the discrete symbol-based approaches of the past.

<figure><img src="/files/bXULn1J8p3TnkiOfqW5O" alt=""><figcaption><p>A simplistic representation of encoder-decoder architecture</p></figcaption></figure>

The figure above roughly demonstrates an encoder-decoder type NMT architecture, a typically common structure in this family of MT systems [introduced in 2014](https://arxiv.org/abs/1409.0473). The encoder is responsible for "reading" the source sentence and creating continuous embeddings, while the decoder generates target tokens. This architecture served as a precursor to the development of the [Transformers architecture](https://research.google/pubs/pub46201/), which is now widely used both within and outside Machine Translation (MT). Transformers introduced self-attention mechanisms that enable models to weigh the relevance of different words in the source and target languages, significantly improving translation quality by capturing long-range dependencies and context in an unparalleled way.

This paradigm shift represented a breakthrough in MT, offering more fluent and context-aware translations by leveraging neural networks to handle the entire translation process, transcending the limitations of earlier rule-based and statistical methods. This new way of modeling introduced in 2014 made 50% fewer word order mistakes, 17% fewer lexical mistakes, and 19% fewer grammar mistakes compared to earlier models as [a 2018 study shows](https://arxiv.org/ftp/arxiv/papers/1803/1803.08409.pdf).

Machine translation services like Google Translate and DeepL have made their way into reliable tools for translators and also regular folk in the recent years. The principle uses of machine translation are as follows:

1. **Assimilation**, emulating a certain document in another language. This use-case enables e.g. reading a news site or technical paper in a language that we don’t understand. We know that it’s not a 100% accurate translation, but it gives the gist to explore further.
2. **Communication**, enabling the communication between individuals and organizations e.g. in chat, tourism, and e-commerce, lowering the need for a lingua franca.
3. **Monitoring**, enabling tracing of information in large-scale multilingual documents e.g. discovering international trends in Twitter.
4. **Assistance**, in improving translation workflows e.g. computer-assisted translation, and post-editing.

### Parallel data (bitext)

The type of data that is needed to build a machine translation system is **parallel data**, which consists of a collection of sentences in a language together with their translations. Historically, parallel data were sourced from translations in multilingual public spaces like the United Nations, and European Parliament. Now, the greatest resource of parallel text is the multilingual web.

#### Sourcing parallel data

As we presented in [Section 6.3](/6.-language-technology-implementation/6.3-open-source-data-and-models), [OPUS](https://opus.nlpl.eu/) is a collection of almost all publicly available parallel data. It is the go-to point for many researchers to publish their parallel data or source data for the development of MT models. In addition to OPUS, the Common Crawl initiative plays a pivotal role in the creation of modern MT systems. Common Crawl is a vast web archive that captures a snapshot of the internet, including multilingual content from diverse sources. Researchers and developers can tap into this rich resource to extract parallel text from websites, news articles, and other online content. This approach provides a wealth of real-world linguistic diversity, enabling the training of MT models that are robust and adaptable to different language pairs and domains. In conjunction with traditional sources of parallel data such as multilingual websites, movie subtitles, holy texts, parliament proceedings, and software localization data, these modern initiatives greatly expand the availability of high-quality parallel data, facilitating the advancement of multilingual communication through MT technology.

{% hint style="info" %}
**Note:** It's important to recognize that the quality of training data can significantly impact the performance of machine translation models, especially for less-resourced languages. As highlighted in the paper "[Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets](https://arxiv.org/abs/2103.12028)," the success of large-scale pre-training and multilingual modeling in NLP has led to the emergence of numerous web-mined text datasets for various languages. However, lower-resource corpora often exhibit systematic issues, including unusable text, sentences of subpar quality, mislabeling, and nonstandard/ambiguous language codes. This phenomenon is particularly prevalent in languages with limited available resources. The study demonstrates that these quality issues are discernible even to non-proficient speakers, and automated analyses further support these findings. As you work with multilingual corpora, it's essential to consider these challenges and explore techniques to evaluate and enhance the quality of data sources.
{% endhint %}


# 7.2.1 Building your own MT models

In the field of machine translation, pre-trained models have brought a significant shift, offering valuable starting points for creating translation systems. Platforms like Hugging Face provide access to a range of pre-trained models suitable for various language pairs and directions. For instance, the [Helsinki-NLP](https://huggingface.co/Helsinki-NLP) repository houses both unidirectional and multilingual models trained using parallel data sourced from OPUS. Evaluations of these models across different benchmark datasets are accessible at[ https://opus.nlpl.eu/dashboard/](https://opus.nlpl.eu/dashboard/).

Recently, multilingual models have gained prominence, accommodating multiple language directions simultaneously. [Meta AI's NLLB model](https://ai.meta.com/blog/nllb-200-high-quality-machine-translation/) stands out, supporting a remarkable 200 languages. However, it's crucial to acknowledge that translation quality can differ substantially among various language pairs as the data quality and size differ substantially.

The BLEU scores of the [FLORES-200 benchmark set](https://github.com/facebookresearch/flores/blob/main/flores200/README.md) for the NLLB model offer insights into its performance across different language directions. While these models offer impressive capabilities, their practicality necessitates customization to suit specific language pairs and domains.

These pre-trained models are advantageous due to their adaptability. Researchers can download these models and further fine-tune them using their collected data given that they have access to the right computational resources.

In 2022, CLEAR Global has collaborated with Digital Umuganda to fine-tune the NLLB model in Kinyarwanda on two domains: Finance education and Tourism. You can find the source code for training and evaluation scripts together with the results in [this link.](https://github.com/Digital-Umuganda/twb_nllb_finetuning)

Several other well-known machine translation frameworks have also gained traction in the MT landscape. Each of these frameworks brings its own strengths and features to the table, catering to different needs and preferences in the realm of machine translation.

[**OpenNMT**](https://4bcplaybook.clearglobal.org/7-development-and-deployment-guidelines/7.2-machine-translation/www.opennmt.net), for instance, stands out as an open-source toolkit that provides comprehensive support for neural machine translation. With its modular design, OpenNMT offers flexibility in building and fine-tuning translation models.

You can consult [CLEAR Global’s codebase](https://github.com/translatorswb/TWB-MT) for training OpenNMT models. It contains the necessary scripts and short instructions for training and evaluating a neural machine translation system from scratch.

[**HuggingFace Transformers library**](https://huggingface.co/docs/transformers/index) was created with the objective of providing a unified interface for loading, training, and storing Transformer models, streamlining the process for NLP practitioners. Notably, the library boasts features such as ease of use, enabling users to download and employ cutting-edge NLP models for inference using just a few lines of code. Moreover, the library seamlessly integrates with the open model and datasets in their platform, making loading of a dataset or a model as easy as one line in Python. Refer to their [NLP course](https://huggingface.co/learn/nlp-course) for a general overview of NLP in practice, and the [Translation section](https://huggingface.co/learn/nlp-course/en/chapter7/4?fw=pt) for a deeper understanding of developing translation models.

[**MarianNMT**](https://marian-nmt.github.io/) is another notable framework that focuses on efficiency and high performance. It leverages advanced optimization techniques to achieve fast and accurate translations, making it a preferred choice for various MT applications.

[**JoeyNMT**](https://joeynmt.readthedocs.io/en/latest/tutorial.html) is recognized for its user-friendly interface and ease of use. Built on top of PyTorch, JoeyNMT simplifies the process of creating, training and deploying translation models. Its straightforward configuration and accessibility have made it popular among developers and researchers.

[**Sockeye**](https://github.com/awslabs/sockeye) is an open-source sequence-to-sequence framework for Neural Machine Translation built on[ PyTorch](https://pytorch.org/). It implements distributed training and optimized inference for state-of-the-art models, powering[ Amazon Translate](https://aws.amazon.com/translate/) and other MT applications.


# 7.2.2 Deploying your own scalable Machine Translation API

[**TWB-MT-fastapi**](https://github.com/translatorswb/TWB-MT-fastapi) is an open-source REST API designed for serving multiple machine translation models in production environments. It offers a range of features and supports various translation system types, including ctranslate2 models, Transformers-based models from Hugging Face, and custom rule-based translators.

#### **Features**

* Minimal REST API interface for using many models
* Batch translation
* Low-code specification for loading various types of models at start-up
* Automatic downloading of huggingface models
* Multilingual model support
* Model Chaining
* Translation pipeline specification (lowercasing, tokenization, subwording, recasing)
* Automatic sentence splitting (uses nltk library)
* Manual sentence splitting with punctuation
* Supports sentence piece and byte-pair-encoding models
* GPU support
* Easy deployment with docker-compose

#### **Usage Instructions**

1. **Model Configuration:** In the `config.json` file, you can specify the translation models, their settings, pre- and post-processors. You can configure various translation model types such as ctranslate2, Hugging Face transformer models, and custom rule-based translators.
2. **Custom Configuration Examples:** The repository README provides example configurations for different model types, including [ctranslate2](https://github.com/OpenNMT/CTranslate2), OPUS and OPUS-big models of [Helsinki-NLP](https://huggingface.co/Helsinki-NLP), [NLLB](https://huggingface.co/docs/transformers/v4.28.1/en/model_doc/nllb) models, and [M2M100](https://huggingface.co/docs/transformers/model_doc/m2m_100) models. These examples demonstrate how to set up the config.json file for each model type.
3. **Advanced Configuration Features:** The README explains how to use advanced features like alternative model loading, model chaining, and sentence splitting. These features allow you to customize your translation pipelines and choose appropriate models for various translation tasks.
4. **Building and Running:** The README guides you through the process of building and running the API. You can run it locally using a virtual environment and Python 3.8, or you can utilize Docker Compose for easy deployment.
5. **Example API Calls:** The README provides example cURL and Python code snippets to demonstrate how to use the API for simple translation, using alternative models, and performing batch translations. Additionally, it explains how to retrieve the list of supported languages and model pairs.


# 7.2.3 Evaluation and continuous improvement of machine translation

Evaluating the effectiveness of a machine translation (MT) system is a crucial step in ensuring the quality of its outputs and guiding its improvement over time. Automatic evaluation metrics, such as COMET, chrF++, BLEU and TER which is mentioned in [6.4 Assessing data and model maturity](/6.-language-technology-implementation/6.4-assessing-data-and-model-maturity), provide quantitative insights into the performance of an MT system by comparing its output to reference translations. While these metrics offer a quick and convenient way to assess translation quality, they might not always capture the nuances of language and the specific context of the use case. Therefore, while automatic evaluation methods are valuable, manual evaluation remains an essential component in gauging MT quality accurately.

Manual evaluation entails human judgment and understanding, making it particularly valuable for assessing the suitability of an MT system in real-world scenarios. When designing an evaluation strategy, it’s essential to align it with the specific use case of the MT system. For instance, in scenarios where the MT output is post-edited for translation by linguists, seeking feedback directly from these linguists is highly recommended.

One question we find simple to understand and useful for assessment is “*Has MT helped you with your work?*” with the answers:

1. No, it was useless;
2. No, it was not very helpful;
3. Yes, it was sometimes helpful;
4. Yes, it was very helpful

For direct quality assessment of MT output, one way of a rating of sentences or paragraphs by your linguists from a scale from 1 to 5 where:

1. The MT does not communicate the source meaning at all or is incomprehensible.
2. The MT communicates only part of the source meaning, important information is missing and the text is difficult to follow.
3. The MT communicates most of the source meaning but some details are incorrect or there is an awkward style.
4. The MT communicates the source meaning, but terminology needs to be improved.
5. The MT conveys the source meaning and sounds natural.

However, it’s important to note that manual evaluation goes beyond quantitative metrics. Qualitative assessment from linguists can provide invaluable insights into where the MT system consistently falls short. Examples of consistent errors, such as gender mix-ups or term mistranslations, shed light on areas that need refinement. By analyzing these qualitative assessments, developers can uncover patterns of errors and focus on addressing specific linguistic challenges. This assessment can also guide your next data collection strategy.

Additionally, both the process of evaluation and deployment should be designed to facilitate continuous improvement. In a post-editing setup, collecting post-edited MT outputs forms a valuable dataset for training and refining the MT model in subsequent iterations. This data can be used to fine-tune the model’s performance and address the specific challenges identified during manual evaluation. Deployment, in turn, should make it possible to ask for user feedback to flag mistranslations and errors.

In conclusion, while automatic evaluation metrics provide a preliminary assessment of MT quality, manual evaluation and feedback from linguists are essential for a comprehensive understanding of its effectiveness. By aligning the evaluation process with the use case, implementing structured quality ratings, and incorporating qualitative feedback, developers can ensure that the MT system evolves to meet the demands of real-world translation scenarios. Furthermore, a well-designed evaluation process becomes an integral part of the continuous improvement loop, enhancing the MT system’s performance over time.


# 7.3 Chatbots

In this section, we’ll detail out all the processes involved in deploying an FAQ chatbot, from NLU data creation to server deployment and channel integrations. In this section, we present RASA as the chatbot framework and walk through a tutorial based on a FAQ chatbot for building climate resilience. Later in the chapter, you can find specific instructions to integrate your chatbot into various communication channels, tips on optimizing your NLU data and continous evaluation methodologies.&#x20;


# 7.3.1 Overview of chatbot technologies and RASA framework

Rasa is an open source machine learning framework for automated text and voice-based conversations. The documentation for the 2.x version, which we use, can be accessed through [this link](https://rasa.com/docs/rasa/2.x/). Here we will give a brief explanation of the workings of a chatbot.

The user interacts with the chatbot through a **messaging interface**, which can be a website, messaging app, or voice assistant. The user sends messages, and the chatbot responds with text or voice.

**Natural language understanding (NLU)** is a crucial component that processes the preprocessed messages. It involves tasks like intent recognition (understanding what the user wants) and entity recognition (identifying specific pieces of information within the user's message).

Based on the recognized intent and entities, the chatbot's dialogue management system decides how to respond. This involves selecting the appropriate response from a pool of predefined responses or generating a dynamic response.

If the response is not a straightforward static message, **Natural Language Generation (NLG)** comes into play. It generates human-like text that is coherent and contextually relevant to the user's query.

In some cases, the chatbot may need to interact with external systems or APIs to fetch information or perform specific actions, like retrieving weather data or booking flights.

Finally, the chatbot assembles the final response, which may include text, images, or links, and sends it back to the user through the messaging interface.

**How does a chatbot learn?**

The learning process of a chatbot involves collecting conversation data from user interactions, which is used to train machine learning models (NLU/NLG) responsible for understanding user intents and generating responses. Human annotators or linguistic tools label the intents and entities in the conversation data to create labeled training sets. Machine learning models are trained on this data to recognize patterns and context in user input. User feedback, both positive and negative, further guides the system's learning. Advanced chatbots may employ reinforcement learning to fine-tune responses based on feedback. Periodic updates to the models incorporate new data and insights, followed by rigorous testing and validation. Continuous monitoring ensures ongoing improvement, while domain experts can provide expertise to enhance accuracy and domain-specific knowledge. This iterative process allows the chatbot to continually refine its understanding and responses, resulting in a more effective and user-friendly conversational experience.

Overall, building a chatbot involves integrating various components such as NLU, dialogue management, and NLG to create a seamless conversational experience for users. It's an iterative process that requires collaboration between developers, linguists, domain experts, and user feedback to refine the chatbot's performance and capabilities.

A chatbot doesn’t really learn the language itself from the provided training data. The most common practice is that the chatbot framework utilizes a foundational language model as its starting point, which serves as a fundamental understanding of the form of the language. This base language model is pre-trained on vast amounts of text from the internet, allowing it to grasp grammar, vocabulary, and context. However, this general language model needs customization to serve a particular purpose. The chatbot developer then fine-tunes the model using curated training data, which includes examples of user interactions specific to the desired domain or task. This adaptation process refines the language model's responses, ensuring that the chatbot can comprehend intents, generate relevant responses, and engage users effectively within the intended context.

**Training data format**

Depending on the chatbot framework you use, the exact format of the training data can vary. RASA, for example, uses YAML format files for storing NLU training data, answers, stories, etc. For more information, refer to the [official documentation](https://rasa.com/docs/rasa/2.x/training-data-format).


# 7.3.2 Building data for a climate change resilience chatbot

In this section, we provide a tutorial for developing all the necessary files for creating a RASA-based climate resilience FAQ chatbot. This tutorial is based on a project in collaboration between CLEAR Global and Gram Vaani for the farmers in the Bihar region of India. For more information you can refer to CLEAR Global's blogpost:&#x20;

{% embed url="<https://clearglobal.org/using-ai-to-support-farmers-to-adapt-to-climate-change/>" %}
Blogpost explaining TILES project in India
{% endembed %}

When preparing an FAQ-based bot, you need to start by defining the following:

1. Topics you want to cover
2. List of possible questions you can receive
3. Proper answers to those questions
4. The different ways those questions can be uttered by your users.

It’s useful to curate this type of data in a spreadsheet where your linguists, interaction designers, and content specialists can easily collaborate. Having it in this structured form also helps your developers pull the content automatically to create the specifically formatted files for training and testing chatbot models.

When starting to prepare chatbot content, it’s convenient to work on a format that’s easily editable by both technical and non-technical profiles. A simple way to do this is through a spreadsheet.

As part of this tutorial, we are providing both a publicly accessible spreadsheet and a codebase that pulls data automatically from this spreadsheet to create RASA format files:

1. To access the Climate Resilience FAQ sheet, [click here](https://docs.google.com/spreadsheets/d/1OpTyjwXZjItPugQpprQ86oaV5xU1fUS0A_QELrQMg5U/edit?usp=sharing).
2. To access the Python-based scripts and instructions on Github, [click here.](https://github.com/translatorswb/rasa_climate_change_data)
3. In our spreadsheet, we find three main topics represented in the three sheets:
   1. Climate Change T1 on definitions,
   2. Climate Change T2 on impact,
   3. Climate Change T3 on adaptation methods and government programs

Each topic has a list of FAQs. Let’s take the first FAQ from the first sheet:

![An example of chatbot datasheet](/files/9KQjFgQ8xvg4FhgCaoDe)

This data encapsulates all the information necessary for a chatbot to learn how to receive a question related to the definition of climate change and how to answer it. These encapsulations are also referred to as intents. We will dive deeper into the technical definition of intents in [7.3.7 How to create effective NLU training data](/7-development-and-deployment-guidelines/7.3-chatbots/7.3.7-how-to-create-effective-nlu-training-data).

In this particular tutorial, we are working with two languages, three main topics, and 25 FAQs in total.


# 7.3.3 How to obtain multilinguality

#### **7.1.3 How to obtain multilinguality** <a href="#heading-h.j9r6yy8tkpob" id="heading-h.j9r6yy8tkpob"></a>

The training data we prepared in the previous section contained both English and Hindi, making the resulting bot multilingual. Multilinguality is not just an added functionality; it depends on the local context. It is sometimes the most logical way. Multilingual chatbots, capable of understanding and responding in multiple languages, hold distinct advantages in certain scenarios. In regions with linguistic diversity or areas where multiple languages are commonly spoken, a single chatbot that accommodates different languages can greatly enhance accessibility and user engagement. Moreover, multilingual bots can facilitate communication in cross-border or multicultural contexts, streamlining interactions and information exchange. In this section, we'll explore how to harness the power of multilinguality for your chatbot, enabling it to serve a wider audience while maintaining a seamless conversational experience.

In terms of training data preparation, one might replicate intents and responses for each language to enable multilingual support in a chatbot. For example:

Intents:

* greet\_eng
* greet\_swh
* greet\_fra

Responses:

* utter\_greet\_eng
* utter\_greet\_swh
* utter\_greet\_fra

While this approach initially seems straightforward, it presents challenges as the number of intents grows. Training the model with a limited number of samples can lead to decreased performance and the potential misclassification of similar intents.

To overcome these challenges, we suggest a robust multilingual architecture that leverages different NLU models for every request. This architecture employs a language classifier, which identifies the language of each incoming query. Based on the detected language, the query is directed to the specific NLU model responsible for intent recognition in that language. Once the intent is recognized, it is passed to the core model, which predicts the appropriate action, such as generating a response or triggering a custom action.

The below diagram shows an example of this that was used in one of CLEAR Global’s chatbot systems:

<figure><img src="/files/ECk8qnGx8RoNKTozMqT8" alt=""><figcaption><p>Multilingual chatbot deployment architecture</p></figcaption></figure>

By implementing this multilingual architecture, you can ensure accurate intent recognition and provide language-specific responses for a seamless user experience. Throughout this documentation, we will explain the architecture, how to train the models, and how to run the chatbot.


# 7.3.4 Components of a chatbot in deployment

The chatbot system presented above consists of the following components:

1. **User Interface:** This component represents the end-user-facing interface where users interact with the chatbot. It can be a web application, mobile app, or any other platform like Telegram, WhatsApp, or Facebook Messenger.
2. **Nginx:** Nginx is a web server and reverse proxy server that acts as the entry point for incoming requests to the chatbot system. It handles SSL termination, load balancing, and routing of requests to the appropriate services.
3. **Rasa Production:** Rasa Production is a service responsible for running the chatbot in production mode. It handles incoming user messages, triggers appropriate actions, and requests responses from NLG server. Rasa Production runs the Rasa core model.
4. **Rasa X:** Rasa X is an open-source conversational AI platform that provides a user interface for training and managing chatbots. It offers features such as conversation management, and interactive chatbot testing.
5. **Duckling:** Duckling is a language parsing server used for extracting structured information from text. It can understand and extract entities like dates, numbers, and durations from user messages. Duckling is integrated into the chatbot system to enhance NLU capabilities.
6. **Redis:** Redis is an in-memory data structure store that acts as a cache and message broker within the chatbot system. It is used to store temporary data, session information, and a message queue for inter-service communication.
7. **RabbitMQ:** RabbitMQ is a message broker that enables asynchronous communication between services. It provides reliable message delivery and ensures decoupling between components. RabbitMQ is used for communication between Rasa Production and Redis.
8. **PostgreSQL:** PostgreSQL is a powerful open-source relational database used for storing chatbot-related data. It is utilized by Rasa X for storing conversation logs, user information, and other metadata.

\
**Docker Compose**

The chatbot system is containerized using [Docker](https://www.docker.com/) and orchestrated using [Docker Compose](https://docs.docker.com/compose/). The docker-compose.yml file defines the services and their configurations. Let's briefly review the key services:

**Rasa Production**

This service runs Rasa in production mode and handles user interactions. It is based on the `rasa/rasa:${RASA_VERSION}-full` Docker image.

Rasa Production serves as the runtime environment for the Rasa Core model. It acts as the message processing unit for the chatbot app. Rasa Production interacts with the NLU-NLG server for natural language understanding (NLU) and natural language generation (NLG) capabilities.

To communicate with the NLU-NLG server, Rasa Production sends messages to the server's address specified in the endpoints.yml file:

```
nlg:
    url: "http://nlu-nlg-server:6001/nlg"
nlu:
    url: "http://nlu-nlg-server:6001"
```

It initiates a call to the nlu-nlg-server, which employs the NLU model tailored for the detected language to classify intents. The nlu-nlg-server then relays this information back to Rasa production, which determines the appropriate course of action based on the received intent.

If the decision involves executing an action, Rasa Production makes a call to the 'app' container, running the Rasa actions, through the exposed port 5055 on the ‘app’ container. The necessary files for the 'app' container are copied from the 'data' folder defined by the GIT\_DIR variable.

Similarly, for generating a response, Rasa production interacts with the NLG server. It sends the message to the NLG server(operating in the container ‘nlu-nlg-server), which, based on the detected language, relays the response back to Rasa Production. Once the nlu-nlg-server generates the response, Rasa Core relays it to the connected service (e.g., Telegram) as a reply to the user's query.

In summary, Rasa Production acts as a message processing unit and decision maker based on parameters set in the Rasa core model. It orchestrates the processes involved in understanding the message sent by the user and generating a response.

### **PORT: 5005**

### **nlu-nlg-server**

The nlu-nlg-server is a container that plays a crucial role in enabling multilingual capabilities for the chatbot. Within this container, two servers are running: the NLU server and the NLG server. Communication between these servers takes place through port 6001.

To achieve seamless language switching, the nlu-nlg-server utilizes the functionality provided by Rasa Helpers. The configuration for Rasa Helpers is defined in the 'rh\_config.yml' file, located in the GIT\_DIR directory.

Rasa Helpers offers a convenient way to handle language classification, loading NLU models, and generating NLG responses. With this functionality, you can define and provide models for language classification. Based on the classified language, the NLU-NLG server selects the appropriate NLU model specifically designed for that language.

Similarly, response files are specified and provided, allowing Rasa Helpers to generate NLG responses using the detected language. This integration of language classification, loading NLU models, and NLG generation is all facilitated by Rasa Helpers, ensuring efficient multilingual capabilities for the chatbot.

### **PORT:6001**

**app**: The app container runs Rasa actions server on port 5055

**duckling**: This service runs the Duckling language parsing server to extract entities from user messages.

**redis:** This service provides an in-memory cache and message broker using the Redis server.

**Rasa X**: This service runs Rasa X and provides the user interface for managing the chatbot. It is based on the Rasa/Rasa Docker image.

**rabbit**: This service sets up RabbitMQ as the message broker for inter-service communication.

**db:** This service sets up the PostgreSQL database for storing chatbot-related data.


# 7.3.5 Deploying a RASA chatbot

The instructions below assume a certain level of technical knowledge and familiarity with concepts like Docker, Git, and Rasa. The text aims to guide readers through the process of deploying and managing a Rasa chatbot, ensuring they have the necessary system requirements, and providing instructions for each step of the deployment process.

### System requirements:

#### Docker Compose Development

<table data-header-hidden><thead><tr><th width="217"></th><th></th></tr></thead><tbody><tr><td><strong>Operating System</strong></td><td><em>Ubuntu 18.04 / 20.04</em> Debian 9 / 10 <em>CentOS 7 / 8</em> RHEL 8</td></tr><tr><td><strong>vCPUs</strong></td><td><em>Minimum: 2 vCPUs</em> Recommended: 2-6 vCPUs</td></tr><tr><td><strong>RAM</strong></td><td>8 - 16 GB RAM</td></tr><tr><td><strong>Disk Space</strong></td><td>* Recommended: 100 GB disk space available</td></tr></tbody></table>

### Port requirements

When creating the server, make sure the following ports are open:

<table data-header-hidden><thead><tr><th width="138"></th><th width="166"></th><th></th></tr></thead><tbody><tr><td><strong>Port</strong></td><td><strong>Service</strong></td><td><strong>Description</strong></td></tr><tr><td>22</td><td>SSH</td><td>SSH access</td></tr><tr><td>80</td><td>HTTP</td><td>Web application access</td></tr><tr><td>443</td><td>HTTPS</td><td>Web application over HTTPS access</td></tr><tr><td>5432</td><td>TCP</td><td>External applications accessing the database</td></tr></tbody></table>

The minimum configuration should support up to 50 concurrent users, so it's ideal to use the minimum configuration for development and staging environments. The maximum configuration can support up to 400 concurrent users, so it can be used in a production environment. If the production load goes beyond 400 users, then it's better to look for a Kubernetes deployment.

#### Kubernetes deployment

<table data-header-hidden><thead><tr><th width="220.33333333333331"></th><th width="192"></th><th></th></tr></thead><tbody><tr><td><strong>Deployment</strong></td><td><strong>CPU</strong></td><td><strong>Memory</strong></td></tr><tr><td>rasa-x</td><td>1</td><td>1 GiB</td></tr><tr><td>event-service</td><td>2</td><td>1 GiB</td></tr><tr><td>rasa-production</td><td>2</td><td>2 GiB</td></tr><tr><td>rasa-worker</td><td>4</td><td>4 GiB</td></tr><tr><td>nginx</td><td>0.2</td><td>200 MiB</td></tr><tr><td>app</td><td>0.5</td><td>200 MiB</td></tr><tr><td>duckling</td><td>0.5</td><td>200 MiB</td></tr><tr><td>postgresql</td><td>1</td><td>250 MiB</td></tr><tr><td>rabbit</td><td>0.2</td><td>250 MiB</td></tr><tr><td>redis</td><td>0.2</td><td>250 MiB</td></tr></tbody></table>

We recommend a size of 10 GiB for the Rasa X volume claim and at least 30 GiB for the database volume claim.

### Installation process:

Install git and clone the following repository:

`git clone`[ `https://github.com/translatorswb/chatbot_deployment.git`](https://github.com/translatorswb/chatbot_deployment.git)

The default installation path is /etc/rasa, however, you can install in your desired path if you set:

```
export RASA_HOME=~/rasa/dir
cd chatbot_deployment
sudo chmod 777 install.sh
sudo -E bash ./install.sh # -E will preserve the environment variable you set
```

Details about install.sh file:

This script will install:

* Python
* Docker
* Docker Compose
* Ansible

You can skip this script if you have all of them installed, or you can comment on the specific commands and continue the rest of the script.

You can skip this section if you want everything to be installed by the script automatically.

1. `set -Eeuo pipefail`: This line sets some options for the script:
   * `-E`: Causes the script to exit immediately if any command within it exits with a non-zero status.
   * `-e`: Causes the script to exit immediately if any command within it returns an error.
   * `-u`: Treats unset variables as an error and causes the script to exit.
   * `-o pipefail`: Causes a pipeline to fail if any command within it fails.
2. `source /etc/os-release`: This line sources the `/etc/os-release` file, which contains information about the operating system distribution.
3. `RUN_ANSIBLE=${RUN_ANSIBLE:-true}`: This line assigns the value `true` to the variable `RUN_ANSIBLE` if it is not already set.
4. The following lines check the value of the variable ID, which is obtained from the sourced `/etc/os-release` file, and install necessary dependencies based on the operating system:
   * If the ID matches "`centos`" or "`rhel`" (Red Hat Enterprise Linux), it performs a series of commands using yum package manager to update the system, install Python 3, and install python3-distutils if available.
   * If the ID matches "`ubuntu`" or "`debian`", it performs a series of commands using apt-get package manager to update the system, install Python 3, and install python3-distutils if available. It also installs `wget` if it is not already installed.
5. `curl -O https://bootstrap.pypa.io/get-pip.py`: This line uses curl to download the `get-pip.py` script from the given URL.
6. `sudo python3 get-pip.py`: This line executes the downloaded `get-pip.py` script with sudo privileges to install `pip` for Python 3.
7. `sudo /usr/local/bin/pip install "ansible>-2.9, <2.10"`: This line uses pip to install Ansible within the specified version range.
8. `sudo /usr/local/bin/ansible-galaxy install geerlingguy.docker`: This line uses `ansible-galaxy` to install the "`geerlingguy.docker`" Ansible role.
9. The following section starts with `if [[ "$RUN_ANSIBLE" == "true" ]]; then` and ends with `fi`. It checks if the `RUN_ANSIBLE` variable is set to `true` and, if so, executes the following commands related to running the Ansible playbook.
10. `sudo /usr/local/bin/ansible-playbook -i "localhost," -c local rasa_x_playbook.yml`: This line uses ansible-playbook to execute the `rasa_x_playbook.yml` playbook, targeting the `localhost` inventory, and using the "local" connection method.

Then, from the command line go to `/etc/rasa` or `$RASA_HOME` and follow the steps below:

```
cd /etc/rasa # or cd $RASA_HOME
sudo docker-compose build
```

Make sure all the images are built successfully.

To run the server:

`sudo docker-compose up -d`

Now open localhost or the server ip/domain in your browser, and you should be able to access Rasa X.

If the server is still not up, wait for a few seconds and then try it, as some containers take some time to be up. If the problem persists, then follow the steps below:

Check running containers:

```
docker ps
```

There must be a container that might be stuck in a restart loop; just copy the name of the container and check its logs. Please make sure the name of the container matches the name of the service defined in `docker-compose.yml` or `docker-compose.override.yml`. For example: if the name of the container is rasa\_nginx\_1 then the name of the service is nginx. To check the logs, run the following command:

```
sudo docker-compose logs -f nginx
```

Once the app has started and database migrations are complete, visit <http://YOUR\\_IP\\_OR\\_DOMAIN> on your browser.

To reset the default password, run `sudo python rasa_x_commands.py create --update admin me YOUR_PASSWORD`

### Integrated version control (IVC)

1\. Ensure you have a deployed Rasa Enterprise instance and a Git repository with the default Rasa Open Source project layout.

2\. Take caution, as connecting your remote Git repository will overwrite the existing training data in Rasa Enterprise. If you want to keep the old training data, either use a fresh Rasa Enterprise instance or export the data before connecting the repository.

### Creating your Rasa project data repository

Make sure your project follows the default Rasa Open Source project layout with files like config.yml, data/nlu.yml, data/stories.yml, and domain.yml.

![](/files/onJLElSCU4mVI1CTOdSd)

Some NLU data to get started with different use cases:

<https://github.com/RasaHQ/NLU-training-data>

If you want all the data, including stories, rules, and domains; you can use the data provided here by Rasa:

<https://github.com/RasaHQ/rasa-demo/tree/main/data>

Just make sure you download all the files and reformat your repo with the same structure as shown in the picture above.

### Linking to Git Repository

RasaX connects to a git repository where all training data is maintained. Steps are mentioned below to link your repository.

In the Rasa Enterprise UI, click on the branch icon and select "Connect to a repository" to begin configuring the repository connection.

![](/files/3UeQGHM3MfOoFCTJUTQ9)

Rasa Enterprise supports GitHub, GitLab, and Bitbucket as Git platforms. Choose the appropriate platform and provide the necessary repository URL.

Configure your credentials based on the chosen connection method:

* SSH: Provide the SSH URL and configure the SSH key authentication.

In case you haven't added or generated an SSH key yet, you can follow these steps:

Generate an SSH key pair using the ssh-keygen command in your terminal:

```
ssh-keygen -t rsa -b 4096 -C "email@example.com"
```

Specify the path where you want to save the SSH key. For example:

Enter file in which to save the key `(/home/yourusername/.ssh/id_rsa): /path/to/your/ssh/key`

You can leave the passphrase empty for no passphrase or provide a passphrase for additional security. After generating the SSH key pair, you can proceed with adding the public key to your Git server by following the aforementioned instructions.

#### **Target Branch:**

Set the target branch, which will be used to show initial data, branch off for new changes, and return to after discarding or pushing changes. Users can choose to push changes directly to the target branch or create a new branch. If you want to disable direct pushing, select the option to require users to add changes to a new branch.

After configuring the repository credentials and branch options, click the "Verify Connection" button to establish the connection between Rasa Enterprise and your Git repository.

![](/files/M9ydLo6Jv1JhpPKVokwH)

By following these steps, you will successfully connect your Git repository to Rasa.

### **Rasa training**

Once you connect your Rasa project with GitHub, you can start training your model by going into ‘Models’ and clicking ‘train model’.

![](/files/kISd8TL6J8bubORBIC6e)

Once the training is complete, you first need to activate your model in order to talk to your bot.

In the Rasa X UI, you can explore and interact with your assistant, review conversations, and improve its performance through the Conversations and Training Data sections.

4\. Configure endpoints

To configure endpoints, open endpoints.yml in /etc/rasa or $RASA\_HOME directory and add the desired configurations. A server restart is required after changing that file.

### **NLG Server**

```
nlg:
    url: "http://nlg-server:6001/nlg"
```


# 7.3.6 Channel integrations

In this section, we explain integration instructions for three of the most commonly used chat platforms: Facebook Messenger, WhatsApp, and Telegram.

{% content-ref url="/pages/mplfA7dDUQHFdpKPVXXK" %}
[7.3.6.1 Facebook Messenger](/7-development-and-deployment-guidelines/7.3-chatbots/7.3.6-channel-integrations/7.3.6.1-facebook-messenger)
{% endcontent-ref %}

{% content-ref url="/pages/N17UN3jRINVpTWHigzAS" %}
[7.3.6.2 WhatsApp](/7-development-and-deployment-guidelines/7.3-chatbots/7.3.6-channel-integrations/7.3.6.2-whatsapp)
{% endcontent-ref %}

{% content-ref url="/pages/EvaCUsPSUbdqIvLhFYe5" %}
[7.3.6.3 Telegram](/7-development-and-deployment-guidelines/7.3-chatbots/7.3.6-channel-integrations/7.3.6.3-telegram)
{% endcontent-ref %}


# 7.3.6.1 Facebook Messenger

To serve your chatbot through Facebook Messenger, you first need to set up a Facebook page and app to get credentials. Once you have them, you can add them to your credentials.yml.

### **How to get the Facebook credentials**

You need to set up a Facebook app and a page.

1. To create the app, head over to [Facebook for Developers](https://developers.facebook.com/) and click on **My Apps** → **Add New App**.
2. Go onto the dashboard for the app, and under **Products**, find the **Messenger** section and click **Set Up**. Scroll down to **Token Generation** and click on the link to create a new page for your app.
3. Create your page and select it in the dropdown menu for the **Token Generation**. The shown **Page Access Token** is the page-access-token needed later on.
4. Locate the **App Secret** in the app dashboard under **Settings** → **Basic**. This will be your secret.
5. Use the collected secret and page-access-token in your credentials.yml, and add a field called verify containing a string of your choice. Start Rasa run with the --credentials credentials.yml option.
6. Set up a **Webhook** and select at least the **messaging** and **messaging\_postback** subscriptions. Insert your callback URL, which will look like `https://<host>:<port>/webhooks/facebook/webhook`, replacing the host and port with the appropriate values from your running Rasa X or Rasa Open Source server.
7. Insert the **Verify token,** which has to match the verify entry in your credentials.yml.

**CONFIGURE HTTPS**

Facebook Messenger only forwards messages to endpoints via https, so take appropriate measures to add it to your setup.

For more detailed steps, visit the [Messenger docs](https://developers.facebook.com/docs/graph-api/webhooks).

To use https on localhost, please check: [Setup HTTPS on localhost for testing the API/Channels](https://translatorswb.slab.com/posts/setup-https-on-localhost-for-testing-the-api-channels-uevn4qbd)

### Running On Facebook Messenger

Add the Facebook credentials to your credentials.yml:

```
facebook:
verify: "rasa-bot"
secret: "<secret>"
page-access-token: "<token>"
```

Restart your Rasa X or Rasa Open Source server to make the new channel endpoint available for Facebook Messenger to send messages to.

Debugging

To check the activity on the webhook, we can see the nginx logs:

```
sudo docker-compose logs -f nginx
```

Try a message on messenger and you should be able to see the following:

```
nginx_1 | some_ip - - [09/Sep/2021:06:12:06 +0000] "POST /webhooks/facebook/webhook HTTP/1.1" 200 7 "-" "facebookexternalua" "-"
```

If you are not able to see a log like the one above, then the integration is not done properly, so we need to check the credentials and try again.

Making the bot available to public

With the above integration, the bot can only be tested by the[ test users](https://developers.facebook.com/docs/development/build-and-test/test-users/). To make it public, we need to submit the application for App Review by requesting **pages\_messaging** permission/feature. Furthermore, check out a sample submission here:

<https://developers.facebook.com/docs/app-review/resources/sample-submissions/messenger-platform>


# 7.3.6.2 WhatsApp

Before setting up a WhatsApp Business Account, you first need a Facebook Business Manager Account (refer to Messenger Integration)

To start, you need WhatsApp Business Solution Provider (Infobip, Turn.io,etc.)

They are a global community of third-party solution providers with expertise in the WhatsApp Business API. These BSPs can help you communicate with your customers on WhatsApp for the approved use cases of customer support and time-sensitive, personalized notifications.

* Once you have a WhatsApp Business Solution Provider, if you are making a chatbot, you need to add the name of the bot to the list of our projects on our website
* Then request for a WhatsApp Business Application from our BSP and fill in the presented form
* After this process, you need to send this filled form to the BSP, wherein there will be some contracts to sign and the integration of the bot will be done


# 7.3.6.3 Telegram

You first have to create a Telegram bot to get credentials. Once you have them you can add these to your credentials.yml.

**How to get the Telegram credentials**

You need to set up a Telegram bot.

1. To create the bot, go to [Bot Father](https://web.telegram.org/#/im?p=@BotFather), enter `/newbot`, and follow the instructions. The URL that Telegram should send messages to will look like `http://<host>:<port>/webhooks/telegram/webhook`, replacing the host and port with the appropriate values from your running Rasa X or Rasa Open Source server.
2. At the end you should get your access\_token and the username you set will be your verify.
3. If you want to use your bot in a group setting, it's advisable to turn on group privacy mode by entering `/setprivacy`. Then the bot will only listen when a user's message starts with /bot.

For more information, check out the [Telegram HTTP API](https://core.telegram.org/bots/api).

### Running on Telegram

Add the Telegram credentials to your credentials.yml:

```
telegram:
access_token: "<token>"
verify: "your_bot"
webhook_url: "https://your_url.com/webhooks/telegram/webhook"
```

Restart your Rasa X or Rasa Open Source server to make the new channel endpoint available for Telegram to send messages to.

Also check:[ Setup HTTPS on localhost for testing the API/Channels](https://translatorswb.slab.com/posts/setup-https-on-localhost-for-testing-the-api-channels-uevn4qbd)

Supported Response Attachments

In addition to standard text: responses, this channel also supports the following components from the [Telegram API](https://core.telegram.org/bots/api/#message):

* button arguments:
  * button\_type: inline | vertical | reply
* custom arguments:
  * photo
  * audio
  * document
  * sticker
  * video
  * video\_note
  * animation
  * voice
  * media
  * latitude, longitude (location)
  * latitude, longitude, title, address (venue)
  * phone\_number
  * game\_short\_name
  * action

Examples:

```
utter_ask_transfer_form_confirm:
- buttons:
- payload: /affirm
title: Yes
- payload: /deny
title: No, cancel the transaction
button_type: vertical
text: Would you like to transfer {currency}{amount_of_money} to {PERSON}?
image: "https://i.imgur.com/nGF1K8f.jpg"
Copy
utter_giraffe_sticker:
- text: Here's my giraffe sticker!
custom:
sticker: "https://github.com/TelegramBots/book/raw/master/src/docs/sticker-fred.webp"
```


# 7.3.7 How to create effective NLU training data

In this section, we provide practical insights and tips for creating robust Natural Language Understanding (NLU) training data that empowers chatbots to accurately interpret user intent. From understanding the pivotal role of intents to ensuring a diverse set of training examples, we delve into intent merging, entity extraction, and maintaining balanced training data. Join us as we uncover key strategies for enhancing the NLU capabilities of your chatbots.

**The&#x20;*****intent*****&#x20;of a message is what a person wants to achieve**

When we say something, we often try to achieve a specific goal. These may be some ways of *greeting someone*:

* *Hello there!*
* *Good morning.*
* *Hi!*

These are some ways to *book a table at a restaurant*:

* *I'd like to book a table for two at noon, please.*
* *Do you have space for two people for lunch?*
* *I want a table for two at lunchtime.*

The *intent* (or intention) of a message is the goal that a person is trying to achieve with this message. In the two series of examples above, the way the message is delivered varies, but the *intent* stays the same.

**Humans understand the intent of a message intuitively**

For us humans, it is easy (most of the time) to understand another human's intention based on what they say.

Would it be possible for you to set aside a flat surface in your establishment, for one plus one people, so that they can consume a bit of food when the sun is at its peak?

This is a convoluted way of expressing the same goal as in the previous section: to *book a table at a restaurant*. Yet it is still understandable, we do not need to have seen the exact same wording before in order to understand the intent of that sentence.

We are capable of doing that because we have outside knowledge like:

* We know tables are flat
* We know 1 + 1 = 2
* We know that the sun is (usually) at its peak at around 12:00

This makes us capable of deducing other facts about this person's request, even if they are not (as) explicit: the number of people (2), the time of the reservation (12:00), etc.

We could even *try* to guess some things about the speaker based on the way they wrote: what kind of people they are, their age, etc.

**Intents and chatbots**

Unlike us humans, chatbots only have access to messages and do not understand its context. A chatbot does not have access to the same knowledge as a human, all of its knowledge comes from the textual information it receives.

Its entire world consists of messages, and they are the only things it can use to recognize their intents.

There are two stages that we need to distinguish when working with a chatbot:

* The *learning* or *training* phase: the chatbot is given *training data* in order to become better at recognizing the intents of messages
* The *prediction* phase: the chatbot has been trained, and we're now asking it to *predict* (recognize) the intents of *unseen* messages

**Chatbots need to&#x20;*****learn*****&#x20;how to recognize the intent of a message**

When a chatbot *learns*, it looks at many messages and the intent associated with each of those messages. These messages are called *training examples*.

Let's recap, using the same message as before:

Would it be possible for you to set aside a flat surface in your establishment, for one plus one people, so that they can consume a bit of food when the sun is at its peak?

* A human has seen, and heard many things, has been in many different situations and talked with many different people. It has access to all of this information in order to guess the intent of the message above.
* A chatbot has *only ever seen messages* and has no access to other kinds of information. It guesses the intent of the message above by finding out *how similar* the message is to all the *previous messages it has seen before*.

They are the only things that the chatbot can use to understand the intent of as it only has access to the message itself.

Its entire world consists of the messages, and they are the only things it can use to deduce their intents.

Unfortunately, understanding what a human means is not so intuitive for chatbots.

A chatbot's entire world consists of the messages it receives

It would be very hard to make chatbots process information in the same way humans do since we don't even know how exactly humans do it.

For a chatbot, *learning* to understand users means to correctly guess a user's intent based on a message they sent.

**A good intent has&#x20;*****a lot*****&#x20;of&#x20;*****diverse*****&#x20;training examples**

The chatbot needs to learn there are different ways of saying the same thing:

**Bad:**

* *Can covid be spread by animals?*
* *Can covid be spread by mosquitoes*
* *Can covid be spread by flies?*

**Better:**

* *Can covid be spread by animals?*
* *Do mosquitoes transmit corona?*
* *Are flies capable of giving me the coronavirus?*

Annotating user messages is a good way to gather diverse training examples. Here are real user messages in Uji, CLEAR Global’s COVID-19 related chatbot deployed in DRC:

* *Je voudries connaitre le cas confirm de ventre 19*
* *Savoir actuellement le nombre de cas dans chaque province...*
* *La situation épidémiologique de COVID-19 en ce jour*

These are all asking about `disease_stats`.

It's okay that **some** training examples look similar to one another, but they must not **all** be the same.

**A lot of training examples**

Chatbots need a lot of data to learn well.

Rasa recommends a **minimum** of 75 training examples for each intent. It really is *a lot* for our purposes, but we should still aim for it.

There are a few ways to increase the number of training examples in an intent:

* Write more training examples by hand
* Observe real user messages
* Merge similar intents together

The easiest way is to annotate user messages. Annotating also helps with making the training data more diverse.

{% hint style="info" %}
**Training examples checklist**

When looking at an intent's training data, check that:

* There are **at least** **35** training examples. More is even better.
* The training examples are **diverse**.
* Some (as many as possible) training examples come from **real users**.
  {% endhint %}

**An intent's meaning should be distinct from other intents**

**Merge similar intents together**

The chatbot gets confused when the meaning of two or more intents is too similar. For example:

`covid_myth_heat_kills`

* *Can corona survive in the heat?*
* *Is there a temperature that kills COVID-19?*
* *Can corona not be transmitted when it's hot?*
* *Can COVID-19 survive in humid temperatures?*

`covid_myth_cold_kills`

* *Does the cold kill the coronavirus?*
* *Does snow kill corona?*
* *What temperature is worst for COVID-19?*
* *Does the coronavirus die in snow?*

The meaning of the intents covid\_myth\_heat\_kills and covid\_myth\_cold\_kills are similar. It could cause issues to the chatbot, and so we could try to merge the two intents. It can help in two ways:

* Removing the potential for confusion between the two intents
* Increasing the number of training examples in the new, merged intent

There are two ways of merging intents with overlapping meanings.

**The intents to merge have similar answers**

The answers to covid\_myth\_heat\_kills and covid\_myth\_cold\_kills are:

`covid_myth_heat_kills`

*COVID-19 MAY be transmitted in areas with a hot and humid climate. Exposure to the sun or high temperatures DOES NOT PREVENT against contracting coronavirus disease.*

`covid_myth_cold_kills`

*Cold weather and snow CANNOT kill COVID-19.*

They could be rewritten and merged into a single answer, for example:

`covid_myth_weather_kills`

*You can catch COVID-19 regardless of an area's climate. Hot and humid weather and exposure to the sun do not prevent the transmission of the coronavirus. Cold weather and snow do not kill the coronavirus either.*

In that case, we can simply put all the training examples from before into a single intent:

`covid_myth_weather_kills`

* *Can corona survive in the heat?*
* *Is there a temperature that kills COVID-19?*
* *Can corona not be transmitted when it's hot?*
* *Can COVID-19 survive in humid temperatures?*
* *Does the cold kill the coronavirus?*
* *Does snow kill corona?*
* *What temperature is worst for COVID-19?*
* *Does the coronavirus die in snow?*

**The intents to merge need distinct answers**

**Use entity extraction to trigger the correct answer**

Let's say that two intents have a similar meaning and should be merged. However we want to keep two separate answers: for example, a single answer would be too long.

We can use a mechanism called *entity extraction*. An entity is a "thing of interest" in the user message: for example, it could be a date or time, a person's name, a location, etc.

It is possible to trigger a specific answer based on the intent detected, but also on the entities present in the user's message.

This is what happens when answering questions about a specific disease:

* *Can mosquitoes transmit corona?*
* *Are mosquitoes a vector of Ebola?*

The first message has the intent disease\_myth\_mosquitoes and the chatbot has found the disease entity with the value covid.

The second message also has the intent disease\_myth\_mosquitoes, but the chatbot found the disease entity with the value ebola instead.

The answers triggered are different:

| Intent predicted          | disease entity found | Answer triggered                |
| ------------------------- | -------------------- | ------------------------------- |
| disease\_myth\_mosquitoes | covid                | answer\_covid\_myth\_mosquitoes |
| disease\_myth\_mosquitoes | ebola                | answer\_ebola\_myth\_mosquitoes |

If the disease entity is missing, we can prompt the user for more information.

**Example scenario**

We can use the same idea to merge `disease_myth_mosquitoes` and `disease_myth_flies` together.

`disease_myth_flies`

* *Can flies transmit corona?*
* *Are flies a vector for Covid?*
* *Can a fly give me ebola*
* *Should I stay away from flies?*

`disease_myth_mosquitoes`

* Can mosquitoes transmit ebola?
* Is it possible to get the virus from a mosquito bite?
* Are mosquitoes capable of giving me COVID?
* When a mosquito bites me, can I get EBOLA

If we merge those intents together, we get:

`disease_myth_insects`

* *Can flies transmit corona?*
* *Are flies a vector for Covid?*
* *Can a fly give me ebola*
* *Should I stay away from flies?*
* *Can mosquitoes transmit ebola?*
* *Is it possible to get the virus from a mosquito bite?*
* *Are mosquitoes capable of giving me COVID?*
* *When a mosquito bites me, can I get EBOLA*

The entity annotation component will transform those training examples before building the model. **The annotation is done automatically, the training examples should not be annotated manually.**

This is what the training examples with all the annotations look like:

`disease_myth_insects`

* *Can `[flies]{"entity": "insect", "value": "fly"}` transmit `[corona]{"entity": "disease", "value": "covid"}`?*
* *Are `[flies]{"entity": "insect", "value": "fly"}` a vector for `[Covid]{"entity": "disease", "value": "covid"}`?*
* *Can a `[fly]{"entity": "insect", "value": "fly"}` give me `[ebola]{"entity": "disease", "value": "ebola"}`*
* *Should I stay away from `[flies]{"entity": "insect", "value": "fly"}`?*
* *Can `[mosquitoes]{"entity": "insect", "value": "mosquito"}` transmit `[ebola]{"entity": "disease", "value": "ebola"}`?*
* *Is it possible to get the virus from a `[mosquito]{"entity": "insect", "value": "mosquito"}` bite?*
* *Are `[mosquitoes]{"entity": "insect", "value": "mosquito"}` capable of giving me `[COVID]{"entity": "disease", "value": "covid"}`?*
* *When a `[mosquito]{"entity": "insect", "value": "mosquito"}` bites me, can I get `[EBOLA]{"entity": "disease", "value": "ebola"}`*

The intents are now merged, and the entities are annotated. After adding conditions to the stories, the combination of intent entities should trigger the desired answer.

<table><thead><tr><th width="240">Intent predicted</th><th width="128">disease entity found</th><th width="167">insect entity found</th><th>Answer triggered</th></tr></thead><tbody><tr><td><code>disease_myth_insects</code></td><td>covid</td><td>mosquito</td><td><code>answer_covid_myth_mosquitoes</code></td></tr><tr><td><code>disease_myth_insects</code></td><td>covid</td><td>fly</td><td><code>answer_covid_myth_flies</code></td></tr><tr><td><code>disease_myth_insects</code></td><td>ebola</td><td>mosquito</td><td><code>answer_ebola_myth_mosquitoes</code></td></tr><tr><td><code>disease_myth_insects</code></td><td>ebola</td><td>fly</td><td><code>answer_ebola_myth_flies</code></td></tr></tbody></table>

If an entity is missing, we can prompt the user for more information.

{% hint style="success" %}
**Entity extraction checklist**

To merge intents but retain distinct answers using entity extraction, you would need to create patterns to find the entities. When it was decided that two or more intents should be merged together, I need to know:

* Which entities will need to be created (for example, the entity insect or animal)
* Which values can be assigned to those entities (for example, fly, mosquito, pet)
* Which words are used to designate those values, **in each language**: ("corona", "covid", "c19", "coronavirus" are all valid ways of designating covid)
* Which answer should be triggered by a given combination of intent and entities

With this information, I will set up the entities and modify the stories accordingly.
{% endhint %}

{% hint style="success" %}
**Intent merging checklist**

* Are there intents with similar meaning? It might be possible to merge them.
* Are the answers to those intents similar? Merging can be done just by rewriting the answer and relabeling the training examples.
* Should distinct answers be kept? Merging can still be achieved, and distinct answers can be triggered using entity extraction. Check the previous section for details.
  {% endhint %}

### **Each intents should have around the same amount of training data**

**Unbalanced training data can cause class bias**

Sometimes, some intents can have a lot more training examples than others:

`disease_myth_fruits`

* Can bananas cure covid-19?
* Are bananas a cure for coronavirus?
* Are bananas a cure?
* Can I eat bananas to treat Corona?
* Can I eat bananas to protect me?
* Can I eat fruit to protect myself?
* Will eating lemons protect me?
* Are oranges good against Corona?
* Tea with orange helps against corona?
* Lemon tea protects me from the virus?

`disease_myth_spices`

* Can chili prevent covid?
* Should I eat more chili to cure corona?

When that is the case, the chatbot can get confused and too often select the intent with many training examples. This is known as *class bias*. This effect is likely to be stronger when intents are close together in meaning.

**Prioritize adding training examples for intents with little data**

The easiest way to address class bias is to add more training examples to "small" intents.

If `disease_stats` has more than 100 examples, and `disease_myth_spices` has just 10, then the priority is to increase the number of training examples in `disease_myth_spices`. Another option of course could be to merge the two together encompassing the two topics.

Keeping the intents size balanced is important, but not as much as increasing the number of training examples or merging similar intents.


# 7.3.8 Evaluation and continuous improvement of chatbots

Developing a successful chatbot involves more than just creating it – it’s an ongoing process that requires careful evaluation and continuous improvement. Understanding how effective your chatbot is in real-world scenarios is essential for enhancing its performance and user satisfaction. This iterative approach ensures that your chatbot evolves alongside user needs and changing conversation dynamics. By collecting user data and analyzing interactions, you can gain insights into how well your chatbot is meeting user expectations and where improvements are needed. Evaluating your chatbot’s performance not only helps identify areas for refinement but also allows you to make informed decisions on enhancing its capabilities.

<figure><img src="/files/lNEuvQuVuVm47ocSTgUp" alt=""><figcaption></figcaption></figure>

To illustrate this process, consider the diagram above with explanations of the steps below:

1. **Curated data**: This is the NLU training data that will be used in obtaining the first models. This data consists of intents, responses and stories which are manually prepared by linguists. It can be modified once the cycle is running with the insights we get from evaluations.
2. **NLU training data**: This is the NLU training data that grows at each cycle step with new labeled user input. Initially, it is equal to the curated data.
3. **Train**: This step is where the model is trained with the current NLU training data. This should be done on a development server. Rasa tests are performed right after this step.
4. **Deploy**: The newly trained model should be deployed on the production server if it performs at least as good as the previous model.
5. **Collect user data**: User input to the production bot is collected from the bot’s database. Information needed for the following steps is: `<user input, bot’s classification, confidence score>`. Also, the user metrics database should be updated in this step.
6. **Manual labeling**: The user data that was collected in the previous step is useful for two purposes: (1) Evaluate how the model is performing in real life, (2) Capture occasions where the model is least sure of its classifications and improve them. It is however needed first that the user input is manually classified by linguists. This is adding a new field to the data collected in the previous step: `<user input, bot’s classification, manual classification confidence-score>`\
   Since it is not possible to label all incoming user data, two subsets of sizes proportionate to the linguist capacity should be selected:
   1. A random set of user input, to be used for evaluation
   2. Set of classifications with low confidence scores (how sure the model is with its classification of user’s input).
   3. Optionally, user data can be collected for intents that are commonly mistaken, which can be detected in the confusion matrix, to strengthen up modeling of those specific intents.

At the end of this step, both sets should be fed back into the NLU training data. This will help improve the model for the next cycle.

Rasa Open Source offers a suite of powerful evaluation tools to facilitate this continuous improvement journey. These tools enable you to validate your data, test dialogues, assess the quality of your natural language understanding (NLU) model, and compare different pipeline configurations. By utilizing these evaluation functionalities, you can make data-driven decisions to fine-tune your chatbot’s performance and enhance its effectiveness over time.

Cross-validation testing is the best way to assess the maturity of your bot before rolling it out and collecting any user feedback. It measures [recall, precision & F1 scores](https://en.wikipedia.org/wiki/F1_score) directly from the curated training data you prepare.

End-to-end testing plays out pre-programmed conversations with the bot and reports if it goes as planned.

Confusion matrix gives information which intents (what user questions are “about”) are often confused with each other, thus needing more training data.

![](/files/NMQuzAnyiA2WkpTNhpW1)

For a detailed guide on how to leverage Rasa’s evaluation capabilities, please refer to the[ official Rasa documentation on testing your assistant](https://rasa.com/docs/rasa/2.x/testing-your-assistant). For higher-level strategies for evaluating language solutions please refer to [6.5 Key Metrics for Evaluating Language Solutions](/6.-language-technology-implementation/6.5-key-metrics-for-evaluating-language-solutions).


# 8 Sources and further bibliography

Uszkoreit, H. (1997). “[Language Technology A First Overview](https://www.dfki.de/~hansu/LT.pdf)”. Deutsches Forschungszentrum für Künstliche Intelligenz.

Ideo. (2015). “[The field guide to Human-Centered Design](https://www.designkit.org/resources/1.html)”. Journal design kit the human-centered design toolkit.

Mark, D. (Aug 13, 2020). “[Serverless ML: Deploying Lightweight Models at Scale](https://mark.douthwaite.io/serverless-machine-learning/)”. Mark Douthwaite.

Mark, D. (Jul 22, 2020). “[A Brief Introduction to Serverless Computing](https://mark.douthwaite.io/a-brief-introduction-to-serverless-computing/)”. Mark Douthwaite.

Jan, K. (2023). [Creating community-driven datasets: Insights from Mozilla Common Voice activities in East Africa.](https://assets.mofoprod.net/network/documents/Creating-Community-Driven-Datasets-Report-032023-GIZ-Mozilla.pdf) Deutsche Gesellschaft für Internationale Zusammenarbeit (GIZ) GmbH.&#x20;

Kalika Bali, Monojit Choudhury, Sunayana Sitaram, Vivek Seshadri. (December 2019) “[ELLORA: Enabling Low Resource Languages with Technology”.](https://www.microsoft.com/en-us/research/publication/ellora-enabling-low-resource-languages-with-technology/) UNESCO International Conference on Language Technologies for all (LT4All)

David Adelani, Jesujoba Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Emezue, Colin Leong, Michael Beukman, Shamsuddeen Muhammad, Guyo Jarso, Oreen Yousuf, et al.. 2022.[ A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation](https://aclanthology.org/2022.naacl-main.223). In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3053–3070, Seattle, United States. Association for Computational Linguistics.

Adelani, D.I., Masiak, M., Azime, I.A., Alabi, J.O., Tonja, A.L., Mwase, C., Ogundepo, O., Dossou, B.F., Oladipo, A., Nixdorf, D., Emezue, C.C., Al-Azzawi, S.S., Sibanda, B.K., David, D., Ndolela, L., Mukiibi, J., Ajayi, T.O., Ngoli, T.M., Odhiambo, B., Owodunni, A.T., Obiefuna, N.C., Muhammad, S.H., Abdullahi, S.S., Yigezu, M.G., Gwadabe, T.R., Abdulmumin, I., Bame, M.T., Awoyomi, O.O., Shode, I., Adelani, T.A., Kailani, H.A., Omotayo, A., Adeeko, A., Abeeb, A., Aremu, A., Samuel, O., Siro, C., Kimotho, W., Ogbu, O.R., Mbonu, C.E., Chukwuneke, C.I., Fanijo, S., Ojo, J., Awosan, O.F., Guge, T.K., Sari, S.T., Nyatsine, P., Sidume, F., Yousuf, O., Oduwole, M., Kimanuka, U., Tshinu, K.P., Diko, T., Nxakama, S., Johar, A.T., Gebre, S., Mohamed, M.A., Mohamed, S.A., Hassan, F.M., Mehamed, M.A., Ngabire, E., & Stenetorp, P. (2023). “[MasakhaNEWS: News Topic Classification for African languages”.](https://arxiv.org/abs/2304.09972) ArXiv, abs/2304.09972.

“[Language Specific Peculiarities Document for Sheng as Spoken in Kenya](https://gamayun.translatorswb.org/download/sheng-lsp/)”. (2022) Appen and CLEAR Global.

Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, et al.. 2022.[ “Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets](https://aclanthology.org/2022.tacl-1.4)”. Transactions of the Association for Computational Linguistics, 10:50–72.

Santini, M., Strandqvist, W., Nyström, M., Alirezai, M., Jönsson, A. (2018). “[Can We Quantify Domainhood? Exploring Measures to Assess Domain-Specificity in Web Corpora](https://doi.org/10.1007/978-3-319-99133-7_17)”. In: Elloumi, M., et al. Database and Expert Systems Applications. DEXA 2018. Communications in Computer and Information Science, vol 903. Springer, Cham.&#x20;

Öktem A. (2022) “[Toolkit for marginalized and under-resourced languages](https://language-toolkit.readthedocs.io/en/latest/)”.&#x20;

Öktem, A., DeLuca, E., Bashizi, R., Paquin, E., & Tang, G. (2021). [Congolese Swahili Machine Translation for Humanitarian Response](https://arxiv.org/abs/2103.10734). ArXiv, abs/2103.10734.

Anastasopoulos, A., Cattelan, A., Dou, Z., Federico, M., Federman, C., Genzel, D., Guzm'an, F., Hu, J., Hughes, M., Koehn, P., Lazar, R., Lewis, W., Neubig, G., Niu, M., Oktem, A., Paquin, E., Tang, G., & Tur, S. (2020). [TICO-19: the Translation Initiative for Covid-19](https://arxiv.org/abs/2007.01788). ArXiv, abs/2007.01788.

Nekoto, W., Marivate, V., Matsila, T., Fasubaa, T.E., Kolawole, T., Fagbohungbe, T.H., Akinola, S.O., Muhammad, S.H., KABENAMUALU, S.K., Osei, S., Freshia, S., Niyongabo Rubungo, A., Macharm, R., Ogayo, P., Ahia, O., Meressa, M., Adeyemi, M., Mokgesi-Selinga, M., Okegbemi, L., Martinus, L., Tajudeen, K., Degila, K., Ogueji, K., Siminyu, K., Kreutzer, J., Webster, J., Ali, J.T., Abbott, J.Z., Orife, I., Ezeani, I.U., Dangana, I.A., Kamper, H., ElSahar, H., Duru, G., Kioko, G., Murhabazi, E., Biljon, E.V., Whitenack, D., Onyefuluchi, C., Emezue, C.C., Dossou, B.F., Sibanda, B.K., Bassey, B.I., Olabiyi, A., Ramkilowan, A., Oktem, A., Akinfaderin, A., & Bashir, A.M. (2020). [Participatory Research for Low-resourced Machine Translation: A Case Study in African Languages.](https://arxiv.org/abs/2010.02353) ArXiv, abs/2010.02353

Alp Öktem, Muhannad Albayk Jaam, Eric DeLuca, Grace Tang. (2020) “[Gamayun –  Language Technology for Humanitarian Response”](https://gamayun.translatorswb.org/gamayun-ghtc/) In: 2020 IEEE Global Humanitarian Technology Conference (GHTC) 2020 October 29 - November 1; Virtual. [IEEEXplore link](https://ieeexplore.ieee.org/document/9342939)

∀, Iroro Orife, Julia Kreutzer, Blessing Sibanda, Daniel Whitenack, Kathleen Siminyu, Laura Martinus, Jamiil Toure Ali, Jade Abbott, Vukosi Marivate, Salomon Kabongo, Musie Meressa, Espoir Murhabazi, Orevaoghene Ahia, Elan van Biljon, Arshath Ramkilowan, Adewale Akinfaderin, Alp Öktem, Wole Akin, Ghollah Kioko, Kevin Degila, Herman Kamper, Bonaventure Dossou, Chris Emezue, Kelechi Ogueji, Abdallah Bashir. “[Masakhane -- Machine Translation For Africa](https://arxiv.org/abs/2003.11529)” In: AfricaNLP workshop @ Eighth International Conference on Learning Representations (ICLR 2020) 2020 April 26; Addis Ababa, Ethiopia (Online).

Alp Öktem, Mirko Plitt, Grace Tang. “[Tigrinya Neural Machine Translation with Transfer Learning for Humanitarian Response](https://arxiv.org/abs/2003.11523)”. In: AfricaNLP workshop @ Eighth International Conference on Learning Representations (ICLR 2020) 2020 April 26; Addis Ababa, Ethiopia (Online).

Aimee Ansari (2023) “[Hello, Porcupine! Using AI to support farmers to adapt to climate change](https://clearglobal.org/using-ai-to-support-farmers-to-adapt-to-climate-change/)”. CLEAR Global Blog

Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and André F. T. Martins. 2022.[ Results of WMT22 Metrics Shared Task: Stop Using BLEU – Neural Metrics Are Better and More Robust](https://aclanthology.org/2022.wmt-1.2). In Proceedings of the Seventh Conference on Machine Translation (WMT), pages 46–68, Abu Dhabi, United Arab Emirates (Hybrid). Association for Computational Linguistics.

Kishore Papineni, Salim Roukos, Todd Ward, and WeiJing Zhu. 2002. [Bleu: a method for automatic evaluation of machine translation. ](https://aclanthology.org/P02-1040/)In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318, Philadelphia, Pennsylvania, USA. Association for Computational Linguistics.

Maja Popovic. 2015. [chrF: character n-gram F-score for automatic MT evaluation](https://aclanthology.org/W15-3049/). In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392–395, Lisbon, Portugal. Association for Computational Linguistics.

Way, A. (2018). [Quality Expectations of Machine Translation.](https://arxiv.org/abs/1803.08409) In: Moorkens, J., Castilho, S., Gaspari, F., Doherty, S. (eds) Translation Quality Assessment. Machine Translation: Technologies and Applications, vol 1. Springer, Cham. <https://doi.org/10.1007/978-3-319-91241-7_8>

Müller, Mathias. [Robust Neural Machine Translation Systems.](https://www.zora.uzh.ch/id/eprint/208945/) 2021, University of Zurich, Faculty of Arts.

Bahdanau, D., Cho, K., & Bengio, Y. (2014). [Neural Machine Translation by Jointly Learning to Align and Translate](https://arxiv.org/abs/1409.0473). CoRR, abs/1409.0473.

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. [Attention is All you Need.](https://research.google/pubs/pub46201/) In Advances in Neural Information Processing Systems 30, pages 5998–6008.


