How to improve your task-specific chatbot for better safety, relevancy, and user feedback

How to improve your task-specific chatbot for better safety, relevancy, and user feedback

Building task-specific chatbots requires a structured approach when it comes to improving their everyday usefulness, specifically for better safety, relevancy, and user feedback. Given the highly subjective tasks LLM-powered chat applications are expected to perform, a common denominator for how well they do in real-world settings depend on the availability of reliable high-quality training data and how closely aligned they are to human preferences. Working hand-in-hand with leading AI teams, we've observed a set of best practices that we wanted to share in order to help you improve the performance of your task-specific chatbots.

In this tutorial guide, we'll walk through some of these top considerations and how Labelbox can be used as a platform to help accelerate chatbot development.

Part 1: Trust and Safety — Understanding Intentions

To ensure the best user experience, an LLM-based chatbot must be properly scoped to deliver on its intended area of expertise. For example, a chatbot application for an airline company should not be responding to off topic questions, such as politics. Therefore, a chatbot that understands user intent can steer the user towards its intended areas of expertise and away from potentially harmful or unrelated conversations.

In this section, we'll leverage Labelbox to classify the intent of historical conversations as on-topic (coffee / tea) or off-topic (politics). To start off, let’s load a subset of the Ultrachat dataset to Labelbox Catalog. Ultrachat is an open-source dialogue dataset powered by Turbo APIs to train powerful language models with general conversational capability.

To being, let's first identify political conversations that your chatbot shouldn’t be able to answer.

You can see that using semantic search and labeling functions have produced promising initial results. As a next step, let’s validate your results further with foundation models.

To ensure completeness, you can next leverage a human-in-the-loop (HITL) approach to conduct a final review for intent.

Part 2: Generating Quality Responses for your LLM-based Chatbot

When your intelligent chat application is powered by an underlying large language model, you can customize these LLMs to a defined task with 2 key approaches: Fine-tuning or Retrieval Augmented Generation (RAG).

To start off, let's first generate quality responses to selected prompts. Ideally, we’ll want the chatbot to replicate the responses provided by your annotators.

Your team of annotators can now produce specific responses for the LLM to learn from.

By identifying relevant prompts and generating quality responses, you can now train a task specific LLM model and output predictions for any new requests.

Part 3: Model Evaluation and Deployment

Before deploying your model into production, let’s evaluate the performance of the responses when compared to your ground-truth data (expected output). Using holdout prompts and responses pairs not used to train the model, you can evaluate the performance of your fine-tuned LLM.

Afterwards, you can create a model run to add the holdout dataset containing the ground-truth responses.

Using the Python SDK, you can next export the prompts to generate responses from your fine-tuned model from Google's Vertex AI.

With the metrics and predictions uploaded, you can filter by various metrics to identify the highest and lowest performing prompts, along with overall model performance.

Conclusion

In this guide, we highlighted a few best practices for ensuring better safety, relevancy, and user feedback when building a task-specific chatbot. Regardless of the techniques used and the models chosen, it is crucial to generate and classify data to the highest quality for optimizing chatbot outputs.

By incorporating semantic search, labeling functions, foundation models, and a human-in-the-loop approach, you will be able to generate and classify data in higher qualities consistently. Combined with an SDK-driven approach, you can more easily train models and enhance LLM performance through faster iterations. Give the tutorial a try and we'd love to hear your feedback or ideas on how we can help you improve your LLM-based chatbot applications.