Featurization of Conversations
These docs are for version 1.x of Rasa Open Source.
User Guide
- Installation
- Tutorial: Rasa Basics
- Tutorial: Building Assistants
- Command Line Interface
- Architecture
- Messaging and Voice Channels
- Evaluating Models
- Validate Data
- Configuring the HTTP API
- Deploying your Rasa Assistant
- Cloud Storage
NLU
- About
- Using NLU Only
- Training Data Format
- Choosing a Pipeline
- Language Support
- Entity Extraction
- Components
Core
- About
- Stories
- Domains
- Responses
- Actions
- Reminders and External Events
- Policies
- Slots
- Forms
- Retrieval Actions
- Interactive Learning
- Fallback Actions
- Knowledge Base Actions
Conversation Design
API Reference
- Action Server
- HTTP API
- Jupyter Notebooks
- Agent
- Custom NLU Components
- Rasa SDK
- Events
- Tracker
- Tracker Stores
- Event Brokers
- Lock Stores
- Training Data Importers
- Featurization of Conversations
- TensorFlow Configuration
- Migration Guide
- Rasa Open Source Change Log
Migrate from (beta)
Reference
Versions
viewing: 1.8.1
Featurization of Conversations
In order to apply machine learning algorithms to conversational AI, we need to build up vector representations of conversations.
Each story corresponds to a tracker which consists of the states of the conversation just before each action was taken.
State Featurizers
Every event in a trackers history creates a new state (e.g. running a bot action, receiving a user message, setting slots). Featurizing a single state of the tracker has a couple steps:
- Tracker provides a bag of active features:
features indicating intents and entities, if this is the first state in a turn (e.g.
[intent_restaurant_search, entity_cuisine])features indicating which slots are currently defined (e.g.
slot_location)features indicating the results of any API calls stored in slots (e.g.
slot_matches)features indicating what the last action was (e.g.
prev_action_listen)
- Convert all the features into numeric vectors:
We use the
X, ynotation that’s common for supervised learning, whereXis an array of shape(num_data_points, time_dimension, num_input_features), andyis an array of shape(num_data_points, num_bot_features)where the target class labels are encoded as one-hot vectors.
The target labels correspond to actions taken by the bot. There are different featurizers available:
BinarySingleStateFeaturizercreates a binary one-hot encoding.LabelTokenizerSingleStateFeaturizercreates a vector based on the feature label.
Tracker Featurizers
It’s often useful to include a bit more history than just the current state when predicting an action. The TrackerFeaturizer iterates over tracker states and calls a SingleStateFeaturizer for each state. There are two different tracker featurizers:
1. Full Dialogue
FullDialogueTrackerFeaturizer creates numerical representation of stories. Therefore, X is an array of shape (num_stories, max_dialogue_length, num_input_features) and y is of shape (num_stories, max_dialogue_length, num_bot_features).
2. Max History
MaxHistoryTrackerFeaturizer creates an array of previous tracker states for each bot action with the parameter max_history defining how many states go into each row in X.