Choosing a Pipeline

Choosing a Pipeline

In Rasa Open Source, incoming messages are processed by a sequence of components. These components are executed one after another in a so-called processing pipeline defined in your config.yml. Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.

Tokenization

For tokenization of English input, we recommend the ConveRTTokenizer.

Featurization

You need to decide whether to use components that provide pre-trained word embeddings or not. We recommend using them in cases of small amounts of training data. Once you have larger amounts of data, supervised embeddings can make your model more specific to your domain.

Multi-Intent Classification

You can use Rasa Open Source components to split intents into multiple labels by using the DIETClassifier in your pipeline.

Comparing Pipelines

Rasa gives you the tools to compare the performance of multiple pipelines directly. See Comparing NLU Pipelines for more information.

Handling Class Imbalance

Classification algorithms often do not perform well with a large class imbalance. To mitigate this, you can use a balanced batching strategy which ensures that all classes are represented in every batch.