Choosing a Pipeline

Choosing a Pipeline

In Rasa Open Source, incoming messages are processed by a sequence of components. These components are executed one after another in a so-called processing pipeline defined in your config.yml. Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.

How to Choose a Pipeline

The Short Answer

If your training data is in English, a good starting point is the following pipeline:

language: "en"

pipeline:
  - name: ConveRTTokenizer
  - name: ConveRTFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
    epochs: 100
  - name: EntitySynonymMapper
  - name: ResponseSelector
    epochs: 100

If your training data is not in English, start with the following pipeline:

language: "fr"  # your two-letter language code

pipeline:
  - name: WhitespaceTokenizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
    epochs: 100
  - name: EntitySynonymMapper
  - name: ResponseSelector
    epochs: 100

A Longer Answer

We recommend using the following pipeline, if your training data is in English:

language: "en"

The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.10.10/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance. Pre-trained word embeddings are helpful as they already encode some kind of linguistic knowledge.

#### Choosing the Right Components
There are components for entity extraction, for intent classification, response selection, pre-processing, and others. A pipeline usually consists of three main parts:
- **Tokenization**
- **Featurization**
- **Entity Recognition / Intent Classification / Response Selectors**

### Multi-Intent Classification
You can use Rasa Open Source components to split intents into multiple labels. To do this, use the [DIETClassifier](https://legacy-docs-v1.rasa.com/1.10.10/nlu/components/#diet-classifier) in your pipeline.

### Comparing Pipelines
Rasa gives you the tools to compare the performance of multiple pipelines on your data directly. See [Comparing NLU Pipelines](https://legacy-docs-v1.rasa.com/1.10.10/user-guide/testing-your-assistant/#comparing-nlu-pipelines) for more information.

### Handling Class Imbalance
Classification algorithms often do not perform well if there is a large class imbalance. To mitigate this problem, you can use a `balanced` batching strategy. This algorithm ensures that all classes are represented in every batch, or at least in as many subsequent batches as possible.

### Component Lifecycle
Each component processes an input and/or creates an output. The order of the components is determined by the order they are listed in the `config.yml`; the output of a component can be used by any other component that comes after it in the pipeline.