Choosing a Pipeline

Choosing a Pipeline

Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.

The Short Answer

If your training data is in English, a good starting point is the following pipeline:

language: "en"

pipeline:
  - name: ConveRTTokenizer
  - name: ConveRTFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector

In case your training data is in a different language than English, use the following pipeline:

language: "en"

pipeline:
  - name: WhitespaceTokenizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector

A Longer Answer

We recommend using the following pipeline if your training data is in English:

language: "en"

The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.8.2/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance. Pre-trained word embeddings are helpful as they already encode some kind of linguistic knowledge.

If your training data is not in English, but you still want to use pre-trained word embeddings, we recommend using the following pipeline:

```yaml
language: "en"

pipeline:
  - name: SpacyNLP
  - name: SpacyTokenizer
  - name: SpacyFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector

Choosing the right Components

A pipeline usually consists of three main parts:

  1. Tokenization
  2. Featurization
  3. Entity Recognition / Intent Classification / Response Selectors

Tokenization

If your chosen language is whitespace-tokenized (words are separated by spaces), you can use the WhitespaceTokenizer. If this is not the case, you should use a different tokenizer. We support a number of different tokenizers, or you can create your own custom tokenizer.

Featurization

You need to decide whether to use components that provide pre-trained word embeddings or not. If you don’t use any pre-trained word embeddings inside your pipeline, you are not bound to a specific language and can train your model to be more domain specific.

If your training data is in English, we recommend using the ConveRTFeaturizer.

Entity Recognition / Intent Classification / Response Selectors

Depending on your data, you may want to only perform intent classification, entity recognition, or response selection. We recommend using DIETClassifier for intent classification and entity recognition and ResponseSelector for response selection.

Class Imbalance

Classification algorithms often do not perform well if there is a large class imbalance, for example, if you have a lot of training data for some intents and very little training data for others.

To mitigate this problem, you can use a balanced batching strategy.

language: "en"

pipeline:
# - ... other components
- name: "DIETClassifier"
  batch_strategy: sequence

Multiple Intents

If you want to split intents into multiple labels, you need to use the DIETClassifier in your pipeline.

Understanding the Rasa NLU Pipeline

In Rasa NLU, incoming messages are processed by a sequence of components. There are components for entity extraction, for intent classification, response selection, pre-processing, and others. Each component processes the input and creates an output.

Component Lifecycle

Every component can implement several methods from the Component base class; in a pipeline, these methods will be called in a specific order. The components are executed one after another in a so-called processing pipeline.