Choosing a Pipeline

These docs are for version 1.x of Rasa Open Source.

Choosing a Pipeline

In Rasa Open Source, incoming messages are processed by a sequence of components. These components are executed one after another in a so-called processing pipeline defined in your config.yml. Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.

How to Choose a Pipeline

The Short Answer

If your training data is in English, a good starting point is the following pipeline:

language: "en"

pipeline:
  - name: ConveRTTokenizer
  - name: ConveRTFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
    epochs: 100
  - name: EntitySynonymMapper
  - name: ResponseSelector
    epochs: 100

If your training data is not in English, start with the following pipeline:

language: "fr"  # your two-letter language code

pipeline:
  - name: WhitespaceTokenizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
    epochs: 100
  - name: EntitySynonymMapper
  - name: ResponseSelector
    epochs: 100

A Longer Answer

We recommend using the following pipeline if your training data is in English:

language: "en"

The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.10.21/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance. Pre-trained word embeddings are helpful as they already encode some kind of linguistic knowledge. The advantage of the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.10.21/nlu/components/#convertfeaturizer) is that it doesn’t treat each word of the user message independently, but creates a contextual vector representation for the complete sentence.

If your training data is not in English, you can also use a different variant of a language model which is pre-trained in the language specific to your training data.

## Choosing the Right Components

There are components for entity extraction, for intent classification, response selection, pre-processing, and others. A pipeline usually consists of three main parts:

- Tokenization
- Featurization
- Entity Recognition / Intent Classification / Response Selectors

### Tokenization

For tokenization of English input, we recommend the [ConveRTTokenizer](https://legacy-docs-v1.rasa.com/1.10.21/nlu/components/#converttokenizer).

### Featurization

You need to decide whether to use components that provide pre-trained word embeddings or not. We recommend in cases of small amounts of training data to start with pre-trained word embeddings.

#### Pre-trained Embeddings

The advantage of using pre-trained word embeddings in your pipeline is that if you have a training example like: “I want to buy apples”, your model already knows that the words “apples” and “pears” are very similar. We support a few components that provide pre-trained word embeddings:

1. [MitieFeaturizer](https://legacy-docs-v1.rasa.com/1.10.21/nlu/components/#mitiefeaturizer)
2. [SpacyFeaturizer](https://legacy-docs-v1.rasa.com/1.10.21/nlu/components/#spacyfeaturizer)
3. [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.10.21/nlu/components/#convertfeaturizer)
4. [LanguageModelFeaturizer](https://legacy-docs-v1.rasa.com/1.10.21/nlu/components/#languagemodelfeaturizer)

### Entity Recognition / Intent Classification / Response Selectors

Depending on your data, you may want to only perform intent classification, entity recognition, or response selection. We recommend using [DIETClassifier](https://legacy-docs-v1.rasa.com/1.10.21/nlu/components/#diet-classifier) for intent classification and entity recognition and [ResponseSelector](https://legacy-docs-v1.rasa.com/1.10.21/nlu/components/#response-selector) for response selection.