Choosing a Pipeline

Choosing a Pipeline

In Rasa Open Source, incoming messages are processed by a sequence of components. These components are executed one after another in a so-called processing pipeline defined in your config.yml. Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.

How to Choose a Pipeline

The Short Answer

If your training data is in English, a good starting point is the following pipeline:

language: "en"

pipeline:
  - name: ConveRTTokenizer
  - name: ConveRTFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
    epochs: 100
  - name: EntitySynonymMapper
  - name: ResponseSelector
    epochs: 100

If your training data is not in English, start with the following pipeline:

language: "fr"  # your two-letter language code

pipeline:
  - name: WhitespaceTokenizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
    epochs: 100
  - name: EntitySynonymMapper
  - name: ResponseSelector
    epochs: 100

A Longer Answer

We recommend using following pipeline, if your training data is in English:

language: "en"

The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.9.2/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance.

### [Choosing the Right Components](https://legacy-docs-v1.rasa.com/1.9.2/nlu/choosing-a-pipeline/#id9)

There are components for entity extraction, for intent classification, response selection,
pre-processing, and others. You can learn more about any specific component on the [Components](https://legacy-docs-v1.rasa.com/1.9.2/nlu/components/#components) page.

A pipeline usually consists of three main parts:

- [Tokenization](https://legacy-docs-v1.rasa.com/1.9.2/nlu/choosing-a-pipeline/#tokenization)

- [Featurization](https://legacy-docs-v1.rasa.com/1.9.2/nlu/choosing-a-pipeline/#featurization)

- [Entity Recognition / Intent Classification / Response Selectors](https://legacy-docs-v1.rasa.com/1.9.2/nlu/choosing-a-pipeline/#entity-recognition-intent-classification-response-selectors)

## [Tokenization](https://legacy-docs-v1.rasa.com/1.9.2/nlu/choosing-a-pipeline/#id15)

For tokenization of English input, we recommend the [ConveRTTokenizer](https://legacy-docs-v1.rasa.com/1.9.2/nlu/components/#converttokenizer).

## [Featurization](https://legacy-docs-v1.rasa.com/1.9.2/nlu/choosing-a-pipeline/#id16)

You need to decide whether to use components that provide pre-trained word embeddings or not.

## [Entity Recognition / Intent Classification / Response Selectors](https://legacy-docs-v1.rasa.com/1.9.2/nlu/choosing-a-pipeline/#id17)

Depending on your data you may want to only perform intent classification, entity recognition or response selection.

### [Multi-Intent Classification](https://legacy-docs-v1.rasa.com/1.9.2/nlu/choosing-a-pipeline/#id10)

You can use Rasa Open Source components to split intents into multiple labels.

## [Comparing Pipelines](https://legacy-docs-v1.rasa.com/1.9.2/nlu/choosing-a-pipeline/#id11)

Rasa gives you the tools to compare the performance of multiple pipelines on your data directly.

## [Handling Class Imbalance](https://legacy-docs-v1.rasa.com/1.9.2/nlu/choosing-a-pipeline/#id12)

To mitigate this problem, you can use a `balanced` batching strategy.

## [Component Lifecycle](https://legacy-docs-v1.rasa.com/1.9.2/nlu/choosing-a-pipeline/#id13)

Each component processes an input and/or creates an output.

## [Pipeline Templates (deprecated)](https://legacy-docs-v1.rasa.com/1.9.2/nlu/choosing-a-pipeline/#id14)

A template is just a shortcut for a full list of components.