Choosing a Pipeline
These docs are for version 1.x of Rasa Open Source.
Choosing a Pipeline
In Rasa Open Source, incoming messages are processed by a sequence of components. These components are executed one after another in a so-called processing pipeline defined in your config.yml. Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.
How to Choose a Pipeline
The Short Answer
If your training data is in English, a good starting point is the following pipeline:
language: "en"
pipeline:
- name: ConveRTTokenizer
- name: ConveRTFeaturizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: "char_wb"
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
epochs: 100
- name: EntitySynonymMapper
- name: ResponseSelector
epochs: 100
If your training data is not in English, start with the following pipeline:
language: "fr" # your two-letter language code
pipeline:
- name: WhitespaceTokenizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: "char_wb"
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
epochs: 100
- name: EntitySynonymMapper
- name: ResponseSelector
epochs: 100
A Longer Answer
We recommend using the following pipeline, if your training data is in English:
language: "en"
The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.9.7/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance. Pre-trained word embeddings are helpful as they already encode some kind of linguistic knowledge.
An alternative to [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.9.7/nlu/components/#convertfeaturizer) is the [LanguageModelFeaturizer](https://legacy-docs-v1.rasa.com/1.9.7/nlu/components/#languagemodelfeaturizer) which uses pre-trained language models such as BERT, GPT-2, etc. to extract similar contextual vector representations for the complete sentence.
### Choosing the Right Components
There are components for entity extraction, for intent classification, response selection, pre-processing, and others. You can learn more about any specific component on the [Components](https://legacy-docs-v1.rasa.com/1.9.7/nlu/components/#components) page.
A pipeline usually consists of three main parts:
- [Tokenization](https://legacy-docs-v1.rasa.com/1.9.7/nlu/choosing-a-pipeline/#tokenization)
- [Featurization](https://legacy-docs-v1.rasa.com/1.9.7/nlu/choosing-a-pipeline/#featurization)
- [Entity Recognition / Intent Classification / Response Selectors](https://legacy-docs-v1.rasa.com/1.9.7/nlu/choosing-a-pipeline/#entity-recognition-intent-classification-response-selectors)
### Multi-Intent Classification
You can use Rasa Open Source components to split intents into multiple labels. For example, you can predict multiple intents (`thank+goodbye`) or model hierarchical intent structure (`feedback+positive` being more similar to `feedback+negative` than `chitchat`). To do this, you need to use the [DIETClassifier](https://legacy-docs-v1.rasa.com/1.9.7/nlu/components/#diet-classifier) in your pipeline.
### Comparing Pipelines
Rasa gives you the tools to compare the performance of multiple pipelines on your data directly. See [Comparing NLU Pipelines](https://legacy-docs-v1.rasa.com/1.9.7/user-guide/testing-your-assistant/#comparing-nlu-pipelines) for more information.
### Handling Class Imbalance
Classification algorithms often do not perform well if there is a large class imbalance, for example if you have a lot of training data for some intents and very little training data for others. To mitigate this problem, you can use a `balanced` batching strategy.
### Component Lifecycle
Each component processes an input and/or creates an output. The order of the components is determined by the order they are listed in the `config.yml`; the output of a component can be used by any other component that comes after it in the pipeline.