Choosing a Pipeline
Choosing a Pipeline
In Rasa Open Source, incoming messages are processed by a sequence of components. These components are executed one after another in a so-called processing pipeline defined in your config.yml. Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.
How to Choose a Pipeline
The Short Answer
If your training data is in English, a good starting point is the following pipeline:
language: "en"
pipeline:
- name: ConveRTTokenizer
- name: ConveRTFeaturizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: "char_wb"
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
epochs: 100
- name: EntitySynonymMapper
- name: ResponseSelector
epochs: 100
If your training data is not in English, start with the following pipeline:
language: "fr" # your two-letter language code
pipeline:
- name: WhitespaceTokenizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: "char_wb"
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
epochs: 100
- name: EntitySynonymMapper
- name: ResponseSelector
epochs: 100
A Longer Answer
We recommend using the following pipeline if your training data is in English:
language: "en"
### [Choosing the Right Components](https://legacy-docs-v1.rasa.com/1.10.24/nlu/choosing-a-pipeline/#id9)
There are components for entity extraction, intent classification, response selection, pre-processing, and others. Each component's specific details can be found on the [Components](https://legacy-docs-v1.rasa.com/1.10.24/nlu/components/#components) page.
A pipeline usually consists of three main parts:
- [Tokenization](https://legacy-docs-v1.rasa.com/1.10.24/nlu/choosing-a-pipeline/#tokenization)
- [Featurization](https://legacy-docs-v1.rasa.com/1.10.24/nlu/choosing-a-pipeline/#featurization)
- [Entity Recognition / Intent Classification / Response Selectors](https://legacy-docs-v1.rasa.com/1.10.24/nlu/choosing-a-pipeline/#entity-recognition-intent-classification-response-selectors)
#### [Tokenization](https://legacy-docs-v1.rasa.com/1.10.24/nlu/choosing-a-pipeline/#id15)
For tokenization of English input, we recommend the [ConveRTTokenizer](https://legacy-docs-v1.rasa.com/1.10.24/nlu/components/#converttokenizer).
#### [Featurization](https://legacy-docs-v1.rasa.com/1.10.24/nlu/choosing-a-pipeline/#id16)
You need to decide whether to use components that provide pre-trained word embeddings or not. We recommend starting with pre-trained embeddings for small datasets.
#### [Entity Recognition / Intent Classification / Response Selectors](https://legacy-docs-v1.rasa.com/1.10.24/nlu/choosing-a-pipeline/#id17)
Recommended components for intent classification and entity recognition include [DIETClassifier](https://legacy-docs-v1.rasa.com/1.10.24/nlu/components/#diet-classifier) and [ResponseSelector](https://legacy-docs-v1.rasa.com/1.10.24/nlu/components/#response-selector).
### [Multi-Intent Classification](https://legacy-docs-v1.rasa.com/1.10.24/nlu/choosing-a-pipeline/#id10)
You can use Rasa Open Source components to split intents into multiple labels, for example:
language: "en"
pipeline:
- name: "WhitespaceTokenizer" intent_tokenization_flag: True intent_split_symbol: "_"
- name: "CountVectorsFeaturizer"
- name: "DIETClassifier"
### [Comparing Pipelines](https://legacy-docs-v1.rasa.com/1.10.24/nlu/choosing-a-pipeline/#id11)
Rasa provides tools to compare performance of multiple pipelines on your data.
### [Handling Class Imbalance](https://legacy-docs-v1.rasa.com/1.10.24/nlu/choosing-a-pipeline/#id12)
To address large class imbalance, you can use a `balanced` batching strategy which ensures all classes are represented in batches. This is the default setting.