Choosing a Pipeline
Choosing a Pipeline
Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.
The Short Answer
If your training data is in English, a good starting point is the following pipeline:
language: "en"
pipeline:
- name: ConveRTTokenizer
- name: ConveRTFeaturizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: "char_wb"
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
- name: EntitySynonymMapper
- name: ResponseSelector
In case your training data is in a different language than English, use the following pipeline:
language: "en"
pipeline:
- name: WhitespaceTokenizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: "char_wb"
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
- name: EntitySynonymMapper
- name: ResponseSelector
A Longer Answer
We recommend using following pipeline, if your training data is in English:
language: "en"
The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.8.1/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance.
## Choosing the right Components
A pipeline usually consists of three main parts:
1. Tokenization
2. Featurization
3. Entity Recognition / Intent Classification / Response Selectors
### Tokenization
If your chosen language is whitespace-tokenized (words are separated by spaces), you can use the [WhitespaceTokenizer](https://legacy-docs-v1.rasa.com/1.8.1/nlu/components/#whitespacetokenizer). If this is not the case you should use a different tokenizer.
### Featurization
You need to decide whether to use components that provide pre-trained word embeddings or not.
### Entity Recognition / Intent Classification / Response Selectors
Depending on your data you may want to only perform intent classification, entity recognition or response selection. Or you might want to combine multiple of those tasks. We support several components for each of the tasks.
## Class imbalance
Classification algorithms often do not perform well if there is a large class imbalance, for example if you have a lot of training data for some intents and very little training data for others.
## Multiple Intents
If you want to split intents into multiple labels, e.g. for predicting multiple intents or for modeling hierarchical intent structure, you need to use the [DIETClassifier](https://legacy-docs-v1.rasa.com/1.8.1/nlu/components/#diet-classifier) in your pipeline.
### Understanding the Rasa NLU Pipeline
In Rasa NLU, incoming messages are processed by a sequence of components.
Every component can implement several methods from the `Component` base class; in a pipeline these different methods will be called in a specific order.