Choosing a Pipeline
Choosing a Pipeline
In Rasa Open Source, incoming messages are processed by a sequence of components. These components are executed one after another in a so-called processing pipeline defined in your config.yml. Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.
How to Choose a Pipeline
The Short Answer
If your training data is in English, a good starting point is the following pipeline:
language: "en"
pipeline:
- name: ConveRTTokenizer
- name: ConveRTFeaturizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: "char_wb"
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
epochs: 100
- name: EntitySynonymMapper
- name: ResponseSelector
epochs: 100
If your training data is not in English, start with the following pipeline:
language: "fr" # your two-letter language code
pipeline:
- name: WhitespaceTokenizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: "char_wb"
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
epochs: 100
- name: EntitySynonymMapper
- name: ResponseSelector
epochs: 100
A Longer Answer
We recommend using the following pipeline if your training data is in English:
language: "en"
The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.10.25/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance. Pre-trained word embeddings are helpful as they already encode some kind of linguistic knowledge.
If your training data is not in English but you still want to use pre-trained word embeddings, we recommend using the following pipeline:
language: "fr" # your two-letter language code
pipeline:
- name: SpacyNLP
- name: SpacyTokenizer
- name: SpacyFeaturizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer analyzer: "char_wb" min_ngram: 1 max_ngram: 4
- name: DIETClassifier epochs: 100
- name: EntitySynonymMapper
- name: ResponseSelector epochs: 100
If you don’t use any pre-trained word embeddings inside your pipeline, you are not bound to a specific language and can train your model to be more domain specific.
### Choosing the Right Components
There are components for entity extraction, for intent classification, response selection, pre-processing, and others. You can learn more about any specific component on the [Components](https://legacy-docs-v1.rasa.com/1.10.25/nlu/components/#components) page.
A pipeline usually consists of three main parts:
- [Tokenization](https://legacy-docs-v1.rasa.com/1.10.25/nlu/choosing-a-pipeline/#tokenization)
- [Featurization](https://legacy-docs-v1.rasa.com/1.10.25/nlu/choosing-a-pipeline/#featurization)
- [Entity Recognition / Intent Classification / Response Selectors](https://legacy-docs-v1.rasa.com/1.10.25/nlu/choosing-a-pipeline/#entity-recognition-intent-classification-response-selectors)
#### Tokenization
For tokenization of English input, we recommend the [ConveRTTokenizer](https://legacy-docs-v1.rasa.com/1.10.25/nlu/components/#converttokenizer).
#### Featurization
You need to decide whether to use components that provide pre-trained word embeddings or not.
##### Pre-trained Embeddings
The advantage of using pre-trained word embeddings in your pipeline is that if you have a training example like: "I want to buy apples", and Rasa is asked to predict the intent for "get pears", your model already knows that the words "apples" and "pears" are very similar.
- [MitieFeaturizer](https://legacy-docs-v1.rasa.com/1.10.25/nlu/components/#mitiefeaturizer)
- [SpacyFeaturizer](https://legacy-docs-v1.rasa.com/1.10.25/nlu/components/#spacyfeaturizer)
- [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.10.25/nlu/components/#convertfeaturizer)
- [LanguageModelFeaturizer](https://legacy-docs-v1.rasa.com/1.10.25/nlu/components/#languagemodelfeaturizer)
### Entity Recognition / Intent Classification / Response Selectors
Depending on your data you may want to only perform intent classification, entity recognition or response selection. Or you might want to combine multiple of those tasks.