Choosing a Pipeline
Choosing a Pipeline
Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.
The Short Answer
If your training data is in English, a good starting point is the following pipeline:
language: "en"
pipeline:
- name: ConveRTTokenizer
- name: ConveRTFeaturizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: "char_wb"
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
- name: EntitySynonymMapper
- name: ResponseSelector
In case your training data is in a different language than English, use the following pipeline:
language: "en"
pipeline:
- name: WhitespaceTokenizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: "char_wb"
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
- name: EntitySynonymMapper
- name: ResponseSelector
A Longer Answer
We recommend using the following pipeline if your training data is in English:
language: "en"
The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.8.0/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance. Pre-trained word embeddings are helpful as they encode some linguistic knowledge.
If your training data is not in English but you still want to use pre-trained word embeddings, we recommend using the following pipeline:
language: "en"
pipeline:
- name: SpacyNLP
- name: SpacyTokenizer
- name: SpacyFeaturizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer analyzer: "char_wb" min_ngram: 1 max_ngram: 4
- name: DIETClassifier
- name: EntitySynonymMapper
- name: ResponseSelector
If you don’t use any pre-trained word embeddings, you can train your model to be more domain specific. If your domain-specific terminology is not captured by any word embeddings, we recommend:
language: "en"
Choosing the right Components
A pipeline usually consists of three main parts:
- Tokenization
- Featurization
- Entity Recognition / Intent Classification / Response Selectors
Tokenization
If your chosen language is whitespace-tokenized, use the WhitespaceTokenizer. If this is not the case, use a different tokenizer. We support a number of different tokenizers.
Featurization
Decide whether to use components that provide pre-trained word embeddings. If you don’t use them, train your model to be more domain-specific.
Entity Recognition / Intent Classification / Response Selectors
Depending on your data, you may want to perform entity recognition, intent classification, or response selection.
For performance comparisons, Comparing NLU Pipelines provides tools to evaluate different configurations.
Class Imbalance
To address class imbalances, use a balanced batching strategy to ensure class representation.
language: "en"
pipeline:
# - ... other components
- name: "DIETClassifier"
batch_strategy: sequence
Multiple Intents
To predict multiple intents, employ the DIETClassifier and specify intent handling flags.
Understanding the Rasa NLU Pipeline
In Rasa NLU, messages are processed by a sequence of components, creating outputs used by subsequent components.
{
"text": "I am looking for Chinese food",
"entities": [
{
"start": 8,
"end": 15,
"value": "chinese",
"entity": "cuisine",
"extractor": "DIETClassifier",
"confidence": 0.864
}
],
"intent": {
"confidence": 0.6485910906220309,
"name": "restaurant_search"
},
"intent_ranking": [
{"confidence": 0.6485910906220309, "name": "restaurant_search"},
{"confidence": 0.1416153159565678, "name": "affirm"}
]
}
Component Lifecycle
Components implement methods from the Component base class, which are executed in a specific order during the training of a pipeline.
The “entity” object explained
Entities are returned as a dictionary, providing insight into the extractor that identified them and any processes they underwent.
{
"text": "show me chinese restaurants",
"intent": "restaurant_search",
"entities": [
{
"start": 8,
"end": 15,
"value": "chinese",
"entity": "cuisine",
"extractor": "CRFEntityExtractor",
"confidence": 0.854,
"processors": []
}
]
}
Pipeline Templates (deprecated)
A template serves as a shortcut for listing components. The use of templates allows for quicker setups, but consider listing components directly for full control.
pretrained_embeddings_spacy
This pipeline takes advantage of pre-trained word vectors for better semantic understanding.
pretrained_embeddings_convert
Utilizes the ConveRT model for contextual vector representations.
supervised_embeddings
Adapts word vectors to align with the specific domain of your dataset, enhancing performance.