Choosing a Pipeline
Choosing a Pipeline
In Rasa Open Source, incoming messages are processed by a sequence of components. These components are executed one after another in a so-called processing pipeline defined in your config.yml. Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.
How to Choose a Pipeline
The Short Answer
If your training data is in English, a good starting point is the following pipeline:
language: "en"
pipeline:
- name: ConveRTTokenizer
- name: ConveRTFeaturizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: "char_wb"
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
epochs: 100
- name: EntitySynonymMapper
- name: ResponseSelector
epochs: 100
If your training data is not in English, start with the following pipeline:
language: "fr" # your two-letter language code
pipeline:
- name: WhitespaceTokenizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: "char_wb"
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
epochs: 100
- name: EntitySynonymMapper
- name: ResponseSelector
epochs: 100
A Longer Answer
We recommend using the following pipeline, if your training data is in English:
language: "en"
If you don’t use any pre-trained word embeddings inside your pipeline, you are not bound to a specific language and can train your model to be more domain-specific.
### Comparing Pipelines
Rasa gives you the tools to compare the performance of multiple pipelines on your data directly.
### Handling Class Imbalance
To mitigate the problem of class imbalance, use a `balanced` batching strategy.
```yaml
language: "en"
pipeline:
- name: "DIETClassifier"
batch_strategy: sequence
Component Lifecycle
Each component processes an input and/or creates an output. The order of the components is determined by the order they are listed in the config.yml.