Choosing a Pipeline

Choosing a Pipeline

Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.

The Short Answer

If your training data is in English, a good starting point is the following pipeline:

language: "en"

pipeline:
  - name: ConveRTTokenizer
  - name: ConveRTFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector

In case your training data is in a different language than English, use the following pipeline:

language: "en"

pipeline:
  - name: WhitespaceTokenizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector

A Longer Answer

We recommend using the following pipeline if your training data is in English:

language: "en"

The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.8.0/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance. Pre-trained word embeddings are helpful as they encode some linguistic knowledge.

If your training data is not in English but you still want to use pre-trained word embeddings, we recommend using the following pipeline:

language: "en"

pipeline:


If you don’t use any pre-trained word embeddings, you can train your model to be more domain specific. If your domain-specific terminology is not captured by any word embeddings, we recommend:

language: "en"

Choosing the right Components

A pipeline usually consists of three main parts:

  1. Tokenization
  2. Featurization
  3. Entity Recognition / Intent Classification / Response Selectors

Tokenization

If your chosen language is whitespace-tokenized, use the WhitespaceTokenizer. If this is not the case, use a different tokenizer. We support a number of different tokenizers.

Featurization

Decide whether to use components that provide pre-trained word embeddings. If you don’t use them, train your model to be more domain-specific.

Entity Recognition / Intent Classification / Response Selectors

Depending on your data, you may want to perform entity recognition, intent classification, or response selection.

For performance comparisons, Comparing NLU Pipelines provides tools to evaluate different configurations.

Class Imbalance

To address class imbalances, use a balanced batching strategy to ensure class representation.

language: "en"

pipeline:
# - ... other components
- name: "DIETClassifier"
  batch_strategy: sequence

Multiple Intents

To predict multiple intents, employ the DIETClassifier and specify intent handling flags.

Understanding the Rasa NLU Pipeline

In Rasa NLU, messages are processed by a sequence of components, creating outputs used by subsequent components.

{
    "text": "I am looking for Chinese food",
    "entities": [
        {
            "start": 8,
            "end": 15,
            "value": "chinese",
            "entity": "cuisine",
            "extractor": "DIETClassifier",
            "confidence": 0.864
        }
    ],
    "intent": {
        "confidence": 0.6485910906220309,
        "name": "restaurant_search"
    },
    "intent_ranking": [
        {"confidence": 0.6485910906220309, "name": "restaurant_search"},
        {"confidence": 0.1416153159565678, "name": "affirm"}
    ]
}

Component Lifecycle

Components implement methods from the Component base class, which are executed in a specific order during the training of a pipeline.

The “entity” object explained

Entities are returned as a dictionary, providing insight into the extractor that identified them and any processes they underwent.

{
  "text": "show me chinese restaurants",
  "intent": "restaurant_search",
  "entities": [
    {
      "start": 8,
      "end": 15,
      "value": "chinese",
      "entity": "cuisine",
      "extractor": "CRFEntityExtractor",
      "confidence": 0.854,
      "processors": []
    }
  ]
}

Pipeline Templates (deprecated)

A template serves as a shortcut for listing components. The use of templates allows for quicker setups, but consider listing components directly for full control.

pretrained_embeddings_spacy

This pipeline takes advantage of pre-trained word vectors for better semantic understanding.

pretrained_embeddings_convert

Utilizes the ConveRT model for contextual vector representations.

supervised_embeddings

Adapts word vectors to align with the specific domain of your dataset, enhancing performance.