# Choosing a Pipeline

Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.

## The Short Answer

If your training data is in English, a good starting point is the following pipeline:

```
language: "en"

pipeline:
  - name: ConveRTTokenizer
  - name: ConveRTFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector
```

In case your training data is in a different language than English, use the following pipeline:

```
language: "en"

pipeline:
  - name: WhitespaceTokenizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector
```

## A Longer Answer

We recommend using the following pipeline if your training data is in English:

```
language: "en"

The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.8.0/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance. Pre-trained word embeddings are helpful as they encode some linguistic knowledge.

If your training data is not in English but you still want to use pre-trained word embeddings, we recommend using the following pipeline:

```
language: "en"

pipeline:
  - name: SpacyNLP
  - name: SpacyTokenizer
  - name: SpacyFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector
```

If you don’t use any pre-trained word embeddings, you can train your model to be more domain specific. If your domain-specific terminology is not captured by any word embeddings, we recommend:

```
language: "en"

## Choosing the right Components

A pipeline usually consists of three main parts:

> 1. Tokenization  
> 2. Featurization  
> 3. Entity Recognition / Intent Classification / Response Selectors

### Tokenization

If your chosen language is whitespace-tokenized, use the [WhitespaceTokenizer](https://legacy-docs-v1.rasa.com/1.8.0/nlu/components/#whitespacetokenizer). If this is not the case, use a different tokenizer. We support a number of different [tokenizers](https://legacy-docs-v1.rasa.com/1.8.0/nlu/components/#tokenizers).

### Featurization

Decide whether to use components that provide pre-trained word embeddings. If you don’t use them, train your model to be more domain-specific.

### Entity Recognition / Intent Classification / Response Selectors

Depending on your data, you may want to perform entity recognition, intent classification, or response selection.

For performance comparisons, [Comparing NLU Pipelines](https://legacy-docs-v1.rasa.com/1.8.0/user-guide/evaluating-models/#comparing-nlu-pipelines) provides tools to evaluate different configurations.

## Class Imbalance

To address class imbalances, use a balanced batching strategy to ensure class representation.

```
language: "en"

pipeline:
# - ... other components
- name: "DIETClassifier"
  batch_strategy: sequence
```

## Multiple Intents

To predict multiple intents, employ the [DIETClassifier](https://legacy-docs-v1.rasa.com/1.8.0/nlu/components/#diet-classifier) and specify intent handling flags.

## Understanding the Rasa NLU Pipeline

In Rasa NLU, messages are processed by a sequence of components, creating outputs used by subsequent components.

```json
{
    "text": "I am looking for Chinese food",
    "entities": [
        {
            "start": 8,
            "end": 15,
            "value": "chinese",
            "entity": "cuisine",
            "extractor": "DIETClassifier",
            "confidence": 0.864
        }
    ],
    "intent": {
        "confidence": 0.6485910906220309,
        "name": "restaurant_search"
    },
    "intent_ranking": [
        {"confidence": 0.6485910906220309, "name": "restaurant_search"},
        {"confidence": 0.1416153159565678, "name": "affirm"}
    ]
}
```

## Component Lifecycle

Components implement methods from the `Component` base class, which are executed in a specific order during the training of a pipeline.

## The “entity” object explained

Entities are returned as a dictionary, providing insight into the extractor that identified them and any processes they underwent.

```json
{
  "text": "show me chinese restaurants",
  "intent": "restaurant_search",
  "entities": [
    {
      "start": 8,
      "end": 15,
      "value": "chinese",
      "entity": "cuisine",
      "extractor": "CRFEntityExtractor",
      "confidence": 0.854,
      "processors": []
    }
  ]
}
```

## Pipeline Templates (deprecated)
A template serves as a shortcut for listing components. The use of templates allows for quicker setups, but consider listing components directly for full control.

### pretrained_embeddings_spacy

This pipeline takes advantage of pre-trained word vectors for better semantic understanding.

### pretrained_embeddings_convert

Utilizes the [ConveRT](https://github.com/PolyAI-LDN/polyai-models) model for contextual vector representations.

### supervised_embeddings

Adapts word vectors to align with the specific domain of your dataset, enhancing performance.
