# Choosing a Pipeline

Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.

## The Short Answer

If your training data is in English, a good starting point is the following pipeline:

```yaml
language: "en"

pipeline:
  - name: ConveRTTokenizer
  - name: ConveRTFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector
```

In case your training data is in a different language than English, use the following pipeline:

```yaml
language: "en"

pipeline:
  - name: WhitespaceTokenizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector
```

## A Longer Answer

We recommend using the following pipeline if your training data is in English:

```yaml
language: "en"

The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.8.2/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance. Pre-trained word embeddings are helpful as they already encode some kind of linguistic knowledge.

If your training data is not in English, but you still want to use pre-trained word embeddings, we recommend using the following pipeline:

```yaml
language: "en"

pipeline:
  - name: SpacyNLP
  - name: SpacyTokenizer
  - name: SpacyFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector
```

### Choosing the right Components

A pipeline usually consists of three main parts:

> 1. Tokenization
> 2. Featurization
> 3. Entity Recognition / Intent Classification / Response Selectors

### Tokenization

If your chosen language is whitespace-tokenized (words are separated by spaces), you can use the [WhitespaceTokenizer](https://legacy-docs-v1.rasa.com/1.8.2/nlu/components/#whitespacetokenizer). If this is not the case, you should use a different tokenizer. We support a number of different [tokenizers](https://legacy-docs-v1.rasa.com/1.8.2/nlu/components/#tokenizers), or you can create your own [custom tokenizer](https://legacy-docs-v1.rasa.com/1.8.2/api/custom-nlu-components/#custom-nlu-components).

### Featurization

You need to decide whether to use components that provide pre-trained word embeddings or not. If you don’t use any pre-trained word embeddings inside your pipeline, you are not bound to a specific language and can train your model to be more domain specific.

If your training data is in English, we recommend using the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.8.2/nlu/components/#convertfeaturizer).

### Entity Recognition / Intent Classification / Response Selectors

Depending on your data, you may want to only perform intent classification, entity recognition, or response selection. We recommend using [DIETClassifier](https://legacy-docs-v1.rasa.com/1.8.2/nlu/components/#diet-classifier) for intent classification and entity recognition and [ResponseSelector](https://legacy-docs-v1.rasa.com/1.8.2/nlu/components/#response-selector) for response selection.

### Class Imbalance

Classification algorithms often do not perform well if there is a large class imbalance, for example, if you have a lot of training data for some intents and very little training data for others.

To mitigate this problem, you can use a `balanced` batching strategy.

```yaml
language: "en"

pipeline:
# - ... other components
- name: "DIETClassifier"
  batch_strategy: sequence
```

### Multiple Intents

If you want to split intents into multiple labels, you need to use the [DIETClassifier](https://legacy-docs-v1.rasa.com/1.8.2/nlu/components/#diet-classifier) in your pipeline.

### Understanding the Rasa NLU Pipeline

In Rasa NLU, incoming messages are processed by a sequence of components. There are components for entity extraction, for intent classification, response selection, pre-processing, and others. Each component processes the input and creates an output.

### Component Lifecycle

Every component can implement several methods from the `Component` base class; in a pipeline, these methods will be called in a specific order. The components are executed one after another in a so-called processing pipeline.
