# Choosing a Pipeline

Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.

## [The Short Answer](https://legacy-docs-v1.rasa.com/1.8.3/nlu/choosing-a-pipeline/#the-short-answer)

If your training data is in English, a good starting point is the following pipeline:

```yaml
language: "en"

pipeline:
  - name: ConveRTTokenizer
  - name: ConveRTFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector
```

In case your training data is in a different language than English, use the following pipeline:

```yaml
language: "en"

pipeline:
  - name: WhitespaceTokenizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector
```

## [A Longer Answer](https://legacy-docs-v1.rasa.com/1.8.3/nlu/choosing-a-pipeline/#a-longer-answer)

We recommend using the following pipeline if your training data is in English:

```yaml
language: "en"

The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.8.3/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance. Pre-trained word embeddings are helpful as they already encode some kind of linguistic knowledge. For example, if you have a sentence like “I want to buy apples” in your training data, and Rasa is asked to predict the intent for “get pears”, your model already knows that the words “apples” and “pears” are very similar.

If your training data is not in English, but you still want to use pre-trained word embeddings, we recommend using the following pipeline:

```yaml
language: "en"

pipeline:
  - name: SpacyNLP
  - name: SpacyTokenizer
  - name: SpacyFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector
```

### [Choosing the right Components](https://legacy-docs-v1.rasa.com/1.8.3/nlu/choosing-a-pipeline/#choosing-the-right-components)

A pipeline usually consists of three main parts:

1. Tokenization
2. Featurization
3. Entity Recognition / Intent Classification / Response Selectors

#### [Tokenization](https://legacy-docs-v1.rasa.com/1.8.3/nlu/choosing-a-pipeline/#tokenization)

If your chosen language is whitespace-tokenized (words are separated by spaces), you can use the [WhitespaceTokenizer](https://legacy-docs-v1.rasa.com/1.8.3/nlu/components/#whitespacetokenizer). If this is not the case you should use a different tokenizer.

### [Featurization](https://legacy-docs-v1.rasa.com/1.8.3/nlu/choosing-a-pipeline/#featurization)

You need to decide whether to use components that provide pre-trained word embeddings or not. If you don’t use any pre-trained word embeddings inside your pipeline, you are not bound to a specific language and can train your model to be more domain-specific.

### [Entity Recognition / Intent Classification / Response Selectors](https://legacy-docs-v1.rasa.com/1.8.3/nlu/choosing-a-pipeline/#entity-recognition-intent-classification-response-selectors)

Depending on your data you may want to only perform intent classification, entity recognition or response selection, or you might want to combine multiple of those tasks. We support several components for each of the tasks. We recommend using [DIETClassifier](https://legacy-docs-v1.rasa.com/1.8.3/nlu/components/#diet-classifier) for intent classification and entity recognition and [ResponseSelector](https://legacy-docs-v1.rasa.com/1.8.3/nlu/components/#response-selector) for response selection.
