# Choosing a Pipeline  
In Rasa Open Source, incoming messages are processed by a sequence of components. These components are executed one after another in a so-called processing `pipeline` defined in your `config.yml`. Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.

## How to Choose a Pipeline
### The Short Answer
If your training data is in English, a good starting point is the following pipeline:

```
language: "en"

pipeline:
  - name: ConveRTTokenizer
  - name: ConveRTFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
    epochs: 100
  - name: EntitySynonymMapper
  - name: ResponseSelector
    epochs: 100
```

If your training data is not in English, start with the following pipeline:

```
language: "fr"  # your two-letter language code

pipeline:
  - name: WhitespaceTokenizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
    epochs: 100
  - name: EntitySynonymMapper
  - name: ResponseSelector
    epochs: 100
```

### A Longer Answer
We recommend using the following pipeline, if your training data is in English:

```
language: "en"

The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.9.4/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance. Pre-trained word embeddings are helpful as they already encode some kind of linguistic knowledge.

An alternative recommendation for training data not in English:

```
language: "fr"  # your two-letter language code

pipeline:
  - name: SpacyNLP
  - name: SpacyTokenizer
  - name: SpacyFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
    epochs: 100
  - name: EntitySynonymMapper
  - name: ResponseSelector
    epochs: 100
```

### Choosing the Right Components
There are components for entity extraction, for intent classification, response selection, pre-processing, and others. You can learn more about any specific component on the [Components](https://legacy-docs-v1.rasa.com/1.9.4/nlu/components/#components) page.

## Tokenization
For tokenization of English input, we recommend the [ConveRTTokenizer](https://legacy-docs-v1.rasa.com/1.9.4/nlu/components/#converttokenizer).

### Featurization
You need to decide whether to use components that provide pre-trained word embeddings or not. We recommend using pre-trained word embeddings initially if you have a limited amount of training data.

### Entity Recognition / Intent Classification / Response Selectors
We recommend using [DIETClassifier](https://legacy-docs-v1.rasa.com/1.9.4/nlu/components/#diet-classifier) for intent classification and entity recognition and [ResponseSelector](https://legacy-docs-v1.rasa.com/1.9.4/nlu/components/#response-selector) for response selection.
