# Choosing a Pipeline

In Rasa Open Source, incoming messages are processed by a sequence of components. These components are executed one after another in a so-called processing `pipeline` defined in your `config.yml`. Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.

## How to Choose a Pipeline

### The Short Answer

If your training data is in English, a good starting point is the following pipeline:

```
language: "en"

pipeline:
  - name: ConveRTTokenizer
  - name: ConveRTFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
    epochs: 100
  - name: EntitySynonymMapper
  - name: ResponseSelector
    epochs: 100
```

If your training data is not in English, start with the following pipeline:

```
language: "fr"  # your two-letter language code

pipeline:
  - name: WhitespaceTokenizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
    epochs: 100
  - name: EntitySynonymMapper
  - name: ResponseSelector
    epochs: 100
```

### A Longer Answer

We recommend using the following pipeline, if your training data is in English:

```
language: "en"

The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.10.14/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance. Pre-trained word embeddings are helpful as they already encode some kind of linguistic knowledge.

### Choosing the Right Components

There are components for entity extraction, for intent classification, response selection, pre-processing, and others. You can learn more about any specific component on the [Components](https://legacy-docs-v1.rasa.com/1.10.14/nlu/components/#components) page. If you want to add your own component, for example to run a spell-check or to do sentiment analysis, check out [Custom NLU Components](https://legacy-docs-v1.rasa.com/1.10.14/api/custom-nlu-components/#custom-nlu-components).

A pipeline usually consists of three main parts:

- [Tokenization](https://legacy-docs-v1.rasa.com/1.10.14/nlu/choosing-a-pipeline/#tokenization)
- [Featurization](https://legacy-docs-v1.rasa.com/1.10.14/nlu/choosing-a-pipeline/#featurization)
- [Entity Recognition / Intent Classification / Response Selectors](https://legacy-docs-v1.rasa.com/1.10.14/nlu/choosing-a-pipeline/#entity-recognition-intent-classification-response-selectors)

## Component Lifecycle

Each component processes an input and/or creates an output. The order of the components is determined by the order they are listed in the `config.yml`; the output of a component can be used by any other component that comes after it in the pipeline.

For example, for the sentence "I am looking for Chinese food", the output is:

```
{
    "text": "I am looking for Chinese food",
    "entities": [
        {
            "start": 8,
            "end": 15,
            "value": "chinese",
            "entity": "cuisine",
            "extractor": "DIETClassifier",
            "confidence": 0.864
        }
    ],
    "intent": {"confidence": 0.6485910906220309, "name": "restaurant_search"},
    "intent_ranking": [
        {"confidence": 0.6485910906220309, "name": "restaurant_search"},
        {"confidence": 0.1416153159565678, "name": "affirm"}
    ]
}
```

This is created as a combination of the results of the different components in the following pipeline:

```
pipeline:
  - name: WhitespaceTokenizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector
```
