# Choosing a Pipeline

In Rasa Open Source, incoming messages are processed by a sequence of components.
These components are executed one after another in a so-called processing `pipeline` defined in your `config.yml`.
Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.

## [How to Choose a Pipeline](https://legacy-docs-v1.rasa.com/1.10.3/nlu/choosing-a-pipeline/#id6)
### [The Short Answer](https://legacy-docs-v1.rasa.com/1.10.3/nlu/choosing-a-pipeline/#id7)
If your training data is in English, a good starting point is the following pipeline:

```
language: "en"

pipeline:
  - name: ConveRTTokenizer
  - name: ConveRTFeaturizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
    epochs: 100
  - name: EntitySynonymMapper
  - name: ResponseSelector
    epochs: 100
```

If your training data is not in English, start with the following pipeline:

```
language: "fr"  # your two-letter language code

pipeline:
  - name: WhitespaceTokenizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
    epochs: 100
  - name: EntitySynonymMapper
  - name: ResponseSelector
    epochs: 100
```

### [Choosing the Right Components](https://legacy-docs-v1.rasa.com/1.10.3/nlu/choosing-a-pipeline/#id9)
There are components for entity extraction, for intent classification, response selection,
pre-processing, and others. You can learn more about any specific component on the [Components](https://legacy-docs-v1.rasa.com/1.10.3/nlu/components/#components) page.

A pipeline usually consists of three main parts:

- [Tokenization](https://legacy-docs-v1.rasa.com/1.10.3/nlu/choosing-a-pipeline/#tokenization)
- [Featurization](https://legacy-docs-v1.rasa.com/1.10.3/nlu/choosing-a-pipeline/#featurization)
- [Entity Recognition / Intent Classification / Response Selectors](https://legacy-docs-v1.rasa.com/1.10.3/nlu/choosing-a-pipeline/#entity-recognition-intent-classification-response-selectors)

### [Tokenization](https://legacy-docs-v1.rasa.com/1.10.3/nlu/choosing-a-pipeline/#id15)
For tokenization of English input, we recommend the [ConveRTTokenizer](https://legacy-docs-v1.rasa.com/1.10.3/nlu/components/#converttokenizer).
### [Featurization](https://legacy-docs-v1.rasa.com/1.10.3/nlu/choosing-a-pipeline/#id16)
You need to decide whether to use components that provide pre-trained word embeddings or not.
#### [Pre-trained Embeddings](https://legacy-docs-v1.rasa.com/1.10.3/nlu/choosing-a-pipeline/#id18)
The advantage of using pre-trained word embeddings in your pipeline is that if you have a training example like:
“I want to buy apples”, and Rasa is asked to predict the intent for “get pears”, your model already knows that the
words “apples” and “pears” are very similar. We support a few components that provide pre-trained word embeddings:

1. [MitieFeaturizer](https://legacy-docs-v1.rasa.com/1.10.3/nlu/components/#mitiefeaturizer)
2. [SpacyFeaturizer](https://legacy-docs-v1.rasa.com/1.10.3/nlu/components/#spacyfeaturizer)
3. [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.10.3/nlu/components/#convertfeaturizer)
4. [LanguageModelFeaturizer](https://legacy-docs-v1.rasa.com/1.10.3/nlu/components/#languagemodelfeaturizer)

If you don’t use any pre-trained word embeddings inside your pipeline, you are not bound to a specific language and can train your model to be more domain specific.

### [Entity Recognition / Intent Classification / Response Selectors](https://legacy-docs-v1.rasa.com/1.10.3/nlu/choosing-a-pipeline/#id17)
Depending on your data you may want to only perform intent classification, entity recognition or response selection.
### [Multi-Intent Classification](https://legacy-docs-v1.rasa.com/1.10.3/nlu/choosing-a-pipeline/#id10)
You can use Rasa Open Source components to split intents into multiple labels.  For example, you can predict
multiple intents (`thank+goodbye`) or model hierarchical intent structure (`feedback+positive` being more similar to `feedback+negative` than `chitchat`).  To do this, you need to use the [DIETClassifier](https://legacy-docs-v1.rasa.com/1.10.3/nlu/components/#diet-classifier) in your pipeline.

## [Comparing Pipelines](https://legacy-docs-v1.rasa.com/1.10.3/nlu/choosing-a-pipeline/#id11)
Rasa gives you the tools to compare the performance of multiple pipelines on your data directly.

## [Handling Class Imbalance](https://legacy-docs-v1.rasa.com/1.10.3/nlu/choosing-a-pipeline/#id12)
Classification algorithms often do not perform well if there is a large class imbalance, for example if you have a lot of training data for some intents and very little training data for others.

## [Component Lifecycle](https://legacy-docs-v1.rasa.com/1.10.3/nlu/choosing-a-pipeline/#id13)
Each component processes an input and/or creates an output. The order of the components is determined by
the order they are listed in the `config.yml`; the output of a component can be used by any other component that
comes after it in the pipeline.

For example, for the sentence "I am looking for Chinese food", the output is:

```
{
    "text": "I am looking for Chinese food",
    "entities": [\
        {\
            "start": 8,\
            "end": 15,\
            "value": "chinese",\
            "entity": "cuisine",\
            "extractor": "DIETClassifier",\
            "confidence": 0.864\
        }\
    ],
    "intent": {"confidence": 0.6485910906220309, "name": "restaurant_search"},
    "intent_ranking": [\
        {"confidence": 0.6485910906220309, "name": "restaurant_search"},\
        {"confidence": 0.1416153159565678, "name": "affirm"}\
    ]
}
```

This is created as a combination of the results of the different components in the following pipeline:

```
pipeline:
  - name: WhitespaceTokenizer
  - name: RegexFeaturizer
  - name: LexicalSyntacticFeaturizer
  - name: CountVectorsFeaturizer
  - name: CountVectorsFeaturizer
    analyzer: "char_wb"
    min_ngram: 1
    max_ngram: 4
  - name: DIETClassifier
  - name: EntitySynonymMapper
  - name: ResponseSelector
```
