# Choosing a Pipeline

In Rasa Open Source, incoming messages are processed by a sequence of components. These components are executed one after another in a so-called processing `pipeline` defined in your `config.yml`. Choosing an NLU pipeline allows you to customize your model and finetune it on your dataset.

- [How to Choose a Pipeline](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#how-to-choose-a-pipeline)

- [The Short Answer](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#the-short-answer)

- [A Longer Answer](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#a-longer-answer)

- [Choosing the Right Components](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#choosing-the-right-components)

- [Multi-Intent Classification](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#multi-intent-classification)
- [Comparing Pipelines](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#comparing-pipelines)

- [Handling Class Imbalance](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#handling-class-imbalance)

- [Component Lifecycle](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#component-lifecycle)

## [How to Choose a Pipeline](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#id6)

### [The Short Answer](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#id7)

If your training data is in English, a good starting point is the following pipeline:

> ```
> language: "en"
>
> pipeline:
>   - name: ConveRTTokenizer
>   - name: ConveRTFeaturizer
>   - name: RegexFeaturizer
>   - name: LexicalSyntacticFeaturizer
>   - name: CountVectorsFeaturizer
>   - name: CountVectorsFeaturizer
>     analyzer: "char_wb"
>     min_ngram: 1
>     max_ngram: 4
>   - name: DIETClassifier
>     epochs: 100
>   - name: EntitySynonymMapper
>   - name: ResponseSelector
>     epochs: 100
> ```

If your training data is not in English, start with the following pipeline:

> ```
> language: "fr"  # your two-letter language code
>
> pipeline:
>   - name: WhitespaceTokenizer
>   - name: RegexFeaturizer
>   - name: LexicalSyntacticFeaturizer
>   - name: CountVectorsFeaturizer
>   - name: CountVectorsFeaturizer
>     analyzer: "char_wb"
>     min_ngram: 1
>     max_ngram: 4
>   - name: DIETClassifier
>     epochs: 100
>   - name: EntitySynonymMapper
>   - name: ResponseSelector
>     epochs: 100
> ```

### [A Longer Answer](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#id8)

We recommend using following pipeline, if your training data is in English:

The pipeline contains the [ConveRTFeaturizer](https://legacy-docs-v1.rasa.com/1.9.5/nlu/components/#convertfeaturizer) that provides pre-trained word embeddings of the user utterance. Pre-trained word embeddings are helpful as they already encode some kind of linguistic knowledge.

### [Choosing the Right Components](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#id9)

There are components for entity extraction, for intent classification, response selection, pre-processing, and others. You can learn more about any specific component on the [Components](https://legacy-docs-v1.rasa.com/1.9.5/nlu/components/#components) page. If you want to add your own component, for example to run a spell-check or to do sentiment analysis, check out [Custom NLU Components](https://legacy-docs-v1.rasa.com/1.9.5/api/custom-nlu-components/#custom-nlu-components).

A pipeline usually consists of three main parts:

- [Tokenization](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#tokenization)

- [Featurization](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#featurization)

- [Entity Recognition / Intent Classification / Response Selectors](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#entity-recognition-intent-classification-response-selectors)

#### [Tokenization](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#id15)

For tokenization of English input, we recommend the [ConveRTTokenizer](https://legacy-docs-v1.rasa.com/1.9.5/nlu/components/#converttokenizer). You can process other whitespace-tokenized (words are separated by spaces) languages with the [WhitespaceTokenizer](https://legacy-docs-v1.rasa.com/1.9.5/nlu/components/#whitespacetokenizer). If your language is not whitespace-tokenized, you should use a different tokenizer.

### [Featurization](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#id16)

You need to decide whether to use components that provide pre-trained word embeddings or not. We recommend in cases of small amounts of training data to start with pre-trained word embeddings.

#### [Entity Recognition / Intent Classification / Response Selectors](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#id17)

Depending on your data you may want to only perform intent classification or entity recognition or response selection. We recommend using [DIETClassifier](https://legacy-docs-v1.rasa.com/1.9.5/nlu/components/#diet-classifier) for intent classification and entity recognition and [ResponseSelector](https://legacy-docs-v1.rasa.com/1.9.5/nlu/components/#response-selector) for response selection.

### [Multi-Intent Classification](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#id10)

You can use Rasa Open Source components to split intents into multiple labels. For example, you can predict multiple intents (`thank+goodbye`) or model hierarchical intent structure (`feedback+positive` being more similar to `feedback+negative` than `chitchat`).

### [Comparing Pipelines](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#id11)

Rasa gives you the tools to compare the performance of multiple pipelines on your data directly. See [Comparing NLU Pipelines](https://legacy-docs-v1.rasa.com/1.9.5/user-guide/testing-your-assistant/#comparing-nlu-pipelines) for more information.

### [Handling Class Imbalance](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#id12)

Classification algorithms often do not perform well if there is a large class imbalance. To mitigate this problem, you can use a `balanced` batching strategy.

### [Component Lifecycle](https://legacy-docs-v1.rasa.com/1.9.5/nlu/choosing-a-pipeline/#id13)

Each component processes an input and/or creates an output. The order of the components is determined by the order they are listed in the `config.yml`; the output of a component can be used by any other component that comes after it in the pipeline.

👋 I can help you get started with Rasa and answer your technical questions.
