# Policies

## Configuring Policies

The `rasa.core.policies.Policy` class decides which action to take at every step in the conversation.

There are different policies to choose from, and you can include multiple policies in a single [`rasa.core.agent.Agent`](https://legacy-docs-v1.rasa.com/1.3.10/api/agent/#rasa.core.agent.Agent).

Note

Per default a maximum of 10 next actions can be predicted by the agent after every user message. To update this value you can set the environment variable `MAX_NUMBER_OF_PREDICTIONS` to the desired number of maximum predictions.

Your project’s `config.yml` file takes a `policies` key which you can use to customize the policies your assistant uses.

```yaml
policies:
  - name: "KerasPolicy"
    featurizer:
    - name: MaxHistoryTrackerFeaturizer
      max_history: 5
      state_featurizer:
        - name: BinarySingleStateFeaturizer
  - name: "MemoizationPolicy"
    max_history: 5
  - name: "FallbackPolicy"
    nlu_threshold: 0.4
    core_threshold: 0.3
    fallback_action_name: "my_fallback_action"
  - name: "path.to.your.policy.class"
    arg1: "..."
```

### Max History

One important hyperparameter for Rasa Core policies is the `max_history`. This controls how much dialogue history the model looks at to decide which action to take next.

You can set the `max_history` by passing it to your policy’s `Featurizer` in the policy configuration yaml file.

Note

Only the `MaxHistoryTrackerFeaturizer` uses a max history, whereas the `FullDialogueTrackerFeaturizer` always looks at the full conversation history.

### Data Augmentation

When you train a model, by default Rasa Core will create longer stories by randomly gluing together the ones in your stories files.

This is because if you have stories like:

```yaml
# thanks
* thankyou
   - utter_youarewelcome

# bye
* goodbye
   - utter_goodbye
```

You actually want to teach your policy to **ignore** the dialogue history when it isn’t relevant and just respond with the same action no matter what happened before.

### Action Selection

At every turn, each policy defined in your configuration will predict a next action with a certain confidence level. The bot’s next action is then decided by the policy that predicts with the highest confidence.

### Keras Policy

The `KerasPolicy` uses a neural network implemented in  [Keras](http://keras.io/) to select the next action.

```python
def model_architecture(
    self, input_shape: Tuple[int, int], output_shape: Tuple[int, Optional[int]]
) -> tf.keras.models.Sequential:
    # Build Model
    model = Sequential()
    ...
    return model
```

### Embedding Policy

Transformer Embedding Dialogue Policy (TEDP)

This policy has a pre-defined architecture which comprises the following steps:

- concatenate user input (user intent and entities), previous system action, slots, and active forms for each time step into an input vector to pre-transformer embedding layer;
- feed it to transformer;
- apply a dense layer to the output of the transformer to get embeddings of a dialogue for each time step;

### Mapping Policy

The `MappingPolicy` can be used to directly map intents to actions. The mappings are assigned by giving an intent the property `triggers`.

### Memoization Policy

The `MemoizationPolicy` just memorizes the conversations in your training data. It predicts the next action with confidence `1.0` if this exact conversation exists in the training data.

### Fallback Policy

The `FallbackPolicy` invokes a fallback action if either of the following occurs:
1. The intent recognition has a confidence below `nlu_threshold`.
2. The highest ranked intent differs in confidence with the second highest ranked intent by less than `ambiguity_threshold`.

### Two-Stage Fallback Policy

The `TwoStageFallbackPolicy` handles low NLU confidence in multiple stages by trying to disambiguate the user input.
