These docs are for version 1.x of Rasa Open Source.

# Policies

## Configuring Policies

The `rasa.core.policies.Policy` class decides which action to take at every step in the conversation.

There are different policies to choose from, and you can include multiple policies in a single [`rasa.core.agent.Agent`](https://legacy-docs-v1.rasa.com/1.10.23/api/agent/#rasa.core.agent.Agent).

Note

Per default a maximum of 10 next actions can be predicted by the agent after every user message. To update this value you can set the environment variable `MAX_NUMBER_OF_PREDICTIONS` to the desired number of maximum predictions.

Your project’s `config.yml` file takes a `policies` key which you can use to customize the policies your assistant uses. In the example below, the last two lines show how to use a custom policy class and pass arguments to it.

```
policies:
  - name: "KerasPolicy"
    featurizer:
    - name: MaxHistoryTrackerFeaturizer
      max_history: 5
      state_featurizer:
        - name: BinarySingleStateFeaturizer
  - name: "MemoizationPolicy"
    max_history: 5
  - name: "FallbackPolicy"
    nlu_threshold: 0.4
    core_threshold: 0.3
    fallback_action_name: "my_fallback_action"
  - name: "path.to.your.policy.class"
    arg1: "..."
```

### Max History

One important hyperparameter for Rasa Core policies is the `max_history`.

You can set the `max_history` by passing it to your policy’s `Featurizer` in the policy configuration yaml file.

Note

Only the `MaxHistoryTrackerFeaturizer` uses a max history, whereas the `FullDialogueTrackerFeaturizer` always looks at the full conversation history. See [Featurization of Conversations](https://legacy-docs-v1.rasa.com/1.10.23/api/core-featurization/#featurization-conversations) for details.

### Data Augmentation

When you train a model, by default Rasa Core will create longer stories by randomly gluing together the ones in your stories files.

You can alter this behavior with the `--augmentation` flag.

### Action Selection

At every turn, each policy defined in your configuration will predict a next action with a certain confidence level. The bot’s next action is then decided by the policy that predicts with the highest confidence.

If you create your own policy, use these priorities as a guide for figuring out the priority of your policy.

### Keras Policy

The `KerasPolicy` uses a neural network implemented in Keras to select the next action.

```
def model_architecture(
    self, input_shape: Tuple[int, int], output_shape: Tuple[int, Optional[int]]
) -> tf.keras.models.Sequential:
    from tensorflow.keras.models import Sequential
    from tensorflow.keras.layers import (
        Masking,
        LSTM,
        Dense,
        TimeDistributed,
        Activation,
    )

model = Sequential()

if len(output_shape) == 1:
        model.add(Masking(mask_value=-1, input_shape=input_shape))
        model.add(LSTM(self.rnn_size, dropout=0.2))
        model.add(Dense(input_dim=self.rnn_size, units=output_shape[-1]))
    elif len(output_shape) == 2:
        model.add(Masking(mask_value=-1, input_shape=(None, input_shape[1])))
        model.add(LSTM(self.rnn_size, return_sequences=True, dropout=0.2))
        model.add(TimeDistributed(Dense(units=output_shape[-1])))
    else:
        raise ValueError(
            "Cannot construct the model because"
            "length of output_shape = {} "
            "should be 1 or 2."
            "".format(len(output_shape))
        )

model.add(Activation("softmax"))
    model.compile(
        loss="categorical_crossentropy", optimizer="rmsprop", metrics=["accuracy"]
    )
    return model
```

### Mapping Policy

The `MappingPolicy` can be used to directly map intents to actions.

```
intents:
 - ask_is_bot:
     triggers: action_is_bot
```

### Memoization Policy

The `MemoizationPolicy` just memorizes the conversations in your training data.

### Fallback Policy

The `FallbackPolicy` invokes a fallback action if at least one of the following occurs:

1. The intent recognition has a confidence below `nlu_threshold`.
2. The highest ranked intent differs in confidence with the second highest ranked intent by less than `ambiguity_threshold`.
3. None of the dialogue policies predict an action with confidence higher than `core_threshold`.

**Configuration:**

```
policies:
   - name: "FallbackPolicy"
     nlu_threshold: 0.3
     ambiguity_threshold: 0.1
     core_threshold: 0.3
     fallback_action_name: 'action_default_fallback'
```

### Two-Stage Fallback Policy

The `TwoStageFallbackPolicy` handles low NLU confidence in multiple stages by trying to disambiguate the user input.

**Configuration:**

```
policies:
   - name: TwoStageFallbackPolicy
     nlu_threshold: 0.3
     ambiguity_threshold: 0.1
     core_threshold: 0.3
     fallback_core_action_name: "action_default_fallback"
     fallback_nlu_action_name: "action_default_fallback"
     deny_suggestion_intent_name: "out_of_scope"
```

### Form Policy

The `FormPolicy` is an extension of the `MemoizationPolicy` which handles the filling of forms. Once a `FormAction` is called, the `FormPolicy` will continually predict the `FormAction` until all required slots in the form are filled.
