# Testing Your Assistant

## End-to-End Testing
Rasa Open Source lets you test dialogues end-to-end by running through test conversations and making sure that both NLU and Core make correct predictions.

To do this, you need some stories in the end-to-end format, which includes both the NLU output and the original text. Here are some examples:

By default Rasa Open Source saves conversation tests to `tests/conversation_tests.md`. You can test your assistant against them by running:

```bash
$ rasa test
```

**Note**: [Custom Actions](https://legacy-docs-v1.rasa.com/1.10.18/core/actions/#custom-actions) are not executed as part of end-to-end tests. If your custom actions append any events to the tracker, this has to be reflected in your end-to-end tests (e.g., by adding `slot` events to your end-to-end story).

If you have any questions or problems, please share them with us in the dedicated [testing section on our forum](https://forum.rasa.com/tags/testing) !

### Evaluating an NLU Model
A standard technique in machine learning is to keep some data separate as a _test set_. You can [split your NLU training data](https://legacy-docs-v1.rasa.com/1.10.18/user-guide/command-line-interface/#train-test-split) into train and test sets using:

```bash
rasa data split nlu
```

If you’ve done this, you can see how well your NLU model predicts the test cases using this command:

```bash
rasa test nlu -u train_test_split/test_data.md --model models/nlu-20180323-145833.tar.gz
```

**If you don’t want to create a separate test set,** you can still estimate how well your model generalizes using cross-validation. To do this, add the flag `--cross-validation`:

```bash
rasa test nlu -u data/nlu.md --config config.yml --cross-validation
```

### Comparing NLU Pipelines
By passing multiple pipeline configurations (or a folder containing them) to the CLI, Rasa will run a comparative examination between the pipelines.

```bash
$ rasa test nlu --config pretrained_embeddings_spacy.yml supervised_embeddings.yml --nlu data/nlu.md --runs 3 --percentages 0 25 50 70 90
```

### Evaluating a Core Model
You can evaluate your trained model on a set of test stories by using the evaluate script:

```bash
rasa test core --stories test_stories.md --out results
```

This will print the failed stories to `results/failed_stories.md`. We count any story as failed if at least one of the actions was predicted incorrectly.

### Comparing Core Configurations
To choose a configuration for your core model, or to choose hyperparameters for a specific policy, you want to measure how well Rasa Core will generalize to conversations which it hasn’t seen before.

Rasa Core has some scripts to help you choose and fine-tune your policy configuration. Once you are happy with it, you can then train your final configuration on your full data set. To do this, you first have to train models for your different configurations. Create two (or more) config files including the policies you want to compare, and then use the `compare` mode of the train script to train your models:

```bash
$ rasa train core -c config_1.yml config_2.yml -d domain.yml -s stories_folder --out comparison_models --runs 3 --percentages 0 5 25 50 70 95
```

Once this script has finished, you can use the evaluate script in `compare` mode to evaluate the models you just trained:

```bash
$ rasa test core -m comparison_models --stories stories_folder --out comparison_results --evaluate-model-directory
```

### Important Notes
- Make sure your model file in `models` is a combined `core` and `nlu` model.
- Check out this [tutorial](https://blog.rasa.com/rasa-nlu-in-depth-part-3-hyperparameters/) for tuning the hyperparameters of your NLU model.
