These docs are for version 1.x of Rasa Open Source.

---

# Testing Your Assistant

## End-to-End Testing
Rasa Open Source lets you test dialogues end-to-end by running through test conversations and making sure that both NLU and Core make correct predictions.

To do this, you need some stories in the end-to-end format, which includes both the NLU output and the original text. Here are some examples:

By default Rasa Open Source saves conversation tests to `tests/conversation_tests.md`. You can test your assistant against them by running:

```
$ rasa test
```

> Note: Custom Actions are not executed as part of end-to-end tests. If your custom actions append any events to the tracker, this has to be reflected in your end-to-end tests (e.g. by adding `slot` events to your end-to-end story).

If you have any questions or problems, please share them with us in the dedicated 
[testing section on our forum](https://forum.rasa.com/tags/testing) !

### Evaluating an NLU Model
A standard technique in machine learning is to keep some data separate as a *test set*. You can [split your NLU training data](https://legacy-docs-v1.rasa.com/1.9.2/user-guide/command-line-interface/#train-test-split) into train and test sets using:

```
rasa data split nlu
```

If you’ve done this, you can see how well your NLU model predicts the test cases using this command:

```
rasa test nlu -u train_test_split/test_data.md --model models/nlu-20180323-145833.tar.gz
```

If you don’t want to create a separate test set, you can still estimate how well your model generalises using cross-validation. To do this, add the flag `--cross-validation`:

```
rasa test nlu -u data/nlu.md --config config.yml --cross-validation
```

The full list of options for the script is:
```
usage: rasa test nlu [-h] [-v] [-vv] [--quiet] [-m MODEL] [-u NLU] [--out OUT]
                     [--successes] [--no-errors] [--histogram HISTOGRAM]
                     [--confmat CONFMAT] [-c CONFIG [CONFIG ...]]
                     [--cross-validation] [-f FOLDS] [-r RUNS]
                     [-p PERCENTAGES [PERCENTAGES ...]] [--no-plot]
```

### Comparing NLU Pipelines
By passing multiple pipeline configurations (or a folder containing them) to the CLI, Rasa will run a comparative examination between the pipelines.

```
$ rasa test nlu --config pretrained_embeddings_spacy.yml supervised_embeddings.yml
  --nlu data/nlu.md --runs 3 --percentages 0 25 50 70 90
```

### Intent Classification
The evaluation script will produce a report, confusion matrix, and confidence histogram for your model.

### Response Selection
The evaluation script will produce a combined report for all response selector models in your pipeline.

### Entity Extraction
The `CRFEntityExtractor` is the only entity extractor which you train using your own data, and so is the only one that will be evaluated.

### Evaluating a Core Model
You can evaluate your trained model on a set of test stories by using the evaluate script:

```
rasa test core --stories test_stories.md --out results
```

### Comparing Core Configurations
To choose a configuration for your core model, or to choose hyperparameters for a specific policy, you want to measure how well Rasa Core will generalise to conversations which it hasn’t seen before.

---  
  
👋 I can help you get started with Rasa and answer your technical questions.
