These docs are for version 1.x of Rasa Open Source.

## Warning
This document is for an old version of Rasa.

# Evaluating Models

### Evaluating an NLU Model
A standard technique in machine learning is to keep some data separate as a _test set_. You can [split your NLU training data](https://legacy-docs-v1.rasa.com/1.4.6/user-guide/command-line-interface/#train-test-split) into train and test sets using:

```
rasa data split nlu
```

If you’ve done this, you can see how well your NLU model predicts the test cases using this command:

```
rasa test nlu -u test_set.md --model models/nlu-20180323-145833.tar.gz
```

If you don’t want to create a separate test set, you can still estimate how well your model generalises using cross-validation. To do this, add the flag `--cross-validation`:

```
rasa test nlu -u data/nlu.md --config config.yml --cross-validation
```

### Comparing NLU Pipelines
By passing multiple pipeline configurations (or a folder containing them) to the CLI, Rasa will run a comparative examination between the pipelines.

```
$ rasa test nlu --config pretrained_embeddings_spacy.yml supervised_embeddings.yml --nlu data/nlu.md --runs 3 --percentages 0 25 50 70 90
```

### Intent Classification
The evaluation script will produce a report, confusion matrix, and confidence histogram for your model. The report logs precision, recall and f1 measure for each intent and entity, as well as providing an overall average. You can save these reports as JSON files using the `--report` argument.

### Response Selection
The evaluation script will produce a combined report for all response selector models in your pipeline. The report logs precision, recall and f1 measure for each response, as well as providing an overall average. You can save these reports as JSON files using the `--report` argument.

### Evaluating a Core Model
You can evaluate your trained model on a set of test stories by using the evaluate script:

```
rasa test core --stories test_stories.md --out results
```

This will print the failed stories to `results/failed_stories.md`. We count any story as failed if at least one of the actions was predicted incorrectly.

### Comparing Core Configurations
To choose a configuration for your core model, or to choose hyperparameters for a specific policy, you want to measure how well Rasa Core will generalise to conversations which it hasn’t seen before. To do this, you first have to train models for your different configurations. Create two (or more) config files including the policies you want to compare, and then use the `compare` mode of the train script to train your models:

```
$ rasa train core -c config_1.yml config_2.yml -d domain.yml -s stories_folder --out comparison_models --runs 3 --percentages 0 5 25 50 70 95
```

### End-to-End Evaluation
Rasa lets you evaluate dialogues end-to-end, running through test conversations and making sure that both NLU and Core make correct predictions. To do this, you need some stories in the end-to-end format, which includes both the NLU output and the original text. Here is an example:

```
## end-to-end story 1
* greet: hello
   - utter_ask_howcanhelp
* inform: show me [chinese](cuisine) restaurants
   - utter_ask_location
* inform: in [Paris](location)
   - utter_ask_price
```

If you’ve saved end-to-end stories as a file called `e2e_stories.md`, you can evaluate your model against them by running:

```
$ rasa test --stories e2e_stories.md --e2e
```
