These docs are for version 1.x of Rasa Open Source.

## User Guide

- [Installation](https://legacy-docs-v1.rasa.com/1.9.6/user-guide/installation/)
- [Tutorial: Rasa Basics](https://legacy-docs-v1.rasa.com/1.9.6/user-guide/rasa-tutorial/)
- [Tutorial: Building Assistants](https://legacy-docs-v1.rasa.com/1.9.6/user-guide/building-assistants/)
- [Command Line Interface](https://legacy-docs-v1.rasa.com/1.9.6/user-guide/command-line-interface/)
- [Architecture](https://legacy-docs-v1.rasa.com/1.9.6/user-guide/architecture/)
- [Messaging and Voice Channels](https://legacy-docs-v1.rasa.com/1.9.6/user-guide/messaging-and-voice-channels/)
- [Testing Your Assistant](https://legacy-docs-v1.rasa.com/1.9.6/user-guide/testing-your-assistant/)
- [Setting up CI/CD](https://legacy-docs-v1.rasa.com/1.9.6/user-guide/setting-up-ci-cd/)
- [Validate Data](https://legacy-docs-v1.rasa.com/1.9.6/user-guide/validate-files/)
- [Configuring the HTTP API](https://legacy-docs-v1.rasa.com/1.9.6/user-guide/configuring-http-api/)
- [Deploying Your Rasa Assistant](https://legacy-docs-v1.rasa.com/1.9.6/user-guide/how-to-deploy/)
- [Cloud Storage](https://legacy-docs-v1.rasa.com/1.9.6/user-guide/cloud-storage/)

## NLU

- [About](https://legacy-docs-v1.rasa.com/1.9.6/nlu/about/)
- [Using NLU Only](https://legacy-docs-v1.rasa.com/1.9.6/nlu/using-nlu-only/)
- [Training Data Format](https://legacy-docs-v1.rasa.com/1.9.6/nlu/training-data-format/)
- [Language Support](https://legacy-docs-v1.rasa.com/1.9.6/nlu/language-support/)
- [Choosing a Pipeline](https://legacy-docs-v1.rasa.com/1.9.6/nlu/choosing-a-pipeline/)
- [Components](https://legacy-docs-v1.rasa.com/1.9.6/nlu/components/)
- [Entity Extraction](https://legacy-docs-v1.rasa.com/1.9.6/nlu/entity-extraction/#)

## Core

- [About](https://legacy-docs-v1.rasa.com/1.9.6/core/about/)
- [Stories](https://legacy-docs-v1.rasa.com/1.9.6/core/stories/)
- [Domains](https://legacy-docs-v1.rasa.com/1.9.6/core/domains/)
- [Responses](https://legacy-docs-v1.rasa.com/1.9.6/core/responses/)
- [Actions](https://legacy-docs-v1.rasa.com/1.9.6/core/actions/)
- [Reminders and External Events](https://legacy-docs-v1.rasa.com/1.9.6/core/reminders-and-external-events/)
- [Policies](https://legacy-docs-v1.rasa.com/1.9.6/core/policies/)
- [Slots](https://legacy-docs-v1.rasa.com/1.9.6/core/slots/)
- [Forms](https://legacy-docs-v1.rasa.com/1.9.6/core/forms/)
- [Retrieval Actions](https://legacy-docs-v1.rasa.com/1.9.6/core/retrieval-actions/)
- [Interactive Learning](https://legacy-docs-v1.rasa.com/1.9.6/core/interactive-learning/)
- [Fallback Actions](https://legacy-docs-v1.rasa.com/1.9.6/core/fallback-actions/)
- [Knowledge Base Actions](https://legacy-docs-v1.rasa.com/1.9.6/core/knowledge-bases/)

## Conversation Design

- [Dialogue Elements](https://legacy-docs-v1.rasa.com/1.9.6/dialogue-elements/dialogue-elements/)
- [Small Talk](https://legacy-docs-v1.rasa.com/1.9.6/dialogue-elements/small-talk/)
- [Completing Tasks](https://legacy-docs-v1.rasa.com/1.9.6/dialogue-elements/completing-tasks/)
- [Guiding Users](https://legacy-docs-v1.rasa.com/1.9.6/dialogue-elements/guiding-users/)

## API Reference

- [Action Server](https://legacy-docs-v1.rasa.com/1.9.6/api/action-server/)
- [HTTP API](https://legacy-docs-v1.rasa.com/1.9.6/api/http-api/)
- [Jupyter Notebooks](https://legacy-docs-v1.rasa.com/1.9.6/api/jupyter-notebooks/)
- [Agent](https://legacy-docs-v1.rasa.com/1.9.6/api/agent/)
- [Custom NLU Components](https://legacy-docs-v1.rasa.com/1.9.6/api/custom-nlu-components/)
- [Rasa SDK](https://legacy-docs-v1.rasa.com/1.9.6/api/rasa-sdk/)
- [Events](https://legacy-docs-v1.rasa.com/1.9.6/api/events/)
- [Tracker](https://legacy-docs-v1.rasa.com/1.9.6/api/tracker/)
- [Tracker Stores](https://legacy-docs-v1.rasa.com/1.9.6/api/tracker-stores/)
- [Event Brokers](https://legacy-docs-v1.rasa.com/1.9.6/api/event-brokers/)
- [Lock Stores](https://legacy-docs-v1.rasa.com/1.9.6/api/lock-stores/)
- [Training Data Importers](https://legacy-docs-v1.rasa.com/1.9.6/api/training-data-importers/)
- [Featurization of Conversations](https://legacy-docs-v1.rasa.com/1.9.6/api/core-featurization/)
- [TensorFlow Configuration](https://legacy-docs-v1.rasa.com/1.9.6/api/tensorflow_usage/)
- [Migration Guide](https://legacy-docs-v1.rasa.com/1.9.6/migration-guide/)
- [Rasa Open Source Change Log](https://legacy-docs-v1.rasa.com/1.9.6/changelog/)

## Migrate from (beta)

- [Dialogflow](https://legacy-docs-v1.rasa.com/1.9.6/migrate-from/google-dialogflow-to-rasa/)
- [Wit.ai](https://legacy-docs-v1.rasa.com/1.9.6/migrate-from/facebook-wit-ai-to-rasa/)
- [LUIS](https://legacy-docs-v1.rasa.com/1.9.6/migrate-from/microsoft-luis-to-rasa/)
- [IBM Watson](https://legacy-docs-v1.rasa.com/1.9.6/migrate-from/ibm-watson-to-rasa/)

## Reference

- [Glossary](https://legacy-docs-v1.rasa.com/1.9.6/glossary/)

## Versions

viewing: 1.9.6

# Entity Extraction

Entity extraction involves parsing user messages for required pieces of information. Rasa Open Source provides entity extractors for custom entities as well as pre-trained ones like dates and locations. Here is a summary of the available extractors and what they are used for:

| Component                 | Requires                 | Model                                     | Notes                             |
| ------------------------- | ------------------------ | ---------------------------------------- | --------------------------------- |
| `CRFEntityExtractor`     | sklearn-crfsuite         | conditional random field                  | good for training custom entities  |
| `SpacyEntityExtractor`   | spaCy                    | averaged perceptron                      | provides pre-trained entities      |
| `DucklingHTTPExtractor`  | running duckling         | context-free grammar                     | provides pre-trained entities      |
| `MitieEntityExtractor`   | MITIE                    | structured SVM                           | good for training custom entities  |
| `EntitySynonymMapper`    | existing entities        | N/A                                      | maps known synonyms                |
| `DIETClassifier`         |                          | conditional random field<br>on top of a transformer | good for training custom entities  |

## [The “entity” Object](https://legacy-docs-v1.rasa.com/1.9.6/nlu/entity-extraction/#the-entity-object)

After parsing, an entity is returned as a dictionary. There are two fields that show information about how the pipeline impacted the entities returned: the `extractor` field of an entity tells you which entity extractor found this particular entity, and the `processors` field contains the name of components that altered this specific entity.

The use of synonyms can cause the `value` field not match the `text` exactly. Instead it will return the trained synonym.

```json
{
  "text": "show me chinese restaurants",
  "intent": "restaurant_search",
  "entities": [
    {
      "start": 8,
      "end": 15,
      "value": "chinese",
      "entity": "cuisine",
      "extractor": "CRFEntityExtractor",
      "confidence": 0.854,
      "processors": []
    }
  ]
}
```

## [Custom Entities](https://legacy-docs-v1.rasa.com/1.9.6/nlu/entity-extraction/#custom-entities)

Almost every chatbot and voice app will have some custom entities. A restaurant assistant should understand `chinese` as a cuisine, but to a language-learning assistant it would mean something very different. The `CRFEntityExtractor` component can learn custom entities in any language, given some training data. See [Training Data Format](https://legacy-docs-v1.rasa.com/1.9.6/nlu/training-data-format/#training-data-format) for details on how to include entities in your training data.

## [Extracting Places, Dates, People, Organisations](https://legacy-docs-v1.rasa.com/1.9.6/nlu/entity-extraction/#extracting-places-dates-people-organisations)

spaCy has excellent pre-trained named-entity recognisers for a few different languages. You can test them out in this [interactive demo](https://demos.explosion.ai/displacy-ent/). We don’t recommend that you try to train your own NER using spaCy, unless you have a lot of data and know what you are doing. Note that some spaCy models are highly case-sensitive.

## [Dates, Amounts of Money, Durations, Distances, Ordinals](https://legacy-docs-v1.rasa.com/1.9.6/nlu/entity-extraction/#dates-amounts-of-money-durations-distances-ordinals)

The [duckling](https://duckling.wit.ai/) library does a great job of turning expressions like “next Thursday at 8pm” into actual datetime objects that you can use, e.g.

```json
"next Thursday at 8pm"
=> {"value":"2018-05-31T20:00:00.000+01:00"}
```

The list of supported languages can be found [here](https://github.com/facebook/duckling/tree/master/Duckling/Dimensions). Duckling can also handle durations like “two hours”, amounts of money, distances, and ordinals. Fortunately, there is a duckling docker container ready to use, that you just need to spin up and connect to Rasa NLU (see [DucklingHTTPExtractor](https://legacy-docs-v1.rasa.com/1.9.6/nlu/components/#ducklinghttpextractor)).

## [Regular Expressions (regex)](https://legacy-docs-v1.rasa.com/1.9.6/nlu/entity-extraction/#regular-expressions-regex)

You can use regular expressions to help the CRF model learn to recognize entities. In your training data (see [Training Data Format](https://legacy-docs-v1.rasa.com/1.9.6/nlu/training-data-format/#training-data-format)) you can provide a list of regular expressions, each of which provides the `CRFEntityExtractor` with an extra binary feature, which says if the regex was found (1) or not (0).

For example, the names of German streets often end in `strasse`. By adding this as a regex, we are telling the model to pay attention to words ending this way, and will quickly learn to associate that with a location entity.

## [Passing Custom Features to `CRFEntityExtractor`](https://legacy-docs-v1.rasa.com/1.9.6/nlu/entity-extraction/#passing-custom-features-to-crfentityextractor)

If you want to pass custom features, such as pre-trained word embeddings, to `CRFEntityExtractor`, you can add any dense featurizer to the pipeline before the `CRFEntityExtractor`. `CRFEntityExtractor` automatically finds the additional dense features and checks if the dense features are an iterable of `len(tokens)`, where each entry is a vector. A warning will be shown in case the check fails. However, `CRFEntityExtractor` will continue to train just without the additional custom features. In case dense features are present, `CRFEntityExtractor` will pass the dense features to `sklearn_crfsuite` and use them for training.

👋 I can help you get started with Rasa and answer your technical questions.
