# Entity Extraction

Entity extraction involves parsing user messages for required pieces of information. Rasa Open Source provides entity extractors for custom entities as well as pre-trained ones like dates and locations. Here is a summary of the available extractors and what they are used for:

| Component                  | Requires                   | Model                                          | Notes                       |
|----------------------------|----------------------------|------------------------------------------------|-----------------------------|
| `CRFEntityExtractor`       | sklearn-crfsuite           | conditional random field                      | good for training custom entities |
| `SpacyEntityExtractor`     | spaCy                      | averaged perceptron                           | provides pre-trained entities |
| `DucklingHTTPExtractor`    | running duckling           | context-free grammar                           | provides pre-trained entities |
| `MitieEntityExtractor`     | MITIE                      | structured SVM                                 | good for training custom entities |
| `EntitySynonymMapper`      | existing entities          | N/A                                            | maps known synonyms          |
| `DIETClassifier`           |                            | conditional random field<br>on top of a transformer | good for training custom entities |

- [The “entity” Object](https://legacy-docs-v1.rasa.com/1.10.12/nlu/entity-extraction/#the-entity-object)
- [Custom Entities](https://legacy-docs-v1.rasa.com/1.10.12/nlu/entity-extraction/#custom-entities)
- [Entities Roles and Groups](https://legacy-docs-v1.rasa.com/1.10.12/nlu/entity-extraction/#entities-roles-and-groups)
- [Extracting Places, Dates, People, Organizations](https://legacy-docs-v1.rasa.com/1.10.12/nlu/entity-extraction/#extracting-places-dates-people-organizations)
- [Dates, Amounts of Money, Durations, Distances, Ordinals](https://legacy-docs-v1.rasa.com/1.10.12/nlu/entity-extraction/#dates-amounts-of-money-durations-distances-ordinals)
- [Regular Expressions (regex)](https://legacy-docs-v1.rasa.com/1.10.12/nlu/entity-extraction/#regular-expressions-regex)
- [Passing Custom Features to `CRFEntityExtractor`](https://legacy-docs-v1.rasa.com/1.10.12/nlu/entity-extraction/#passing-custom-features-to-crfentityextractor)

## [The “entity” Object](https://legacy-docs-v1.rasa.com/1.10.12/nlu/entity-extraction/#id2)

After parsing, an entity is returned as a dictionary. There are two fields that show information about how the pipeline impacted the entities returned: the `extractor` field of an entity tells you which entity extractor found this particular entity, and the `processors` field contains the name of components that altered this specific entity.

The use of synonyms can cause the `value` field not match the `text` exactly. Instead it will return the trained synonym.

```json
{
  "text": "show me chinese restaurants",
  "intent": "restaurant_search",
  "entities": [
    {
      "start": 8,
      "end": 15,
      "value": "chinese",
      "entity": "cuisine",
      "extractor": "CRFEntityExtractor",
      "confidence": 0.854,
      "processors": []
    }
  ]
}
```

Note

The `confidence` will be set by the `CRFEntityExtractor` component. The `DucklingHTTPExtractor` will always return `1`. The `SpacyEntityExtractor` extractor and `DIETClassifier` do not provide this information and return `null`.

Some extractors, like `duckling`, may include additional information. For example:

```json
{
  "additional_info":{
    "grain":"day",
    "type":"value",
    "value":"2018-06-21T00:00:00.000-07:00",
    "values":[
      {
        "grain":"day",
        "type":"value",
        "value":"2018-06-21T00:00:00.000-07:00"
      }
    ]
  },
  "confidence":1.0,
  "end":5,
  "entity":"time",
  "extractor":"DucklingHTTPExtractor",
  "start":0,
  "text":"today",
  "value":"2018-06-21T00:00:00.000-07:00"
}
```

## [Custom Entities](https://legacy-docs-v1.rasa.com/1.10.12/nlu/entity-extraction/#id3)

Almost every chatbot and voice app will have some custom entities. A restaurant assistant should understand `chinese` as a cuisine, but to a language-learning assistant it would mean something very different. The `CRFEntityExtractor` and the `DIETClassifier` component can learn custom entities in any language, given some training data. See [Training Data Format](https://legacy-docs-v1.rasa.com/1.10.12/nlu/training-data-format/#training-data-format) for details on how to include entities in your training data.

## [Entities Roles and Groups](https://legacy-docs-v1.rasa.com/1.10.12/nlu/entity-extraction/#id4)

Warning

This feature is experimental. We introduce experimental features to get feedback from our community, so we encourage you to try it out! However, the functionality might be changed or removed in the future. If you have feedback (positive or negative) please share it with us on the [forum](https://forum.rasa.com/).

Assigning custom entity labels to words, allow you to define certain concepts in the data. For example, we can define what a city is:

```
I want to fly from [Berlin](city) to [San Francisco](city).
```

However, sometimes you want to specify entities even further. Let’s assume we want to build an assistant that should book a flight for us. The assistant needs to know which of the two cities in the example above is the departure city and which is the destination city. `Berlin` and `San Francisco` are still cities, but they play a different role in our example.
To distinguish between the different roles, you can assign a role label in addition to the entity label.

```
- I want to fly from [Berlin]{
