# Entity Extraction

Entity extraction involves parsing user messages for required pieces of information. Rasa Open Source provides entity extractors for custom entities as well as pre-trained ones like dates and locations. Here is a summary of the available extractors and what they are used for:

| Component               | Requires                | Model                              | Notes                                  |
|------------------------|-------------------------|------------------------------------|----------------------------------------|
| `CRFEntityExtractor`   | sklearn-crfsuite        | conditional random field            | good for training custom entities      |
| `SpacyEntityExtractor` | spaCy                   | averaged perceptron                | provides pre-trained entities          |
| `DucklingHTTPExtractor`| running duckling        | context-free grammar                | provides pre-trained entities          |
| `MitieEntityExtractor` | MITIE                   | structured SVM                      | good for training custom entities      |
| `EntitySynonymMapper`  | existing entities       | N/A                                | maps known synonyms                    |
| `DIETClassifier`       |                         | conditional random field<br>on top of a transformer | good for training custom entities      |

## [The “entity” Object](https://legacy-docs-v1.rasa.com/1.10.26/nlu/entity-extraction/#the-entity-object)

After parsing, an entity is returned as a dictionary. There are two fields that show information about how the pipeline impacted the entities returned: the `extractor` field of an entity tells you which entity extractor found this particular entity, and the `processors` field contains the name of components that altered this specific entity.

The use of synonyms can cause the `value` field not to match the `text` exactly. Instead, it will return the trained synonym.

```json
{
  "text": "show me chinese restaurants",
  "intent": "restaurant_search",
  "entities": [
    {
      "start": 8,
      "end": 15,
      "value": "chinese",
      "entity": "cuisine",
      "extractor": "CRFEntityExtractor",
      "confidence": 0.854,
      "processors": []
    }
  ]
}
```

Note: The `confidence` will be set by the `CRFEntityExtractor` component. The `DucklingHTTPExtractor` will always return `1`. The `SpacyEntityExtractor` extractor and `DIETClassifier` do not provide this information and return `null`.

## [Custom Entities](https://legacy-docs-v1.rasa.com/1.10.26/nlu/entity-extraction/#custom-entities)

Almost every chatbot and voice app will have some custom entities. A restaurant assistant should understand `chinese` as a cuisine, but to a language-learning assistant, it would mean something very different. The `CRFEntityExtractor` and the `DIETClassifier` component can learn custom entities in any language, given some training data. See [Training Data Format](https://legacy-docs-v1.rasa.com/1.10.26/nlu/training-data-format/#training-data-format) for details on how to include entities in your training data.

## [Entities Roles and Groups](https://legacy-docs-v1.rasa.com/1.10.26/nlu/entity-extraction/#entities-roles-and-groups)

Assigning custom entity labels to words allows you to define certain concepts in the data. For example, we can define what a city is:

```text
I want to fly from [Berlin](city) to [San Francisco](city).
```

However, sometimes you want to specify entities even further. Let’s assume we want to build an assistant that should book a flight for us. The assistant needs to know which of the two cities in the example above is the departure city and which is the destination city.

```text
- I want to fly from [Berlin]{"entity": "city", "role": "departure"} to [San Francisco]{"entity": "city", "role": "destination"}.
```

## [Extracting Places, Dates, People, Organizations](https://legacy-docs-v1.rasa.com/1.10.26/nlu/entity-extraction/#extracting-places-dates-people-organizations)

spaCy has excellent pre-trained named-entity recognizers for a few different languages. You can test them out in this [interactive demo](https://demos.explosion.ai/displacy-ent/).

## [Passing Custom Features to `CRFEntityExtractor`](https://legacy-docs-v1.rasa.com/1.10.26/nlu/entity-extraction/#passing-custom-features-to-crfentityextractor)

If you want to pass custom features, such as pre-trained word embeddings, to `CRFEntityExtractor`, you can add any dense featurizer to the pipeline before the `CRFEntityExtractor`. `CRFEntityExtractor` automatically finds the additional dense features and checks if the dense features are an iterable of `len(tokens)`, where each entry is a vector.
