# Entity Extraction

Entity extraction involves parsing user messages for required pieces of information. Rasa Open Source provides entity extractors for custom entities as well as pre-trained ones like dates and locations. Here is a summary of the available extractors and what they are used for:

| Component | Requires | Model | Notes |
| --- | --- | --- | --- |
| `CRFEntityExtractor` | sklearn-crfsuite | conditional random field | good for training custom entities |
| `SpacyEntityExtractor` | spaCy | averaged perceptron | provides pre-trained entities |
| `DucklingHTTPExtractor` | running duckling | context-free grammar | provides pre-trained entities |
| `MitieEntityExtractor` | MITIE | structured SVM | good for training custom entities |
| `EntitySynonymMapper` | existing entities | N/A | maps known synonyms |
| `DIETClassifier` |  | conditional random field<br>on top of a transformer | good for training custom entities |

### [The “entity” Object](https://legacy-docs-v1.rasa.com/1.10.4/nlu/entity-extraction/#id2)

After parsing, an entity is returned as a dictionary. There are two fields that show information about how the pipeline impacted the entities returned: the `extractor` field of an entity tells you which entity extractor found this particular entity, and the `processors` field contains the name of components that altered this specific entity.

The use of synonyms can cause the `value` field not match the `text` exactly. Instead it will return the trained synonym.

```json
{
  "text": "show me chinese restaurants",
  "intent": "restaurant_search",
  "entities": [
    {
      "start": 8,
      "end": 15,
      "value": "chinese",
      "entity": "cuisine",
      "extractor": "CRFEntityExtractor",
      "confidence": 0.854,
      "processors": []
    }
  ]
}
```

Note: The `confidence` will be set by the `CRFEntityExtractor` component. The `DucklingHTTPExtractor` will always return `1`. The `SpacyEntityExtractor` extractor and `DIETClassifier` do not provide this information and return `null`.

Some extractors, like `duckling`, may include additional information. For example:

```json
{
  "additional_info":{
    "grain":"day",
    "type":"value",
    "value":"2018-06-21T00:00:00.000-07:00",
    "values":[
      {
        "grain":"day",
        "type":"value",
        "value":"2018-06-21T00:00:00.000-07:00"
      }
    ]
  },
  "confidence":1.0,
  "end":5,
  "entity":"time",
  "extractor":"DucklingHTTPExtractor",
  "start":0,
  "text":"today",
  "value":"2018-06-21T00:00:00.000-07:00"
}
```

### [Custom Entities](https://legacy-docs-v1.rasa.com/1.10.4/nlu/entity-extraction/#id3)

Almost every chatbot and voice app will have some custom entities. A restaurant assistant should understand `chinese` as a cuisine, but to a language-learning assistant it would mean something very different. The `CRFEntityExtractor` and the `DIETClassifier` component can learn custom entities in any language, given some training data. See [Training Data Format](https://legacy-docs-v1.rasa.com/1.10.4/nlu/training-data-format/#training-data-format) for details on how to include entities in your training data.

### [Entities Roles and Groups](https://legacy-docs-v1.rasa.com/1.10.4/nlu/entity-extraction/#id4)

Assigning custom entity labels to words, allow you to define certain concepts in the data. For example, we can define what a city is:

```text
I want to fly from [Berlin](city) to [San Francisco](city).
```

However, sometimes you want to specify entities even further.

```text
- I want to fly from [Berlin]{"entity": "city", "role": "departure"} to [San Francisco]{"entity": "city", "role": "destination"}.
```

You can also group different entities by specifying a group label next to the entity label. The group label can, for example, be used to define different orders. In the following example we use the group label to reference what toppings goes with which pizza and what size which pizza has.

```text
Give me a [small]{"entity": "size", "group": "1"} pizza with [mushrooms]{"entity": "topping", "group": "1"} and a [large]{"entity": "size", "group": "2"} [pepperoni]{"entity": "topping", "group": "2"}
```

See [Training Data Format](https://legacy-docs-v1.rasa.com/1.10.4/nlu/training-data-format/#training-data-format) for details on how to define entities with roles and groups in your training data.

The entity object returned by the extractor will include the detected role/group label.

```json
{
  "text": "Book a flight from Berlin to SF",
  "intent": "book_flight",
  "entities": [
    {
      "start": 19,
      "end": 25,
      "value": "Berlin",
      "entity": "city",
      "role": "departure",
      "extractor": "DIETClassifier"
    },
    {
      "start": 29,
      "end": 31,
      "value": "San Francisco",
      "entity": "city",
      "role": "destination",
      "extractor": "DIETClassifier"
    }
  ]
}
```

### [Extracting Places, Dates, People, Organizations](https://legacy-docs-v1.rasa.com/1.10.4/nlu/entity-extraction/#id5)

spaCy has excellent pre-trained named-entity recognizers for a few different languages. You can test them out in this [interactive demo](https://demos.explosion.ai/displacy-ent/).

### [Dates, Amounts of Money, Durations, Distances, Ordinals](https://legacy-docs-v1.rasa.com/1.10.4/nlu/entity-extraction/#id6)

The [duckling](https://duckling.wit.ai/) library does a great job turning expressions like “next Thursday at 8pm” into actual datetime objects that you can use.

### [Regular Expressions (regex)](https://legacy-docs-v1.rasa.com/1.10.4/nlu/entity-extraction/#id7)

You can use regular expressions to help the CRF model learn to recognize entities.

### [Passing Custom Features to `CRFEntityExtractor`](https://legacy-docs-v1.rasa.com/1.10.4/nlu/entity-extraction/#id8)

If you want to pass custom features, such as pre-trained word embeddings, to `CRFEntityExtractor`, you can add any dense featurizer to the pipeline before the `CRFEntityExtractor`. `CRFEntityExtractor` automatically finds the additional dense features and checks if the dense features are an iterable of `len(tokens)`, where each entry is a vector.
