Entity Extraction

These docs are for version 1.x of Rasa Open Source.

User Guide

NLU

Core

Conversation Design

API Reference

Migrate from (beta)

Reference

Versions

viewing: 1.10.0

Entity Extraction

Entity extraction involves parsing user messages for required pieces of information. Rasa Open Source provides entity extractors for custom entities as well as pre-trained ones like dates and locations. Here is a summary of the available extractors and what they are used for:

Component Requires Model Notes
CRFEntityExtractor sklearn-crfsuite conditional random field good for training custom entities
SpacyEntityExtractor spaCy averaged perceptron provides pre-trained entities
DucklingHTTPExtractor running duckling context-free grammar provides pre-trained entities
MitieEntityExtractor MITIE structured SVM good for training custom entities
EntitySynonymMapper existing entities N/A maps known synonyms
DIETClassifier conditional random field
on top of a transformer
good for training custom entities

The “entity” Object

After parsing, an entity is returned as a dictionary. There are two fields that show information about how the pipeline impacted the entities returned: the extractor field of an entity tells you which entity extractor found this particular entity, and the processors field contains the name of components that altered this specific entity.

The use of synonyms can cause the value field not match the text exactly. Instead it will return the trained synonym.

{
  "text": "show me chinese restaurants",
  "intent": "restaurant_search",
  "entities": [\
    {\
      "start": 8,\
      "end": 15,\
      "value": "chinese",\
      "entity": "cuisine",\
      "extractor": "CRFEntityExtractor",\
      "confidence": 0.854,\
      "processors": []\
    }\
  ]
}

Note

The confidence will be set by the CRFEntityExtractor component. The DucklingHTTPExtractor will always return 1. The SpacyEntityExtractor extractor and DIETClassifier do not provide this information and return null.

Some extractors, like duckling, may include additional information. For example:

{
  "additional_info":{
    "grain":"day",
    "type":"value",
    "value":"2018-06-21T00:00:00.000-07:00",
    "values":[\
      {\
        "grain":"day",
        "type":"value",
        "value":"2018-06-21T00:00:00.000-07:00"\
      }\
    ]
  },
  "confidence":1.0,
  "end":5,
  "entity":"time",
  "extractor":"DucklingHTTPExtractor",
  "start":0,
  "text":"today",
  "value":"2018-06-21T00:00:00.000-07:00"
}

Custom Entities

Almost every chatbot and voice app will have some custom entities. A restaurant assistant should understand chinese as a cuisine, but to a language-learning assistant it would mean something very different. The CRFEntityExtractor and the DIETClassifier component can learn custom entities in any language, given some training data.

Entities Roles and Groups

Assigning custom entity labels to words, allow you to define certain concepts in the data.

I want to fly from [Berlin](city) to [San Francisco](city).

However, sometimes you want to specify entities even further. Let’s assume we want to build an assistant that should book a flight for us. The assistant needs to know which of the two cities in the example above is the departure city and which is the destination city.

- I want to fly from [Berlin]{"entity": "city", "role": "departure"} to [San Francisco]{"entity": "city", "role": "destination"}.

To fill slots from entities with a specific role/group, you need to either define a custom slot mappings using Forms or use Custom Actions to extract the corresponding entity directly from the tracker.

Extracting Places, Dates, People, Organisations

spaCy has excellent pre-trained named-entity recognisers for a few different languages.

Dates, Amounts of Money, Durations, Distances, Ordinals

The duckling library does a great job of turning expressions like “next Thursday at 8pm” into actual datetime objects that you can use, e.g.

"next Thursday at 8pm"
=> {"value":"2018-05-31T20:00:00.000+01:00"}

Regular Expressions (regex)

You can use regular expressions to help the CRF model learn to recognize entities.

Passing Custom Features to CRFEntityExtractor

If you want to pass custom features, such as pre-trained word embeddings, to CRFEntityExtractor, you can add any dense featurizer to the pipeline before the CRFEntityExtractor.CRFEntityExtractor automatically finds the additional dense features and checks if the dense features are an iterable of len(tokens), where each entry is a vector.

👋 I can help you get started with Rasa and answer your technical questions.