Custom NLU Components

Custom NLU Components

You can create a custom component to perform a specific task which NLU doesn’t currently offer (for example, sentiment analysis). Below is the specification of the rasa.nlu.components.Component class with the methods you’ll need to implement.

Note

There is a detailed tutorial on building custom components here.

You can add a custom component to your pipeline by adding the module path. So if you have a module called sentiment containing a SentimentAnalyzer class:

pipeline:
- name: "sentiment.SentimentAnalyzer"

Also be sure to read the section on the Component Lifecycle.

To get started, you can use this skeleton that contains the most important methods that you should implement:

<br> 1<br> 2<br> 3<br> 4<br> 5<br> 6<br> 7<br> 8<br> 9<br>10<br>11<br>12<br>13<br>14<br>15<br>16<br>17<br>18<br>19<br>20<br>21<br>22<br>23<br>24<br>25<br>26<br>27<br>28<br>29<br>30<br>31<br>32<br>33<br>34<br>35<br>36<br>37<br>38<br>39<br>40<br>41<br>42<br>43<br>44<br>45<br>46<br>47<br>48<br>49<br>50<br>51<br>52<br>53<br>54<br>55<br>56<br>57<br>58<br>59<br>60<br>61<br>62<br>63<br>64<br>65<br>66<br>67<br>68<br>69<br>70<br>71<br>72<br>73<br>74<br>75<br>76<br>77<br>78<br>79<br>80<br>81<br>82<br>83<br>84<br>85<br>86<br>87<br>88<br>89<br>90<br>91<br>92<br>93<br>94<br> <br>import typing<br>from typing import Any, Optional, Text, Dict, List, Type<br>from rasa.nlu.components import Component<br>from rasa.nlu.config import RasaNLUModelConfig<br>from rasa.nlu.training_data import Message, TrainingData<br>if typing.TYPE_CHECKING:<br> from rasa.nlu.model import Metadata<br>class MyComponent(Component):<br> """A new component"""<br> @classmethod<br> def required_components(cls) -> List[Type[Component]]:<br> return []<br> defaults = {}<br> supported_language_list = None<br> not_supported_language_list = None<br> def __init__(self, component_config: Optional[Dict[Text, Any]] = None) -> None:<br> super().__init__(component_config)<br> def train(self, training_data: TrainingData, config: Optional[RasaNLUModelConfig] = None, **kwargs: Any) -> None:<br> pass<br> def process(self, message: Message, **kwargs: Any) -> None:<br> pass<br> def persist(self, file_name: Text, model_dir: Text) -> Optional[Dict[Text, Any]]:<br> pass<br> @classmethod<br> def load(cls, meta: Dict[Text, Any], model_dir: Optional[Text] = None, model_metadata: Optional["Metadata"] = None, cached_component: Optional["Component"] = None, **kwargs: Any) -> "Component":<br> if cached_component:<br> return cached_component<br> else:<br> return cls(meta)<br>

Note

If you create a custom tokenizer you should implement the methods of rasa.nlu.tokenizers.tokenizer.Tokenizer. The train and process methods are already implemented and you simply need to overwrite the tokenize method.

Note

If you create a custom featurizer you should return a sequence of features. E.g. your featurizer should return a matrix of size (number-of-tokens x feature-dimension). The feature vector of the __CLS__ token should contain features for the complete message.

Component

class rasa.nlu.components.``Component(component_config=None)

A component is a message processing unit in a pipeline.

Components are collected sequentially in a pipeline. Each component is called one after another. This holds for initialization, training, persisting and loading the components.

classmethod required_components()

Specify which components need to be present in the pipeline.

Returns

The list of class names of required components.

classmethod required_packages()

Specify which python packages need to be installed.

classmethod create(component_config, config)

Creates this component (e.g. before training is started).

Parameters

provide_context()

Initialize this component for a new pipeline.

train(training_data, config=None, **kwargs)

Train this component.

Parameters

process(message, **kwargs)

Process an incoming message.

Parameters

persist(file_name, model_dir)

Persist this component to disk for future loading.

Parameters

prepare_partial_processing(pipeline, context)

Sets the pipeline and context used for partial processing.

partially_process(message)

Allows the component to process messages during training (e.g. external training data).

classmethod can_handle_language(language)

Check if the component supports a specific language.