Training Data Importers
These docs are for version 1.x of Rasa Open Source.
Training Data Importers
By default, you can use command line arguments to specify where Rasa should look for training data on your disk. Rasa then loads any potential training files and uses them to train your assistant.
If needed, you can also customize how Rasa imports training data. Potential use cases for this might be:
- using a custom parser to load training data in other formats
- using different approaches to collect training data (e.g. loading them from different resources)
You can instruct Rasa to load and use your custom importer by adding the section importers to the Rasa configuration file and specifying the importer with its full class path:
importers:
- name: "module.CustomImporter"
parameter1: "value"
parameter2: "value2"
- name: "module.AnotherCustomImporter"
The name key is used to determine which importer should be loaded. Any extra parameters are passed as constructor arguments to the loaded importer.
Note: You can specify multiple importers. Rasa will automatically merge their results.
RasaFileImporter (default)
By default Rasa uses the importer RasaFileImporter. If you want to use it on its own, you don’t have to specify anything in your configuration file. If you want to use it together with other importers, add it to your configuration file:
importers:
- name: "RasaFileImporter"
MultiProjectImporter (experimental)
This feature is currently experimental and might change or be removed in the future.
With this importer you can build a contextual AI assistant by combining multiple reusable Rasa projects. You might, for example, handle chitchat with one project and greet your users with another. These projects can be developed in isolation, and then combined at train time to create your assistant.
An example directory structure could look like this:
.
├── config.yml
└── projects
├── GreetBot
│ ├── data
│ │ ├── nlu.md
│ │ └── stories.md
│ └── domain.yml
└── ChitchatBot
├── config.yml
├── data
│ ├── nlu.md
│ └── stories.md
└── domain.yml
In this example the contextual AI assistant imports the ChitchatBot project which in turn imports the GreetBot project. Project imports are defined in the configuration files of each project.
To instruct Rasa to use the MultiProjectImporter module, put this section in the config file of your root project:
importers:
- name: MultiProjectImporter
Then specify which projects you want to import.
Writing a Custom Importer
If you are writing a custom importer, this importer has to implement the interface of TrainingDataImporter:
from typing import Optional, Text, Dict, List, Union
import rasa
from rasa.core.domain import Domain
from rasa.core.interpreter import RegexInterpreter, NaturalLanguageInterpreter
from rasa.core.training.structures import StoryGraph
from rasa.importers.importer import TrainingDataImporter
from rasa.nlu.training_data import TrainingData
class MyImporter(TrainingDataImporter):
"""Example implementation of a custom importer component."""
def __init__(
self,
config_file: Optional[Text] = None,
domain_path: Optional[Text] = None,
training_data_paths: Optional[Union[List[Text], Text]] = None,
**kwargs: Dict
):
"""Constructor of your custom file importer."
pass
async def get_domain(self) -> Domain:
path_to_domain_file = self._custom_get_domain_file()
return Domain.load(path_to_domain_file)
# Additional methods follow...
commonrasa.importers.importer.TrainingDataImporter provides an interface for loading training data.
asyncget_domain - Retrieves the domain of the bot. Returns Loaded Domain.
asyncget_config - Retrieves the configuration that should be used for training. Returns The configuration as dictionary.
asyncget_nlu_data(language='en') - Retrieves the NLU training data. Parameters: language – Can be used to only load training data for a certain language. Returns Loaded NLU TrainingData.
asyncget_stories(...) - Retrieves the stories that should be used for training. Returns StoryGraph containing all loaded stories.