LLM Configuration for Rasa Pro ≤ 3.10 | Rasa Documentation
LLM Configuration for Rasa Pro 3.11 and above
For Rasa Pro versions 3.11 and above, refer to the LLM Configuration for >=3.11 page.
Overview
This page applies to the following components which use LLMs:
- SingleStepLLMCommandGenerator
- MultiStepLLMCommandGenerator
- EnterpriseSearchPolicy
- IntentlessPolicy
- ContextualResponseRephraser
- LLMBasedRouter
All the above components can be configured to change:
- the LLM provider
- the model to be used
Starting with version Rasa Pro 3.10, CALM uses LiteLLM under the hood to integrate with different LLM providers. Hence, all LiteLLM's integrated providers are supported with CALM as well. We explicitly mention the settings required for the most frequently used ones in the sections below.
Warning: If you want to try a provider other than OpenAI / Azure OpenAI, it is recommended to install Rasa Pro versions >= 3.10.
Recommended Models
The table below documents the versions of each model we recommend for use with various Rasa components. As new models are published, Rasa will test these and where appropriate add them as a recommended model.
| Component | Providing platform | Recommended models |
|---|---|---|
SingleStepLLMCommandGenerator, EnterpriseSearchPolicy, IntentlessPolicy |
OpenAI, Azure | gpt-4-0613 |
ContextualResponseRephraser |
OpenAI, Azure | gpt-4-0613, gpt-3.5-turbo-0125 |
MultiStepLLMCommandGenerator |
OpenAI, Azure | gpt-4-turbo-2024-04-09, gpt-3.5-turbo-0125, gpt-3.5-turbo-1106, gpt-4o-2024-08-06 |
Chat completion models
Default Provider
CALM is LLM agnostic and can be configured with different LLMs, but OpenAI is the default model provider. Majority of our experiments have been with models available on OpenAI or OpenAI Azure service. The performance of your assistant may vary when using other LLMs, but improvements can be made by tuning flow and collect step descriptions.
To configure components that use a chat completion model as the LLM, declare the configuration under the llm key of that component's configuration. For example:
recipe: default.v1
language: en
pipeline:
- name: SingleStepLLMCommandGenerator
llm:
...
Required Parameters
There are certain required parameters under the llm key:
model- Specifies the name of the model identifier available from the LLM provider's documentation, for e.g.gpt-4-0613provider- Unique identifier of the provider to be used for invoking the specified model.
recipe: default.v1
language: en
pipeline:
- name: SingleStepLLMCommandGenerator
llm:
model: gpt-4-0613
provider: openai
Optional Parameters
The llm key also accepts inference time parameters like temperature, etc which are optional but can be useful in extracting the best performance out of the model being used. Please refer to the official LiteLLM documentation for a list of such parameters supported.
When configuring a particular provider, there are a few provider specific settings which are explained under each provider's individual sub-section below.
Important: If you switch to a different LLM provider, all default parameters for the old provider will be overridden with the default parameters of the new provider.
OpenAI
API Token
The API token authenticates your requests to the OpenAI API.
To configure the API token, follow these steps:
If you haven't already, sign up for an account on the OpenAI platform.
Navigate to the OpenAI Key Management page, and click on the "Create New Secret Key" button to initiate the process of obtaining
<your-api-key>.To set the API key as an environment variable, you can use the following command in a terminal or command prompt:
- Linux/MacOS
export OPENAI_API_KEY=<your-api-key>
- Windows
setx OPENAI_API_KEY <your-api-key>
Replace <your-api-key> with the actual API key you obtained from the OpenAI platform.
Configuration
There are no additional OpenAI specific parameters to be configured. However, there could be model specific parameters like temperature that you might want to modify. Names for such parameters can found in OpenAI's API documentation and defined under llm key of the component's configuration.
Model deprecations
OpenAI regularly publishes a deprecation schedule for its models. This schedule can be accessed in the documentation published by OpenAI.
Azure OpenAI Service
API Token
The API token authenticates your requests to the Azure OpenAI Service.
Set the API token as an environment variable. You can use the following command in a terminal or command prompt:
- Linux/MacOS
export AZURE_API_KEY=<your-api-key>
- Windows
setx AZURE_API_KEY <your-api-key>
Replace <your-api-key> with the actual API key you obtained from the Azure OpenAI Service platform.
Configuration
To access models provided by Azure OpenAI Service, there are a few additional parameters that need to be configured:
provider- Set toazure.api_type- The type of API to use. This should be set to "azure" to indicate the use of Azure OpenAI Service.api_base- The URL for your Azure OpenAI instance. An example might look like this:https://my-azure.openai.azure.com/.api_version- The API version to use for this operation. This follows the YYYY-MM-DD format and the value should be enclosed in single or double quotes.engine/deployment_name- Alias fordeploymentparameter. Name of the deployment on Azure.
Amazon Bedrock
Requirements:
- Make sure you have
rasa-pro>=3.10.xinstalled. - Install
boto3>=1.28.57. - Set the following environment variables -
AWS_ACCESS_KEY_ID,AWS_SECRET_ACCESS_KEY,AWS_REGION_NAME. - (Optional) Might have to set
AWS_SESSION_TOKENif your organisation mandates the usage of temporary credentials for security.
Once the above steps are complete, edit config.yaml to use an appropriate model and set provider to bedrock:
- name: SingleStepLLMCommandGenerator
llm:
provider: bedrock
model: anthropic.claude-instant-v1
Other Providers
Other than the above mentioned providers, we have also tested support for the following providers:
| Platform | provider |
API-KEY variable |
|---|---|---|
| Anthropic | anthropic |
ANTHROPIC_API_KEY |
| Cohere | cohere |
COHERE_API_KEY |
| Mistral | mistral |
MISTRAL_API_KEY |
| Together AI | together_ai |
TOGETHERAI_API_KEY |
| Groq | groq |
GROQ_API_KEY |
Embedding models
To configure components that use an embedding model, declare the configuration under the embeddings key of that component's configuration. For example:
pipeline:
- name: "SingleStepLLMCommandGenerator"
llm:
model: gpt-4-0613
flow_retrieval:
embeddings:
...
The embeddings property needs two mandatory parameters:
model- Specifies the name of the model identifier available from the LLM provider's documentation, for e.g.text-embedding-3-large.provider- Unique identifier of the provider to be used for invoking the specified model, for e.g.openai
Configuring self-signed SSL certificates
In environments where a proxy performs TLS interception, Rasa may need to be configured to trust the certificates used by your proxy. By default, certificates are loaded from the OS certificate store. However, if your setup involves custom self-signed certificates, you can specify these by setting the RASA_CA_BUNDLE environment variable.
This variable points to the path of the certificate file that Rasa should use to validate SSL connections:
export RASA_CA_BUNDLE="path/to/your/certificate.pem"
Configuring Proxy URLs
In environments where LLM requests need to be routed through a proxy, Rasa relies on LiteLLM to handle proxy configurations. LiteLLM supports configuring proxy URLs through the HTTP_PROXY and HTTPS_PROXY environment variables.
To ensure that all LLM requests are routed through the proxy, you can set the environment variables as follows:
export HTTP_PROXY="http://your-proxy-url:port"
export HTTPS_PROXY="https://your-proxy-url:port"
FAQ
Does OpenAI use my data to train their models?
No. OpenAI does not use your data to train their models. From their website:
We do not train our models on your business data by default.
Example Configurations
Azure
A comprehensive example which includes:
llmandembeddingsconfiguration for components inconfig.yml:IntentlessPolicyEnterpriseSearchPolicySingleStepLLMCommandGeneratorflow_retrievalin 3.8.x
llmconfiguration for rephrase inendpoints.yml(ContextualResponseRephraser)
endpoints.yml
nlg:
type: rasa_plus.ml.ContextualResponseRephraser
llm:
engine: rasa-gpt-4
api_type: azure
api_version: "2024-02-15-preview"
api_base: https://my-azure.openai.azure.com
request_timeout: 7
config.yml
recipe: default.v1
language: en
pipeline:
- name: LLMCommandGenerator
llm:
engine: rasa-gpt-4
api_type: azure
api_base: https://my-azure.openai.azure.com/
api_version: "2024-02-15-preview"
request_timeout: 7
policies:
- name: FlowPolicy
- name: rasa_plus.ml.IntentlessPolicy
llm:
engine: rasa-gpt-4
api_type: azure
api_base: https://my-azure.openai.azure.com/
api_version: "2024-02-15-preview"
request_timeout: 7
embeddings:
model: text-embedding-3-small
engine: rasa-embedding-small
api_type: azure
api_base: https://my-azure.openai.azure.com/
api_version: "2024-02-15-preview"
request_timeout: 7
- name: rasa_plus.ml.EnterpriseSearchPolicy
vector_store:
type: "faiss"
threshold: 0.0
llm:
model: gpt-4-0613
engine: rasa-gpt-4
api_type: azure
api_base: https://my-azure.openai.azure.com/
api_version: "2024-02-15-preview"
request_timeout: 7
embeddings:
model: text-embedding-3-small
engine: rasa-embedding-small
api_type: azure
api_base: https://my-azure.openai.azure.com/
api_version: "2024-02-15-preview"
request_timeout: 7