End-to-end Training [Experimental] - Announcements / Important Updates - Rasa Community Forum
End-to-end Training [Experimental]
post by Emma on Dec 17, 2020
Hey Rasa @community,
One year ago we wrote that it’s about time we get rid of intents, and about how we see a future beyond the limitations of them. In order to build level 5 assistants, we believe that we should not remain stuck in the mindset that every user message has to neatly fit into one of our predefined intents.
With Rasa Open Source 2.2, we released a new experimental feature called end-to-end training, it allows you to train the dialogue policy directly on user text without separate NLU data. So instead of a two step process (an NLU prediction followed by a dialogue policy choosing the next action), Rasa can now directly predict the next action the bot should take by looking at the message the user sent.
As you work on an assistant over time to make it more sophisticated, end-to-end learning allows you to keep evolving and improving without being limited by a rigid set of intents. The benefit of this approach is that it makes intents optional.
To get a more in depth understanding of how End-to-End Training works in Rasa Open Source, check our latest blog post:
It’s important to note that, since this is an experimental feature, we don’t have full support yet across all Rasa features like interactive learning, or Rasa X.
For now, think of end-to-end learning as a feature for advanced teams who want to push the limits of what Rasa Open Source can do. This has been a massive joint effort from our research and engineering teams, and we believe it’s a major piece of the puzzle towards better conversational AI.
Examples:
- How to let users use the latest intent without explicitly ask for it
- The inform intent will be the death of me
- Any way to recreate BinarySingleStateFeaturizer performance with TEDPolicy?
- DENSE_FEATURIZABLE_ATTRIBUTES
post by tatianaf on Dec 17, 2020
This looks like a really exciting feature! Can’t wait to experiment with this next year!
post by inthematrix on Dec 17, 2020
As part of We’re a step closer to getting rid of intents, the training example towards the end of the post:
version: "2.0"
stories:
- story: end to end happy path
steps:
- user: “hi”
- bot: “hi!”
- user: “I’m looking for a restaurant”
- bot: “how about Chinese food?”
- user: “sure”
- bot: “here’s what I found ...”
Was wondering how to extract entities like cuisine (chinese in this case)?
post by sebastian on Dec 17, 2020
Great progress, this sounds very interesting. Especially for handling ambiguous, context-sensitive utterances. One question for now: Let’s say I have more than one concrete user utterance I would like to use in story at a specific point as an alternative to an intent. Do I in this case create two (or more) completely separate stories or can I somehow combine the concrete utterance variants in one single story?
post by Nasnl on Dec 17, 2020
Sounds super exciting. I’m just about to start the design of a new bot next week. Not super critical, so would you say that I could best start with this feature right away rather than first using the ‘traditional way’ and later rebuild?
post by aymen on Dec 18, 2020
exciting
post by Ghostvv on Dec 18, 2020
Please take a look at the docs for how to create stories for e2e training: Training Data Format
You can mark entities in user text in the same way, you mark entities in the NLU data
version: "2.0"
stories:
- story: end to end happy path
steps:
- user: “hi”
- bot: “hi!”
- user: “I’m looking for a restaurant”
- bot: “how about [Chinese](cuisine) food?”
- user: “sure”
- bot: “here’s what I found ...”
post by Ghostvv on Dec 18, 2020
you need to create as many stories as you have utterances. This is why we keep intents, because quite often, if you can imagine a lot of such utterances, it is better to create an intent for it
post by Ghostvv on Dec 18, 2020
I don’t suggest to use it in production. I’d recommend to do it in a ‘traditional way’. Why do you want to rebuild it later?
post by Nasnl on Dec 18, 2020
Thanks @Ghostvv. At first I didn’t pick up on that intents will still exist next to it; that’s clear now and no rebuild is required to ‘a next version’.
post by Ghostvv on Dec 18, 2020
glad to clarify. Intents are very useful, but we expect e2e to be useful as well, we’re trying to find a “mixed” approach (you can have intents and texts in the same story, but in different dialogue turns) in the same sense as we implemented rules so that they work together with ML.
post by kearnsw on Dec 18, 2020
I am wondering if you all have thought of a hybrid approach instead of deprecating intents in later releases. That is, you train the model end-to-end, but use scaffolding as with the MOSS framework that trains on all available data (using them as intermediate outputs):
In this way, the user can provide data in the end-to-end format, traditional format, or a mixture.
post by Ghostvv on Dec 21, 2020
as I said above, we’re not deprecating intents, you can have hybrid stories:
stories:
- story: full end-to-end story
steps:
- intent: greet
entities:
- name: Ivan
- bot: Hello, a person with a name!
- intent: search_restaurant
- action: utter_suggest_cuisine
- user: I can always go for [sushi](cuisine)
- bot: Personally, I prefer pizza, but sure let's search sushi restaurants
- action: utter_suggest_cuisine
- user: Have a beautiful day!
- action: utter_goodbye
post by Ghostvv on Dec 21, 2020
in a sense our current approach with intents is a modular supervision approach
post by Ghostvv on Dec 21, 2020
with our hybrid e2e approach, we try to solve the problem when intermediate labels are unknown, therefore there is no supervision signal from intent labels at all. And we implement it so that you can have all the different mixtures.
post by DeqianBai on Dec 23, 2020
what the meaning of " user turns"?
post by Rajendra9 on Dec 30, 2020
when are you going to plan this to do practically.
post by Ghostvv on Jan 5, 2021
“user turn” - is the dialogue turn that contain an input from the user
post by Ghostvv on Jan 5, 2021
what do you mean? end-to-end training is released in rasa open source 2.2
post by Aspirinkb on Jan 12, 2021
Great!
post by mikeymms on Feb 7, 2021
Hi there, If one builds a bot entirely based on pre-existing live chat messages, would it make sense to:
- “Clean” these conversations (remove redundancies, non useful paths etc)
- use e2e training only
- introduce intent later only if it saves significant training time?
Thank you!
post by vikrant67 on Feb 15, 2021
I have 400 pure e2e stories from historic conversational data. I am getting OOM error while training core ( rasa train --augmentation 0 ). I am using 1 GPU with 12 GB RAM. I am using this config.yml.
pipeline:
- name: WhitespaceTokenizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: "char_wb"
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
epochs: 100
- name: EntitySynonymMapper
- name: ResponseSelector
epochs: 100
policies:
- name: TEDPolicy
epochs: 10
max_history: 5
- name: RulePolicy
I have tried changing batch_size but nothing helped.
PS. Its working file for smaller data set. It worked fine for 50 e2e stories.
post by TomV on Jul 1, 2021
@vikrant67 Were you able to resolve this challenge with getting your OOM error? I’m interested in trying a similar experiment, and would love to learn from your experience so far.
post by dingusagar on Jul 21, 2021
was trying out this feature, but got an error during the training stage of rasa core
File "/media/dingusagar/rasa_2_8/lib/python3.7/site-packages/rasa/core/featurizers/tracker_featurizers.py", line 960, in <listcomp>
[intent for intent in tracker_intents]
ValueError: 'my order is late' is not in list
My stories look like this:
- story: Order late
steps:
- user: “my order is late”
- action: utter_sorry_to_hear
I am using rasa 2.8, the config is the default one in rasa 2.8. from the error i feel like the user utterance is treated like an intent and its complaining that such intent is not present. Could someone help me understand what am i doing wrong?
post by inthematrix on Aug 9, 2021
What does your config.yml look like?
post by dingusagar on Aug 13, 2021
the config is the default one in rasa 2.8. no changes made to it.
language: en
pipeline:
- name: WhitespaceTokenizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: char_wb
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
epochs: 100
constrain_similarities: true
- name: EntitySynonymMapper
- name: ResponseSelector
epochs: 100
constrain_similarities: true
policies:
- name: MemoizationPolicy
- name: RulePolicy
- name: UnexpecTEDIntentPolicy
max_history: 5
epochs: 100
- name: TEDPolicy
max_history: 5
epochs: 100
constrain_similarities: true
post by j.mosig on Sep 9, 2021
@dingusagar Let’s deal with your question in this separate thread.
@tatianaf @inthematrix @sebastian @Nasnl @aymen @kearnsw @DeqianBai @Rajendra9 @Aspirinkb @mikeymms @vikrant67 I’d be interested in hearing about your experiences so far. In particular, what kinds of “conversation phenomena” do you struggle to solve with intents where e2e is useful? (e.g. “multi-intents”, “sarcasm”, “long user inputs”, etc.)
post by TomV on Sep 27, 2021
For me, the huge potential for e2e is the ability to use big historical chat logs of Human 2 Human chat conversations as a basis for training.
This is really interesting, and huge potential. Especially combined with some “power” annotation tools like Prodigy from Explosion.ai ( https://explosion.ai/software#prodigy )
post by yadap on Jan 11, 2022
Has anyone experimented with this feature successfully? My main concerns about e2e training is in controllability - the model failing to learn how to handle various flows, especially for complicated dialogues that have multiple branches and subflows depending on slot values and such.
post by bhanrobert99 on May 4, 2022
Let’s say I have more than one concrete user utterance I would like to use in story at a specific point as an alternative to an intent.
post by HermanH on Jun 9, 2022
I think end-to-end training is a great concept, as classifying user utterances as intents and dialogues as stories is a hard job. Especially, when you’ve got a huge mass of real dialogues, with a lot of variations. In my opinion, Rasa’s force is the ability to train dialogue models on real dialogues, not only train the NLU!!
I also believe, end-to-end training appears to be a solution for the multiple intents problems and for user utterances with meanings depending on the context in dialogue history.
When applying CDD and to be in control, SME needs a tool like (former) Rasa X to manage training and test data.
So, I’m wondering, when Rasa Enterprise is going to support end-to-end training? This should include: talking to the bot with mixed stories - all combinations of {intents, responses/actions, end-to-end} -, analyse and annotate, save as either training story or as test story.
Note: I’m aware of Alans news about stopping Rasa X community edition.
Beside this news, my organisation is still in exploring phase, when concerning Conversational Agents and choosing a platform.
As I’m very exited on both Rasa Open Source and Rasa X, Rasa Enterprise could be a candidate.
So, the answer to this question is of great importance in our decision making.
post by Jiaheng212 on Aug 11, 2022
i have a problem when used end to end models like this.
whats wrong?
post by sibbsnb on Aug 30, 2022
is this feature in invest?
post by HermanH on Sep 21, 2022
Still wondering, when this feature will be available in #Rasa#Enterprise.
And if it is going to beat other CAI platforms!
@amn41 Do you have the answer?