Rasa classifies random input as intents with high probability - Rasa Open Source - Rasa Community Forum
post by Jakub on Jul 30, 2021
Hello, I have trained my model using pipeline:
- name: "DucklingHTTPExtractor" url: "http://localhost:8000" dimensions: ["duration"]
- name: WhitespaceTokenizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer analyzer: char_wb min_ngram: 1 max_ngram: 4
- name: DIETClassifier epochs: 100
- name: EntitySynonymMapper
- name: ResponseSelector epochs: 100
- name: FallbackClassifier threshold: 0.3 ambiguity_threshold: 0.1
However random words like ‘asdf’, ‘qwerty’ and even single letters are classified as intents (‘greet’, ‘neutral’ respectively). I could increase FallbackClassifier threshold, but some examples have confidence over 0.9. I’m using Rasa 2.8. How to deal with it?
post by nik202 on Jul 30, 2021
@Jakub Hi! Can you please share some example or file or screenshot?
@Jakub Why you are using ambiguity_threshold: 0.1 inside FallbackClassifier or share the link from where you get this idea?
@Jakub Please share complete config.yml or update the above one.
post by thorty on Jul 30, 2021
@Jakub Same on my site. No Matter what kind of input. There is always a intent recognition. Eg “123456”, “dfjjaspfj” or something like “i want pizza”.
I’m looking forward for some help. Thanks!
Using different versions with docker:
- rasa/rasa:main-spacy-de
- rasa/rasa:2.8.1-spacy-de
pipeline: language: de
pipeline:
- name: SpacyNLP
model: de_core_news_sm
- name: SpacyTokenizer
- name: SpacyFeaturizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer
analyzer: char_wb
min_ngram: 1
max_ngram: 4
- name: DIETClassifier
epochs: 100
constrain_similarities: true
- name: EntitySynonymMapper
- name: FallbackClassifier
threshold: 0.3
ambiguity_threshold: 0.1
nlu.yml: version: "2.0"
nlu:
- intent: greet
examples: |
- hey
- hallo
- hi
- hallo du
- guten morgen
- guten abend
- morgen
- guten tag
- intent: goodbye
examples: |
- tschüss
- wiederhören
- ciao
- bis dann
- bye
result:
{
"text": "12345",
"intent": {
"id": 8761605359927853060,
"name": "goodbye",
"confidence": 0.9780691862106323
},
"entities": [],
"intent_ranking": [
{
"id": 8761605359927853060,
"name": "goodbye",
"confidence": 0.9780691862106323
},
{
"id": -3994034142759590561,
"name": "greet",
"confidence": 0.02193082869052887
}
]
}
post by nik202 on Jul 30, 2021
@thorty Means whatever input i.e “I want pizza” , “I need drinks” do such examples are in your nlu.yml file?; it’s giving you goodbye output?
post by thorty on Jul 30, 2021
@nik202: No they are not into my nlu.yml as examples. In this case I’m expecting a result like: intent: none. But what I’m getting here is something like intent: goodbye with a conf of over 90.
my nlu.yml is exactly like in the post above. the pipeline also. and the result is an example with this config.
thank you!
post by rctatman on Jul 30, 2021
@thorty All your inputs will be given an intent. If you have very high confidence and a lot of incorrect classifications it’s often the case that this is due to your data not being a good representation of what your assistant is seeing in production.
You should create an out of scope intent if you want to be able to capture things your assistant can’t do. This page goes over how to do it: Fallback and Human Handoff
post by thorty on Jul 31, 2021
@Nik202: Thanks, but this does not change anything.
@rctatman: Thanks for your explanation and the Link to the right Step in Documentation. This helps a lot in understanding rasa.
In my case I do not really know the utterances from users. So I will start with a small set of examples and train the NLU after going live with real-world examples. To handle the “non” trained utterances the right way it would be better to get no intent rather than the wrong intent. Especially when utterances are non-sense like in my example. Is there any chance to achieve this?
post by nik202 on Jul 31, 2021
@thorty ok
post by nik202 on Jul 31, 2021
@Jakub Any progress about your error?
post by Jakub on Aug 2, 2021
@nik202 sorry for late response. I removed ambiguity_threshold: 0.1 from my config file (the one I sent in first message is my whole config.yml, I have default policies), but there’s only a slight difference (a little bit more of inputs are classified as nlu_fallback). I upload examples and neutral intent below.
post by thorty on Aug 2, 2021
Short Update from my side:
when I also use the KeywordIntentClassifier in my pipline the NLU response as follows:
"intent": {
"name": null,
"confidence": 0.0
},
pipeline
- name: KeywordIntentClassifier
But I think that makes the DIET Classifier with all its advantages useless. ?!?
post by thorty on Aug 4, 2021
Short Update:
Like @rctatman explained. With more training data I get better confidence values for my utterances. Also for “nonsense” input like “askjfnqiurz”.
post by SowmyaBalam on Oct 12, 2021
hi, I am also facing the same issue, for whatever nonsense query "abcdefghijkl’ it is going to greet with high confidence of 0.99 in diet classifier. can anybody help me resolve this issue.
post by nik202 on Oct 12, 2021
@SowmyaBalam delete the older trained model first, sometimes it helps and did you mention the default fallback, please share some supporting files, so that we can see.
post by SowmyaBalam on Oct 12, 2021
@nik202 this is the config file. i am using:
language: en
pipeline:
- name: WhitespaceTokenizer
- name: RegexFeaturizer
- name: LexicalSyntacticFeaturizer
- name: CountVectorsFeaturizer
- name: CountVectorsFeaturizer analyzer: char_wb min_ngram: 1 max_ngram: 4
- name: DIETClassifier epochs: 300 evaluate_on_number_of_examples: 80 evaluate_every_number_of_epochs: 5 tensorboard_log_directory: "tensorboard" tensorboard_log_level: "epoch" checkpoint_model: True constrain_similarities: True
- name: EntitySynonymMapper
policies:
- name: MemoizationPolicy
- name: TEDPolicy epochs: 300 evaluate_on_number_of_examples: 80 evaluate_every_number_of_epochs: 5 tensorboard_log_directory: "tensorboard" tensorboard_log_level: "epoch" checkpoint_model: True constrain_similarities: True
- name: RulePolicy
post by nik202 on Oct 12, 2021
@SowmyaBalam can you please format the code using "</>" or """ """ ?
post by SowmyaBalam on Oct 12, 2021
language: en
policies:
- name: MemoizationPolicy
- name: TEDPolicy max_history: 5 epochs: 300 evaluate_on_number_of_examples: 80 evaluate_every_number_of_epochs: 5 tensorboard_log_directory: "tensorboard" tensorboard_log_level: "epoch" checkpoint_model: True constrain_similarities: True
- name: RulePolicy
post by nik202 on Oct 12, 2021
@SowmyaBalam Hey! why there is a \ (backward slash) ? in front of the pipelines and policies ?
post by Harvey1976 on May 26, 2022
any update on this thread? I face exactly the same situation here when we use rasa 2.8 to replace the rasa nlu 0.13.7.
post by vapdev on Aug 2, 2022
Facing the same situation… I can’t believe no one knows a good way to handle this
post by vapdev on Aug 2, 2022
Actually, this shouldn’t even happen! I type ‘.’ and I get: Received user message ‘.’ with intent ‘{‘name’: ‘affirm’, ‘confidence’: 0.9790935516357422}’ and entities ‘[]’… Intent Affirm??? Really??? A full stop???
post by vapdev on Aug 2, 2022
I know I can make an out of scope intent or trash intent with this kind of message. But I can’t cover every single out of scope word… Wtf
post by AlTester on Apr 20, 2023
Any updates on this? Getting the same issue.