Mapping FAQ with RASA for large dataset (2000+) - Rasa Open Source - Rasa Community Forum

👋 Introducing the Rasa Playground

Mapping FAQ with RASA for large dataset (2000+)

RASA is consist of RASA NLU + Core, I have tested around I understand some part about it. I try to put it into sample practise, and its working perfect.

I plan to bring it into the next level, I wish to create a FAQ system based on RASA stack with help of “tensorflow” backend.

I got over 1200+ pair of Questions and Answers. 1st, NLU would take role to understand and classify the intent along with entity extraction. 2nd it will pass the json response to RASA core where Answers will map or reponse back to the users. It sounds simple, but as I go and check the RASA it give something different. Normally, RASA core will response to the User back based on pre-define story along with ==> “utter_”. Pre-defined story is good, but for small amount of dataset only. we have to write it manually.

How to deal when dataset or Knowledge based is growing larger such as 1000+ or 5000+, We cannot manually mapping it. I try to look around but cannot find any proper way to deal with it yet.

Previously, I used [Retrieval Model] Sklean Tfidf-vectorizer as bags of word along with consine-similairy to compare and return the most similar question index, when index is found Answer will select based on index, but this kind of solution is not effective since the meaning will lost and much more problem.

Anyone got such a good solution for this ??


post by khut on Nov 3, 2018

Rasa is not support such that kind of thing yet. Maybe I just try using RASA NLU to categorized the type of request 1st and return “Intent” back to FAQ database. LSTM (encoder-decoder) will use at knowledge based in order to find the pair questions and answers.

post by souvikg10 on Nov 3, 2018

I don’t think for FAQ, rasa core would help you that much. Unless your FAQ’s are contextual and personalized, it won’t make much sense to use rasa core.

If you catch an intent that links to an answer in a knowledge base, just use if-then-else statements. I would even use elasticsearch and use simple tokenisation pipeline of ELK to find the right answer. It is very powerful.

However, if you would like add context to your conversation, for example a conversation about general questions like product pricing can be answered in two ways


Q - How much an AWS Lightsail subscription cost?
A - You provide a full answer regarding all types of cost

Here you could just use the NLU and use retrieval based dialogue system.

However, the same conversation can be handled as


U - How much an AWS lightsail subscription cost?
B-  How much memory would you look for?
U- 500mb
B- How much core would you need?
U- 1 core
B - It will cost around 3,50$ a month

They are both FAQs, once is simple while the other is contextual where Rasa core brings value.

Now for a very large dataset - i think you are leaning towards Question Answering system- you provide a corpus to a neural network - it learns the embeddings and allows the token to map or point to the right statements in the text. I am not sure how well it works but this is not you can achieve with Rasa at this moment.

post by khut on Nov 4, 2018

Thank you @souvikg10. I think RASA is not support such that kind of thing yet. Maybe I just try using RASA NLU to categorized the type of request 1st and return “Intent” back to FAQ database.

post by znat on Nov 5, 2018

Look at the Rasa Addons FAQ example. I thinks it does what you need (let you manage your KB without having to modify the Core model every time).

post by souvikg10 on Nov 5, 2018

Doesn’t this make your use of rasa core kind of moot? maybe i am wrong.

I mean if NLU detects an intent- i might as well write a custom logic to retrieve the information using code instead because i know it is an FAQ type question.

post by znat on Nov 5, 2018

A project starts with a FAQ and becomes more complicated with contextualized conversations in time, so why not starting with the right stack? Plus FAQ can be asked as side questions in more complex flows, and having all one turn questions grouped makes dealing with those side questions easier.

post by khut on Nov 5, 2018

I got your suggestion, but as I investigate your mentioned RASA add-on FAQ. I start to be puzzled. I have attached one sample image of sheet which consist of Question | Answer | Type. Example:

        Question             |             Answer             |             Type.

1 How to login | Go to login Page | login

2 Reset password | Click on Reset Button | password

I first train Rasa to be able in identify the Type(Intent) of the input, return response as json to CORE. Next is RASA core role, I already got response as “intent” and “score”. now, I want RASA core to able to map the input Question to correct Answer. Possible case, pair of number of FAQ increase up to 2000+. We cannot manually sit and label everything.

post by znat on Nov 6, 2018

Are you saying that the intent is mapped to a topic and not a particular question? How would you identify the right response then within a topic?

post by khut on Nov 6, 2018

It would something that start like this.

1st, I trained Questions to Intent in NLU so that by given particular question --> intent would response correctly. It will pass to the 2nd stage.

2nd stage, Based on given intent, I got 3 options still under-consideration,

  1. Externally, I would like to use some deep learning algorithm such as LSTM, RNN (encoder/decoder) or Siamese, by training them in form of questions and answers pair, so that given particular question then answer would be given. But until now, I couldn’t find any properly proof about how to really implement it.
  2. Externally, It does compare given Question with all Questions in Type or Class, and return index with the most similarity. Index will use to query the Answer. I would get the help from Sklearn, pairwise consine-similarity and using Hot-encode or Word vector presentation (Word2Vec).
  3. Internally, RASA core, as I see potentially. RASA will map the question to answer in form of Story. But we have to write it manually, that limit the capability of large dataset, You have mentioned add-on. I try to check on it but cannot really find a way to figure out my problem yet.

post by znat on Nov 6, 2018

It is certainly ambitious. How about starting from the simplest implementation possible, see how it goes and make a benchmark, and then trying to improve?

post by JoeTorino on Nov 7, 2018

Thanks for the info, I’m trying to develop a deeper understanding of the different ways of using RASA NLU and RASA Core. From what I understand RASA core is used for retaining memory of previous states (context) when needing to carry out further operations related to previous intents, entities and actions. So if you have a question related to the previous one then the computer will be able to understand what you are saying?

In a service chatbot I would assume this to be important.