Feedback on ConveRT Model + Rasa NLU - Rasa Open Source - Rasa Community Forum

👋 Introducing the Rasa Playground

Feedback on ConveRT Model + Rasa NLU

post by dakshvar22 on Nov 15, 2019

Hey, Recently the folks at PolyAI have open-sourced a new sentence encoding model called ConveRT which is pretrained on a large conversational dataset and hence claims better conversational representations over traditional large language models like BERT, etc. The idea resonates very well with what we also believe and have been working on internally at Rasa. Since they open sourced their model as a TFHub model, we decided to build a quick featurizer based on this to extract representations and use them with downstream intent classification models already existing inside Rasa.

In our internal tests, the model does give a significant boost to intent classification accuracy on multiple datasets. We would love the community to try it out and share evaluation numbers on their test sets.

We have released the featurizer as part of Rasa 1.5.0. You can try by doing a pip install of Rasa-


pip install rasa==1.5.0

The featurizer uses an optional dependency tensorflow_text. Install it with -


pip install --no-deps tensorflow_text==1.15.1

Now you are ready to use the ConveRTFeaturizer. In the project directory, we recommend using a config along these lines -


language: en
pipeline:

- name: WhitespaceTokenizer
  - name: ConveRTFeaturizer
  - name: EmbeddingIntentClassifier

This uses the ConveRT model as a feature extractor alone and we do not fine-tune it along with our intent classifier. Please note that this model can only be used for a dataset in the English language. Feel free to change the config params of EmbeddingIntentClassifier according to your dataset. It would be great if you can share the numbers for evaluation metrics on your test dataset here.

post by psds01 on Nov 15, 2019

Is this language specific? i.e. does it only work on English vocabulary words? What happens when the corpus is heavily dominated by Out Of English Vocabulary words?

post by JulianGerhard on Nov 15, 2019

Hi @dakshvar22, as far as I am currently experiencing, this is only possible on Linux machines - Windows seems to not be supported currently. I’ll test it on a Linux machine on Monday and give you feedback on one of our medium datasets with really great quality. Any metric in specific that you are interested in?

post by dakshvar22 on Nov 15, 2019

@psds01 This is only for English language for now. Regarding out of vocabulary words, if you have loads of them and your dataset is dominated by them, this may not help you much but ConveRT does apply a clever trick for handling OOV words. So I would still suggest to give it a try.

post by JulianGerhard on Nov 18, 2019

I did three different experiments for now:

  1. Using the given pipeline on an English dataset with 5 distinct intents, 162 samples, doing a 5-fold cross-validation with the following results:

test Accuracy: 0.916 (0.044)
test F1-score: 0.897 (0.048)
test Precision: 0.882 (0.049)

This result can be seen as really good because the dataset is really tricky in terms of mostly similar sentences leading to different intents.

  1. Using the given pipeline on an English dataset with 52 distinct intents, 3842 samples, doing a 5-fold cross-validation with the following results:

test Accuracy: 0.934 (0.004)
test F1-score: 0.916 (0.003)
test Precision: 0.911 (0.003)
  1. Using the given pipeline and setup on the German version of the set used in experiment #2

test Accuracy: 0.902 (0.004)
test F1-score: 0.902 (0.003)
test Precision: 0.910 (0.003)

post by dakshvar22 on Nov 28, 2019

Thanks for sharing the numbers. Very interesting that simple supervised embedding config is on par or even a bit better in some cases.

post by maulikmadhavi on Nov 27, 2019

Hello all! Here is my result on small size data.

Intent examples: 209 (11 distinct intents)
Entity examples: 80 (4 distinct entities)

Sklearn pipeline

rasa.nlu.test  - train Accuracy: 0.995 (0.002)
rasa.nlu.test  - train F1-score: 0.995 (0.002)
rasa.nlu.test  - train Precision: 0.995 (0.002)
rasa.nlu.test  - test Accuracy: 0.828 (0.075)
rasa.nlu.test  - test F1-score: 0.806 (0.081)
rasa.nlu.test  - test Precision: 0.813 (0.078)

ConveRT pipeline

rasa.nlu.test  - train Accuracy: 1.000 (0.000)
rasa.nlu.test  - train F1-score: 1.000 (0.000)
rasa.nlu.test  - train Precision: 1.000 (0.000)
rasa.nlu.test  - test Accuracy: 0.923 (0.024)
rasa.nlu.test  - test F1-score: 0.919 (0.025)
rasa.nlu.test  - test Precision: 0.937 (0.020)

@matthenWhitespaceTokenizer can be omitted if you want to classify intent; for entity extraction you should include.