I’m having difficulty in pre-processing the training dataset of 40000 reviews of text with max length of review exceeding 5000 words. Which approach should I follow to get embedding layer output from text input?
Challenge -Movie Rating Prediction using RNN
Hey @preetishvij, you need to train your model in batches, or use something like train_generator. One way is to use pretrained embeddings like glove etc. In this case. You can refer this link,
https://stanford.edu/~shervine/blog/keras-how-to-generate-data-on-the-fly
Hope this resolved your doubt.
Plz mark the doubt as resolved.
Do I need to apply lemmatization and stopwords removal on the dataset. I wonder it will remove important words like ‘not’ from review and change the sentiment.
I was also thinking of using glove or word2vec. Correct me if i’m wrong, it will change my 2D text review like (40000,500) into 3D num matrix like (40000,500,50) which will be by input for each RNN cell,… where 40000 = no of reviews in my dataset, 500 = max len of each review I’m deciding and 50 is the size of compact vector for each word given by glove or word2vec.
But it require trimming of review to some maxlen like 500 as defined above. Will it not lead to some essential data loss for my model training as this dataset contains more than 1000 words for some reviews.
Hey @preetishvij, you could prefer to not remove stopwords, or you can remove stopwords but leave ‘not’ in the reviews. You can skip lemmatization part as word embedding will contain the word itself.
Second part, yes model input will be (40000,500,50) for one example, reason has already been well explained by you.
Third part, if you will plot graph of length of reviews and their frequency, you will see there are very less reviews of length greater than 500. Its sure some of the information may be lost, but we can’t afford to increase computation to much greater level for these few examples. So better keep it 500.
Hope this resolved your doubt.
Plz mark the doubt as resolved.
https://colab.research.google.com/drive/1Q1Wi4Rni1RGy6SKZxdgxfyN8yqXNjPj6
I tried the above approach but my training accuracy is very bad (~52%). I used the same model used on Imdb dataset. I just use pretrained word embedding. What is the problem?
Hey @preetishvij, extremely sorry for the late reply from my end, actually the thread just got skipped from my notifications.
If you find your training accuracy is very low, this means that model is very simple and we need to increase the complexity of the modely, by adding more number of layers etc. So try again by adding some layers, also add activation to layers as well.
Hope this resolved your doubt.
Plz mark the doubt as resolved. 
I hope I’ve cleared your doubt. I ask you to please rate your experience here
Your feedback is very important. It helps us improve our platform and hence provide you
the learning experience you deserve.
On the off chance, you still have some questions or not find the answers satisfactory, you may reopen
the doubt.