{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Sentiment Analysis with RNNs\n============================\n\nThis notebook is part of a series:\n\n- Sentiment Analysis with RNNs - Disaster Tweets \n- Sentiment Analysis with RNNs - TripAdvisor Reviews\n- Sentiment Analysis with RNNs - News Headlines Sarcasm\n\nIn this notebook we reviewed some techniques to perform sentiment analysis on a text dataset. \nWe have found 4 different datasets on which we could conduct a sentiment analysis, each one containing a text and a sentiment (positive or negative).\n\n- [IMDB Movies Dataset](https://www.kaggle.com/datasets/lakshmi25npathi/imdb-dataset-of-50k-movie-reviews)\n- [Disaster tweets dataset](https://www.kaggle.com/competitions/nlp-getting-started)\n- [Tripadvisor reviews](https://www.kaggle.com/datasets/andrewmvd/trip-advisor-hotel-reviews)\n- [News headlines sarcasm detection](https://www.kaggle.com/datasets/rmisra/news-headlines-dataset-for-sarcasm-detection)\n\nThis problem can be formulated as a binary classification of a text, based on the sentiment associated to each word and their structure in the sentence.\n\n<a id=\"toc\">Table of Contents</a>\n\n1. [Data Analysis](#1)  \n    [a. Dataset info](#1a)  \n    [b. Dirty characters](#1b)\n2. [Text Preprocessing](#2a)\n3. [Model Experiments](#3)  \n    [a. Simple FFN](#3a)  \n    [b. Simple RNN](#3b)  \n    [c. LSTM](#3c)  \n    [d. GRU](#3d)  \n    [e. 2 Layers LSTM](#3e)  \n4. [Embedding Experiments](#4)  \n    [a. LSTM - GloVe 6B300D](#4a)  \n    [b. GRU - GloVe 6B300D](#4b)  \n    [c. LSTM - GloVe Twitter 27B200D](#4c)  \n    [d. GRU - GloVe 27B200D](#4d)  \n5. [Contextual Embeddings](#5)","metadata":{"execution":{"iopub.execute_input":"2022-04-19T19:34:33.506616Z","iopub.status.busy":"2022-04-19T19:34:33.506084Z","iopub.status.idle":"2022-04-19T19:34:34.098425Z","shell.execute_reply":"2022-04-19T19:34:34.097684Z","shell.execute_reply.started":"2022-04-19T19:34:33.506569Z"},"papermill":{"duration":0.02384,"end_time":"2022-07-16T14:59:28.565656","exception":false,"start_time":"2022-07-16T14:59:28.541816","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow import keras\nimport pandas as pd\nimport numpy as np\n\nfrom tensorflow.keras.preprocessing.text import Tokenizer\nfrom tensorflow.keras.preprocessing.sequence import pad_sequences\nfrom sklearn.model_selection import train_test_split\n\n#Uncomment the desired dataset\n\n##IMDB dataset\n#imdb_path = '../input/imdb-dataset-of-50k-movie-reviews/IMDB Dataset.csv'\n# data = pd.read_csv(imdb_path)\n# X = data['review']\n# y = pd.Series(np.where(data['sentiment'].str.contains(\"positive\"), 1, 0))\n\n##Disaster tweets dataset\ndis_path = '../input/nlp-getting-started/train.csv'\ndata  = pd.read_csv(dis_path)\nX = data['text']\ny = data['target']\n\n##TripAdvisor reviews\n# trip_path = '../input/trip-advisor-hotel-reviews/tripadvisor_hotel_reviews.csv'\n# data = pd.read_csv(trip_path)\n# X = data['Review']\n# y = pd.Series(np.where(data['Rating'] >3, 1, 0))\n\n##News headlines sarcasm detection\n# news_path = '../input/news-headlines-dataset-for-sarcasm-detection/Sarcasm_Headlines_Dataset.json'\n# data = pd.read_json(news_path, lines=True)\n# X = data['headline']\n# y = data['is_sarcastic']\n\n# print(data.columns)\nprint(data.info())\n","metadata":{"papermill":{"duration":7.151092,"end_time":"2022-07-16T14:59:35.736575","exception":false,"start_time":"2022-07-16T14:59:28.585483","status":"completed"},"tags":[],"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-22T17:30:32.815374Z","iopub.execute_input":"2022-07-22T17:30:32.815755Z","iopub.status.idle":"2022-07-22T17:30:32.858467Z","shell.execute_reply.started":"2022-07-22T17:30:32.815719Z","shell.execute_reply":"2022-07-22T17:30:32.857471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Data Analysis\n<a href=\"#toc\" id=\"1\">Table Of Contents</a>  \n\nFirst we determine if our dataset is balanced or not between positive and negative examples","metadata":{"papermill":{"duration":0.011243,"end_time":"2022-07-16T14:59:35.759888","exception":false,"start_time":"2022-07-16T14:59:35.748645","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib.pyplot as plt\n\ny_count =y.value_counts()\nsns.barplot(y_count.index,y_count)\nplt.gca().set_ylabel('samples')\nplt.gca().set_title('Number of positive and negative samples')\n\nprint(y_count)","metadata":{"papermill":{"duration":0.381241,"end_time":"2022-07-16T14:59:36.152746","exception":false,"start_time":"2022-07-16T14:59:35.771505","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:30:32.877045Z","iopub.execute_input":"2022-07-22T17:30:32.877624Z","iopub.status.idle":"2022-07-22T17:30:33.051001Z","shell.execute_reply.started":"2022-07-22T17:30:32.877583Z","shell.execute_reply":"2022-07-22T17:30:33.050001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The twitter dataset is slightly unbalanced, with positive examples (tweets relative to disasters) being the 40% of the total.","metadata":{}},{"cell_type":"markdown","source":"## Characters/Words count\n<a href=\"#toc\" id=\"1a\">Table Of Contents</a>\n\nWe can use some histograms to determine how the character and word count of the samples is distributed, in samples with positive or negative sentiment. Are positive or negative samples distributed in the same way or there is a difference in length?","metadata":{"papermill":{"duration":0.011606,"end_time":"2022-07-16T14:59:36.176336","exception":false,"start_time":"2022-07-16T14:59:36.164730","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import math\nfig,(ax1,ax2)=plt.subplots(1,2,figsize=(10,5))\nX_len=X[y==1].str.len()\nax1.hist(X_len,color='green')\nax1.set_ylabel('samples')\nax1.set_xlabel('characters')\nax1.set_title('positive samples')\n\nprint(max(X_len))\nprint(int(X_len.mean()))\n\nX_len=X[y==0].str.len()\nax2.hist(X_len,color='red')\nax2.set_ylabel('samples')\nax2.set_xlabel('characters')\nax2.set_title('negative samples')\nfig.suptitle('Characters in samples')\nplt.show()\n","metadata":{"papermill":{"duration":0.406707,"end_time":"2022-07-16T14:59:36.595435","exception":false,"start_time":"2022-07-16T14:59:36.188728","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:30:33.053189Z","iopub.execute_input":"2022-07-22T17:30:33.053875Z","iopub.status.idle":"2022-07-22T17:30:33.359806Z","shell.execute_reply.started":"2022-07-22T17:30:33.053837Z","shell.execute_reply":"2022-07-22T17:30:33.358847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Tweets have an average length of 108 characters, with a maximum length of 151.","metadata":{}},{"cell_type":"code","source":"fig,(ax1,ax2)=plt.subplots(1,2,figsize=(10,5))\nX_len=X[y==1].str.split().map(lambda x: len(x))\nax1.hist(X_len,color='green')\nax1.set_ylabel('samples')\nax1.set_xlabel('words')\nax1.set_title('positive samples')\nX_len=X[y==0].str.split().map(lambda x: len(x))\nax2.hist(X_len,color='red')\nax2.set_ylabel('samples')\nax2.set_xlabel('words')\nax2.set_title('negative samples')\nfig.suptitle('Words in samples')\nplt.show()","metadata":{"papermill":{"duration":0.418952,"end_time":"2022-07-16T14:59:37.027217","exception":false,"start_time":"2022-07-16T14:59:36.608265","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:30:33.361119Z","iopub.execute_input":"2022-07-22T17:30:33.361895Z","iopub.status.idle":"2022-07-22T17:30:33.679624Z","shell.execute_reply.started":"2022-07-22T17:30:33.361855Z","shell.execute_reply":"2022-07-22T17:30:33.678641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dirty characters\n<a href=\"#toc\" id=\"1b\">Table Of Contents</a>\n\nSome datasets can contain samples with some dirty characters.\n- HTML tags\n- Links\n- Hashtags\n- Emoticons (both in single character :smile: or with composed characters :) )\n- Non-ascii characters\n\nThose characters have been removed to improve the quality of the dataset and obtain overall better results.","metadata":{"papermill":{"duration":0.012586,"end_time":"2022-07-16T14:59:37.052685","exception":false,"start_time":"2022-07-16T14:59:37.040099","status":"completed"},"tags":[]}},{"cell_type":"code","source":"y_count = X.str.contains('<.*?>', regex= True, na=False).value_counts()\nsns.barplot(y_count.index,y_count)\nplt.gca().set_ylabel('samples')\nplt.gca().set_title('HTML tags in samples')\n\nX = X.str.replace(r'<.*?>','', regex= True)","metadata":{"papermill":{"duration":0.206198,"end_time":"2022-07-16T14:59:37.270992","exception":false,"start_time":"2022-07-16T14:59:37.064794","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:30:33.682116Z","iopub.execute_input":"2022-07-22T17:30:33.682860Z","iopub.status.idle":"2022-07-22T17:30:33.833634Z","shell.execute_reply.started":"2022-07-22T17:30:33.682817Z","shell.execute_reply":"2022-07-22T17:30:33.832626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_count = X.str.contains('http://+', regex= True, na=False).value_counts()\nsns.barplot(y_count.index,y_count)\nplt.gca().set_ylabel('samples')\nplt.gca().set_title('HTTP links in samples')\n\nX = X.str.replace(r'http://+','', regex= True)","metadata":{"papermill":{"duration":0.229761,"end_time":"2022-07-16T14:59:37.514292","exception":false,"start_time":"2022-07-16T14:59:37.284531","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:30:33.838128Z","iopub.execute_input":"2022-07-22T17:30:33.838799Z","iopub.status.idle":"2022-07-22T17:30:34.014459Z","shell.execute_reply.started":"2022-07-22T17:30:33.838760Z","shell.execute_reply":"2022-07-22T17:30:34.013524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_count = X.str.contains('[^[:ascii:]]', regex= True, na=False).value_counts()\nsns.barplot(y_count.index,y_count)\nplt.gca().set_ylabel('samples')\nplt.gca().set_title('Samples with non ascii characters')\n\n# print(X[X.str.contains('[^[:ascii:]]', regex= True, na=False)])\n# print(X[49677])\n\nX = X.str.replace(r'[^[:ascii:]]','', regex= True)","metadata":{"papermill":{"duration":0.323177,"end_time":"2022-07-16T14:59:37.850872","exception":false,"start_time":"2022-07-16T14:59:37.527695","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:30:34.019351Z","iopub.execute_input":"2022-07-22T17:30:34.019965Z","iopub.status.idle":"2022-07-22T17:30:34.320747Z","shell.execute_reply.started":"2022-07-22T17:30:34.019896Z","shell.execute_reply":"2022-07-22T17:30:34.319721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_count = X.str.contains('&/?[a-z]+;', regex= True, na=False).value_counts()\nsns.barplot(y_count.index,y_count)\nplt.gca().set_ylabel('samples')\nplt.gca().set_title('Samples with non ascii characters')\n\nX = X.str.replace(r'&/?[a-z]+;','', regex= True)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:30:34.325965Z","iopub.execute_input":"2022-07-22T17:30:34.328532Z","iopub.status.idle":"2022-07-22T17:30:34.504455Z","shell.execute_reply.started":"2022-07-22T17:30:34.328486Z","shell.execute_reply":"2022-07-22T17:30:34.501908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Text preprocessing\n<a href=\"#toc\" id=\"2\">Table Of Contents</a>\n\nBefore feeding the text to the model, it needs to be tokenized.\nI used the most basic [Tokenizer](https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/text/Tokenizer) to create the vocabulary using the train dataset, and then proceed to tokenize the train and test dataset.\nThere are a few parameters to decide\n* Vocabulary size\n* Dimension of the embedding: usually between 100 and 300\n* Maximum length(in characters)\n* What to do with sentences shorter or longer than maximum length\n* Further reduction of the vocabulary with lemmatization or stemming.\n\nWe took a vocabulary of 1000 words, and chose a max length of 150, in accord with the size of the tweets. \nLonger tweets may be truncated, while shorter have been padded to reach an equal size of each batch. More advanced techniques may use dynamic batching, where maximum size is inferred in each batch and samples are padded accordingly.\nStemming was done with ``PorterStemmer``.","metadata":{"papermill":{"duration":0.013279,"end_time":"2022-07-16T14:59:37.878200","exception":false,"start_time":"2022-07-16T14:59:37.864921","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import nltk\nimport pandas as pd\nimport numpy as np\nimport re\nfrom nltk.stem.porter import *\n\nwpt = nltk.WordPunctTokenizer()\n\nnltk.download('stopwords')\nstop_words = nltk.corpus.stopwords.words('english')\n\n\nvocab_size = 1000\nembedding_dim = 300\nmax_length = 150\ntrunc_type='post'\npadding_type='post'\noov_tok = \"<UNK>\"\n\nstemmer = PorterStemmer()\n\nSTEMMING = True\nBATCH_SIZE = 64\n\ndef normalize_and_tokenize(doc):\n    # lower case and remove special characters\\whitespaces\n    doc = re.sub(r'[^a-zA-Z\\s!]', '', doc, re.I|re.A)\n    doc = doc.lower()\n    doc = doc.strip()\n    # tokenize document\n    tokens = wpt.tokenize(doc)\n    # filter stopwords out of document\n    filtered_tokens = [token for token in tokens if token not in stop_words]\n    if STEMMING:\n        filtered_tokens = [stemmer.stem(token) for token in filtered_tokens]\n    # re-create document from filtered tokens\n    return ' '.join(filtered_tokens)","metadata":{"papermill":{"duration":0.799591,"end_time":"2022-07-16T14:59:38.691471","exception":false,"start_time":"2022-07-16T14:59:37.891880","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:30:34.506073Z","iopub.execute_input":"2022-07-22T17:30:34.506645Z","iopub.status.idle":"2022-07-22T17:30:34.810953Z","shell.execute_reply.started":"2022-07-22T17:30:34.506606Z","shell.execute_reply":"2022-07-22T17:30:34.809945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With a WordCloud we visualize the most frequent words, we'd like that stopwords or dirty characters were not shown here","metadata":{}},{"cell_type":"code","source":"from wordcloud import WordCloud\n\n\nX = X.map(normalize_and_tokenize)\nwordcloud = WordCloud(\n                background_color ='white',\n                stopwords = stop_words,\n                min_font_size = 10).generate(\" \".join(X.to_list()))\n \n# plot the WordCloud image                      \nplt.figure(figsize = (8, 8), facecolor = None)\nplt.imshow(wordcloud)\nplt.axis(\"off\")\nplt.tight_layout(pad = 0)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:30:34.816716Z","iopub.execute_input":"2022-07-22T17:30:34.817505Z","iopub.status.idle":"2022-07-22T17:30:38.207854Z","shell.execute_reply.started":"2022-07-22T17:30:34.817467Z","shell.execute_reply":"2022-07-22T17:30:38.206727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, train_size=0.8, test_size=0.2,\n                                                                random_state=0)\n\ntokenizer = Tokenizer(num_words = vocab_size\n                      ,oov_token=oov_tok\n                     )\n\ntokenizer.fit_on_texts(X_train)\nword_index = tokenizer.word_index\n\ntrain_sequences = tokenizer.texts_to_sequences(X_train)\ntrain_padded = pad_sequences(train_sequences, maxlen=max_length, padding=padding_type, truncating=trunc_type)\n\ntest_sequences = tokenizer.texts_to_sequences(X_test)\ntest_padded = pad_sequences(test_sequences, maxlen=max_length, padding=padding_type, truncating=trunc_type)","metadata":{"papermill":{"duration":3.273493,"end_time":"2022-07-16T14:59:41.980147","exception":false,"start_time":"2022-07-16T14:59:38.706654","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:17:16.413900Z","iopub.execute_input":"2022-07-22T17:17:16.414435Z","iopub.status.idle":"2022-07-22T17:17:16.655674Z","shell.execute_reply.started":"2022-07-22T17:17:16.414401Z","shell.execute_reply":"2022-07-22T17:17:16.654734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Need this block to get it to work with TensorFlow 2.x\nimport numpy as np\ntrain_padded = np.array(train_padded)\ntrain_labels = np.array(y_train)\ntest_padded = np.array(test_padded)\ntest_labels = np.array(y_test)","metadata":{"papermill":{"duration":0.037001,"end_time":"2022-07-16T14:59:42.030777","exception":false,"start_time":"2022-07-16T14:59:41.993776","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:17:16.657141Z","iopub.execute_input":"2022-07-22T17:17:16.657475Z","iopub.status.idle":"2022-07-22T17:17:16.666481Z","shell.execute_reply.started":"2022-07-22T17:17:16.657440Z","shell.execute_reply":"2022-07-22T17:17:16.665353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_graphs(history, string):\n  plt.plot(history.history[string])\n  plt.plot(history.history['val_'+string])\n  plt.xlabel(\"Epochs\")\n  plt.ylabel(string)\n  plt.legend([string, 'val_'+string])\n  plt.show()","metadata":{"papermill":{"duration":0.025948,"end_time":"2022-07-16T14:59:42.070740","exception":false,"start_time":"2022-07-16T14:59:42.044792","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T18:10:49.710526Z","iopub.execute_input":"2022-07-22T18:10:49.710902Z","iopub.status.idle":"2022-07-22T18:10:49.716927Z","shell.execute_reply.started":"2022-07-22T18:10:49.710870Z","shell.execute_reply":"2022-07-22T18:10:49.715805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Experiments\n<a href=\"#toc\" id=\"3\">Table Of Contents</a>\n## Simple Feed Forward Network\n<a href=\"#toc\" id=\"3a\"></a>","metadata":{"papermill":{"duration":0.013333,"end_time":"2022-07-16T14:59:42.132584","exception":false,"start_time":"2022-07-16T14:59:42.119251","status":"completed"},"tags":[]}},{"cell_type":"code","source":"from tensorflow.keras.optimizers import SGD, Adam\nfrom sklearn.metrics import classification_report\n\noptimizer=Adam(learning_rate=1e-4)\n\ndef evaluate(model,X_test,y_test):\n    y_hat = model.predict(X_test,batch_size = BATCH_SIZE)\n    y_hat = (y_hat > 0.5).astype(np.float32)\n    if(len(y_hat.shape)>2):\n        y_hat = y_hat[:,-1,0]\n    report = classification_report(y_test, y_hat)\n    return report","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:50:38.172087Z","iopub.execute_input":"2022-07-22T17:50:38.172441Z","iopub.status.idle":"2022-07-22T17:50:38.180150Z","shell.execute_reply.started":"2022-07-22T17:50:38.172410Z","shell.execute_reply":"2022-07-22T17:50:38.179033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = tf.keras.Sequential([\n    tf.keras.layers.Embedding(vocab_size, embedding_dim, input_length=max_length),\n    tf.keras.layers.GlobalAveragePooling1D(),\n    tf.keras.layers.Dense(24, activation='relu'),\n    tf.keras.layers.Dropout(0.5),\n    tf.keras.layers.Dense(1, activation='sigmoid')\n])\nmodel.compile(loss='binary_crossentropy',optimizer=optimizer,metrics=['accuracy'])\nmodel.summary()","metadata":{"papermill":{"duration":3.344529,"end_time":"2022-07-16T14:59:45.490501","exception":false,"start_time":"2022-07-16T14:59:42.145972","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:17:16.690405Z","iopub.execute_input":"2022-07-22T17:17:16.691161Z","iopub.status.idle":"2022-07-22T17:17:16.732144Z","shell.execute_reply.started":"2022-07-22T17:17:16.691126Z","shell.execute_reply":"2022-07-22T17:17:16.731202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Training","metadata":{"papermill":{"duration":0.012963,"end_time":"2022-07-16T14:59:45.517649","exception":false,"start_time":"2022-07-16T14:59:45.504686","status":"completed"},"tags":[]}},{"cell_type":"code","source":"num_epochs = 40\nhistory = model.fit(train_padded, train_labels, epochs=num_epochs, validation_split=0.2, batch_size = BATCH_SIZE, verbose=2)\nplot_graphs(history, \"accuracy\")\nplot_graphs(history, \"loss\")\n\nreport = evaluate(model,test_padded,test_labels)\nprint(report)","metadata":{"papermill":{"duration":83.21011,"end_time":"2022-07-16T15:01:08.741972","exception":false,"start_time":"2022-07-16T14:59:45.531862","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:17:16.733676Z","iopub.execute_input":"2022-07-22T17:17:16.734032Z","iopub.status.idle":"2022-07-22T17:17:29.667418Z","shell.execute_reply.started":"2022-07-22T17:17:16.733998Z","shell.execute_reply":"2022-07-22T17:17:29.666364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Recurrent Neural Networks\n<a href=\"#toc\" id=\"3b\">Table Of Contents</a>\n\nFor now disabled, they require too much time for training.\nIn addition, simple RNNs are subjected to vanishing gradient and the context is limited to a few time steps.","metadata":{"papermill":{"duration":0.01946,"end_time":"2022-07-16T15:01:08.780923","exception":false,"start_time":"2022-07-16T15:01:08.761463","status":"completed"},"tags":[]}},{"cell_type":"code","source":"model = tf.keras.Sequential([\n    tf.keras.layers.Embedding(vocab_size, embedding_dim, input_length=max_length),\n    tf.keras.layers.Bidirectional(tf.keras.layers.SimpleRNN(50, return_sequences=True,input_shape=(max_length, embedding_dim))),\n    tf.keras.layers.Dropout(0.5),\n    tf.keras.layers.Flatten(),\n    tf.keras.layers.Dense(1, activation='sigmoid')\n])\nmodel.compile(loss='binary_crossentropy',optimizer=optimizer,metrics=['accuracy'])\nmodel.summary()","metadata":{"papermill":{"duration":0.211155,"end_time":"2022-07-16T15:01:09.011578","exception":false,"start_time":"2022-07-16T15:01:08.800423","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:17:29.669386Z","iopub.execute_input":"2022-07-22T17:17:29.669771Z","iopub.status.idle":"2022-07-22T17:17:29.805398Z","shell.execute_reply.started":"2022-07-22T17:17:29.669734Z","shell.execute_reply":"2022-07-22T17:17:29.804357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# num_epochs = 5\n# history = model.fit(train_padded, train_labels, epochs=num_epochs, validation_split=0.2, verbose=2)\n# plot_graphs(history, \"accuracy\")\n# plot_graphs(history, \"loss\")","metadata":{"papermill":{"duration":0.02965,"end_time":"2022-07-16T15:01:09.061483","exception":false,"start_time":"2022-07-16T15:01:09.031833","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:17:29.807065Z","iopub.execute_input":"2022-07-22T17:17:29.807642Z","iopub.status.idle":"2022-07-22T17:17:29.813800Z","shell.execute_reply.started":"2022-07-22T17:17:29.807603Z","shell.execute_reply":"2022-07-22T17:17:29.812128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## LSTMs\n<a href=\"#toc\" id=\"3c\">Table Of Contents</a>","metadata":{"papermill":{"duration":0.020037,"end_time":"2022-07-16T15:01:09.101605","exception":false,"start_time":"2022-07-16T15:01:09.081568","status":"completed"},"tags":[]}},{"cell_type":"code","source":"model = tf.keras.Sequential([\n    tf.keras.layers.Embedding(vocab_size, embedding_dim, input_length=max_length),\n    tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(32, return_sequences=True,input_shape=(max_length, embedding_dim),dropout=0.5)),\n    tf.keras.layers.Dense(24, activation='relu'),\n    tf.keras.layers.Dense(1, activation='sigmoid')\n])\nmodel.compile(loss='binary_crossentropy',optimizer=optimizer,metrics=['accuracy'])\nmodel.summary()","metadata":{"papermill":{"duration":0.639484,"end_time":"2022-07-16T15:01:09.760931","exception":false,"start_time":"2022-07-16T15:01:09.121447","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:17:29.815234Z","iopub.execute_input":"2022-07-22T17:17:29.815663Z","iopub.status.idle":"2022-07-22T17:17:30.312205Z","shell.execute_reply.started":"2022-07-22T17:17:29.815627Z","shell.execute_reply":"2022-07-22T17:17:30.310466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_epochs = 40\ncallback = tf.keras.callbacks.EarlyStopping(monitor='loss', patience=39)\n\nhistory = model.fit(train_padded, train_labels, epochs=num_epochs, validation_split=0.2,batch_size = BATCH_SIZE, verbose=2,callbacks=[callback])\nplot_graphs(history, \"accuracy\")\nplot_graphs(history, \"loss\")\n\nreport = evaluate(model,test_padded,test_labels)\nprint(report)","metadata":{"papermill":{"duration":445.904112,"end_time":"2022-07-16T15:08:35.685596","exception":false,"start_time":"2022-07-16T15:01:09.781484","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:17:30.313565Z","iopub.execute_input":"2022-07-22T17:17:30.314182Z","iopub.status.idle":"2022-07-22T17:18:28.612354Z","shell.execute_reply.started":"2022-07-22T17:17:30.314145Z","shell.execute_reply":"2022-07-22T17:18:28.611011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## GRUs\n<a href=\"#toc\" id=\"3d\">Table Of Contents</a>","metadata":{}},{"cell_type":"code","source":"model = tf.keras.Sequential([\n    tf.keras.layers.Embedding(vocab_size, embedding_dim, input_length=max_length),\n    tf.keras.layers.Bidirectional(tf.keras.layers.GRU(32, return_sequences=True,input_shape=(max_length, embedding_dim),dropout=0.5)),\n    tf.keras.layers.Dense(1, activation='sigmoid')\n])\nmodel.compile(loss='binary_crossentropy',optimizer=optimizer,metrics=['accuracy'])\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:18:28.614126Z","iopub.execute_input":"2022-07-22T17:18:28.615149Z","iopub.status.idle":"2022-07-22T17:18:29.003156Z","shell.execute_reply.started":"2022-07-22T17:18:28.615112Z","shell.execute_reply":"2022-07-22T17:18:29.001367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_epochs = 40\ncallback = tf.keras.callbacks.EarlyStopping(monitor='loss', patience=5)\n\nhistory = model.fit(train_padded, train_labels, epochs=num_epochs, validation_split=0.2,batch_size = BATCH_SIZE, verbose=2,callbacks=[callback])\nplot_graphs(history, \"accuracy\")\nplot_graphs(history, \"loss\")\n\nreport = evaluate(model,test_padded,test_labels)\nprint(report)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:18:29.004674Z","iopub.execute_input":"2022-07-22T17:18:29.005278Z","iopub.status.idle":"2022-07-22T17:19:23.997207Z","shell.execute_reply.started":"2022-07-22T17:18:29.005240Z","shell.execute_reply":"2022-07-22T17:19:23.996196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2 Layers\n<a href=\"#toc\" id=\"3e\">Table Of Contents</a>\n\nWe also tried a more complex model with 2 bidirectional LSTM layers","metadata":{"papermill":{"duration":0.024746,"end_time":"2022-07-16T15:08:35.735885","exception":false,"start_time":"2022-07-16T15:08:35.711139","status":"completed"},"tags":[]}},{"cell_type":"code","source":"model = tf.keras.Sequential([\n    tf.keras.layers.Embedding(vocab_size, embedding_dim, input_length=max_length),\n    tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(32, return_sequences=True,input_shape=(max_length, embedding_dim),dropout=0.5)),\n    tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(32, return_sequences=True,dropout=0.5)),\n#   tf.keras.layers.Flatten(),\n    tf.keras.layers.Dense(24, activation='relu'),\n    tf.keras.layers.Dense(1, activation='sigmoid')\n])\nmodel.compile(loss='binary_crossentropy',optimizer=optimizer,metrics=['accuracy'])\nmodel.summary()","metadata":{"papermill":{"duration":1.172977,"end_time":"2022-07-16T15:08:36.934519","exception":false,"start_time":"2022-07-16T15:08:35.761542","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:19:23.998843Z","iopub.execute_input":"2022-07-22T17:19:23.999192Z","iopub.status.idle":"2022-07-22T17:19:24.816340Z","shell.execute_reply.started":"2022-07-22T17:19:23.999157Z","shell.execute_reply":"2022-07-22T17:19:24.814659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_epochs = 40\nhistory = model.fit(train_padded, train_labels, epochs=num_epochs, validation_split=0.2, verbose=2)\nplot_graphs(history, \"accuracy\")\nplot_graphs(history, \"loss\")\n\nreport = evaluate(model,test_padded,test_labels)\nprint(report)","metadata":{"papermill":{"duration":329.597888,"end_time":"2022-07-16T15:14:06.558678","exception":false,"start_time":"2022-07-16T15:08:36.960790","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:19:24.817730Z","iopub.execute_input":"2022-07-22T17:19:24.818176Z","iopub.status.idle":"2022-07-22T17:22:09.786274Z","shell.execute_reply.started":"2022-07-22T17:19:24.818137Z","shell.execute_reply":"2022-07-22T17:22:09.785260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **GloVe pretrained word embeddings - text-preprocessing WITHOUT stemming**\n<a href=\"toc\" id=\"4\"></a>","metadata":{"papermill":{"duration":0.02833,"end_time":"2022-07-16T15:14:06.804907","exception":false,"start_time":"2022-07-16T15:14:06.776577","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"**text pre-processing**","metadata":{"papermill":{"duration":0.028898,"end_time":"2022-07-16T15:14:06.862446","exception":false,"start_time":"2022-07-16T15:14:06.833548","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom tensorflow.keras.preprocessing.text import Tokenizer\nfrom tensorflow.keras.preprocessing.sequence import pad_sequences\n# TEXT PRE-PROCESSING WITHOUT STEMMING\n \n# import nltk\n# import re\n# from nltk.stem.porter import *\n\n\n# nltk.download('stopwords')\n\n# wpt = nltk.WordPunctTokenizer()\n# stop_words = nltk.corpus.stopwords.words('english')\n\n# vocab_size = 1000\n# embedding_dim = 300\n# max_length = 1000\n# trunc_type='post'\n# padding_type='post'\n# oov_tok = \"<UNK>\"\n\nSTEMMING = False\nX = X.map(normalize_and_tokenize)\n\nX_train, X_test, y_train, y_test = train_test_split(X, y, train_size=0.8, test_size=0.2,\n                                                                random_state=0)\n\ntokenizer = Tokenizer(num_words = vocab_size\n                      ,oov_token= oov_tok\n                     )\ntokenizer.fit_on_texts(X_train)\nword_index = tokenizer.word_index\n\ntrain_sequences = tokenizer.texts_to_sequences(X_train)\ntrain_padded = pad_sequences(train_sequences, maxlen=max_length, padding=padding_type, truncating=trunc_type)\n\ntest_sequences = tokenizer.texts_to_sequences(X_test)\ntest_padded = pad_sequences(test_sequences, maxlen=max_length, padding=padding_type, truncating=trunc_type)\n\nimport numpy as np\ntrain_padded = np.array(train_padded)\ntrain_labels = np.array(y_train)\ntest_padded = np.array(test_padded)\ntest_labels = np.array(y_test)\n","metadata":{"papermill":{"duration":0.627799,"end_time":"2022-07-16T15:14:07.518606","exception":false,"start_time":"2022-07-16T15:14:06.890807","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:30:51.026593Z","iopub.execute_input":"2022-07-22T17:30:51.027303Z","iopub.status.idle":"2022-07-22T17:30:51.502084Z","shell.execute_reply.started":"2022-07-22T17:30:51.027264Z","shell.execute_reply":"2022-07-22T17:30:51.501096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **LSTMs**\n<a href=\"#toc\" id=\"4a\">Table Of Contents</a>","metadata":{"papermill":{"duration":0.027983,"end_time":"2022-07-16T15:14:07.575862","exception":false,"start_time":"2022-07-16T15:14:07.547879","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"**import pre-trained static embeddings**","metadata":{"papermill":{"duration":0.028508,"end_time":"2022-07-16T15:14:07.632568","exception":false,"start_time":"2022-07-16T15:14:07.604060","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# load the GloVe embedding into memory \nimport numpy as np\nembeddings_index = dict()\nf = open(\"../input/glove6b/glove.6B.300d.txt\")\n\nfor line in f: \n    values = line.split()\n    word = values[0]\n    coefs = np.asarray(values[1:], dtype='float32')\n    embeddings_index[word] = coefs \n    \nf.close()\nprint('Loaded %s word vectors.' %len(embeddings_index))\n\n\n# create the weight matrix \nmissing=0\nembedding_matrix = np.zeros((len(tokenizer.word_index.items()) + 1, 300))\n\nfor word, i in tokenizer.word_index.items():\n    embedding_vector = embeddings_index.get(word)\n    if embedding_vector is not None: \n        embedding_matrix[i] = embedding_vector\n    else:\n        missing = missing + 1\nprint('%d embeddings are missing' %missing)","metadata":{"papermill":{"duration":38.461078,"end_time":"2022-07-16T15:14:46.122573","exception":false,"start_time":"2022-07-16T15:14:07.661495","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:22:10.280534Z","iopub.execute_input":"2022-07-22T17:22:10.281140Z","iopub.status.idle":"2022-07-22T17:22:40.729849Z","shell.execute_reply.started":"2022-07-22T17:22:10.281103Z","shell.execute_reply":"2022-07-22T17:22:40.728726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow import keras\n\nmodel = tf.keras.Sequential([\n    tf.keras.layers.Embedding(len(tokenizer.word_index.items()) + 1, embedding_dim, weights = [embedding_matrix], input_length=max_length, trainable=True),\n    tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(32, return_sequences=True,input_shape=(max_length, embedding_dim),dropout=0.5)),\n    tf.keras.layers.Dense(24, activation='relu'),\n    tf.keras.layers.Dense(1, activation='sigmoid')\n])\nmodel.compile(loss='binary_crossentropy',optimizer=optimizer,metrics=['accuracy'])\nmodel.summary()","metadata":{"papermill":{"duration":0.684936,"end_time":"2022-07-16T15:14:46.835366","exception":false,"start_time":"2022-07-16T15:14:46.150430","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:22:40.731295Z","iopub.execute_input":"2022-07-22T17:22:40.732175Z","iopub.status.idle":"2022-07-22T17:22:41.215836Z","shell.execute_reply.started":"2022-07-22T17:22:40.732134Z","shell.execute_reply":"2022-07-22T17:22:41.214830Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_epochs = 40\nBATCH_SIZE = 64\ncallback = tf.keras.callbacks.EarlyStopping(monitor='loss', patience=5)\n\nhistory = model.fit(train_padded, train_labels, epochs=num_epochs, validation_split=0.2,batch_size = BATCH_SIZE, verbose=2,callbacks=[callback])\nplot_graphs(history, \"accuracy\")\nplot_graphs(history, \"loss\")\n\nreport = evaluate(model,test_padded,test_labels)\nprint(report)","metadata":{"papermill":{"duration":146.62994,"end_time":"2022-07-16T15:17:13.494475","exception":false,"start_time":"2022-07-16T15:14:46.864535","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:22:41.217127Z","iopub.execute_input":"2022-07-22T17:22:41.217872Z","iopub.status.idle":"2022-07-22T17:23:43.734834Z","shell.execute_reply.started":"2022-07-22T17:22:41.217835Z","shell.execute_reply":"2022-07-22T17:23:43.733915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## GRUs\n<a href=\"#toc\" id=\"4b\">Table Of Contents</a>","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow import keras\n\nmodel = tf.keras.Sequential([\n    tf.keras.layers.Embedding(len(tokenizer.word_index.items()) + 1, embedding_dim, weights = [embedding_matrix], input_length=max_length, trainable=True),\n    tf.keras.layers.Bidirectional(tf.keras.layers.GRU(32, return_sequences=True,input_shape=(max_length, embedding_dim),dropout=0.5)),\n#     tf.keras.layers.Dense(24, activation='relu'),\n    tf.keras.layers.Dense(1, activation='sigmoid')\n])\nmodel.compile(loss='binary_crossentropy',optimizer=optimizer,metrics=['accuracy'])\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:23:43.736544Z","iopub.execute_input":"2022-07-22T17:23:43.737007Z","iopub.status.idle":"2022-07-22T17:23:44.129726Z","shell.execute_reply.started":"2022-07-22T17:23:43.736962Z","shell.execute_reply":"2022-07-22T17:23:44.128764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_epochs = 40\nBATCH_SIZE = 64\ncallback = tf.keras.callbacks.EarlyStopping(monitor='loss', patience=5)\n\nhistory = model.fit(train_padded, train_labels, epochs=num_epochs, validation_split=0.2,batch_size = BATCH_SIZE, verbose=2,callbacks=[callback])\nplot_graphs(history, \"accuracy\")\nplot_graphs(history, \"loss\")\n\nreport = evaluate(model,test_padded,test_labels)\nprint(report)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:23:44.131316Z","iopub.execute_input":"2022-07-22T17:23:44.131971Z","iopub.status.idle":"2022-07-22T17:24:40.572214Z","shell.execute_reply.started":"2022-07-22T17:23:44.131935Z","shell.execute_reply":"2022-07-22T17:24:40.570948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **GloVe pre-trained static embedding on Twitter data - without STEMMING**\n\nThe quality of embeddings is crucial for the performance: we tried another set of pretrained embedding, of dimension 200 instead of 300 but trained on Twitter data.\nSince these embeddings may be similar to the dataset we are evaluating, they may lead to a better performance.","metadata":{"papermill":{"duration":0.02994,"end_time":"2022-07-16T15:17:13.555620","exception":false,"start_time":"2022-07-16T15:17:13.525680","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## **LSTMs**\n<a href=\"#toc\" id=\"4c\">Table Of Contents</a>","metadata":{"papermill":{"duration":0.030322,"end_time":"2022-07-16T15:17:13.616976","exception":false,"start_time":"2022-07-16T15:17:13.586654","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"**import GloVe static embedding pre trained on twitter data with 200 dimension**","metadata":{"papermill":{"duration":0.030208,"end_time":"2022-07-16T15:17:13.677328","exception":false,"start_time":"2022-07-16T15:17:13.647120","status":"completed"},"tags":[]}},{"cell_type":"code","source":"\n# load the GloVe embedding into memory \nimport numpy as np\ntwitter_embeddings_index = dict()\nf = open(\"../input/glove-global-vectors-for-word-representation/glove.twitter.27B.200d.txt\")\n\n\n\nfor line in f: \n    values = line.split()\n    word = values[0]\n    coefs = np.asarray(values[1:], dtype='float32')\n    twitter_embeddings_index[word] = coefs \n    \nf.close()\nprint('Loaded %s word vectors.' %len(twitter_embeddings_index))\n\n\n# create the weight matrix \nmissing=0\ntwitter_embedding_matrix = np.zeros((len(tokenizer.word_index.items()) + 1, 200))\n\nfor word, i in tokenizer.word_index.items():\n    twitter_embedding_vector = twitter_embeddings_index.get(word)\n    if twitter_embedding_vector is not None: \n        twitter_embedding_matrix[i] = twitter_embedding_vector\n    else:\n        missing = missing + 1\nprint('%d embeddings are missing' %missing)\n","metadata":{"papermill":{"duration":78.996108,"end_time":"2022-07-16T15:18:32.704174","exception":false,"start_time":"2022-07-16T15:17:13.708066","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:24:40.573656Z","iopub.execute_input":"2022-07-22T17:24:40.574342Z","iopub.status.idle":"2022-07-22T17:25:51.336975Z","shell.execute_reply.started":"2022-07-22T17:24:40.574302Z","shell.execute_reply":"2022-07-22T17:25:51.335769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**LSTMs model definition**","metadata":{"papermill":{"duration":0.0523,"end_time":"2022-07-16T15:18:32.808055","exception":false,"start_time":"2022-07-16T15:18:32.755755","status":"completed"},"tags":[]}},{"cell_type":"code","source":"\nimport tensorflow as tf\nfrom tensorflow import keras\ntwitter_embedding_dim = 200\nmodel = tf.keras.Sequential([\n    tf.keras.layers.Embedding(len(tokenizer.word_index.items()) + 1, twitter_embedding_dim, weights = [twitter_embedding_matrix], input_length=max_length, trainable=True),\n    tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(32, return_sequences=True,input_shape=(max_length, twitter_embedding_dim),dropout=0.5)),\n    tf.keras.layers.Dense(24, activation='relu'),\n    tf.keras.layers.Dense(1, activation='sigmoid')\n])\nmodel.compile(loss='binary_crossentropy',optimizer=optimizer,metrics=['accuracy'])\nmodel.summary()","metadata":{"papermill":{"duration":0.749484,"end_time":"2022-07-16T15:18:33.610908","exception":false,"start_time":"2022-07-16T15:18:32.861424","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:25:51.338566Z","iopub.execute_input":"2022-07-22T17:25:51.338935Z","iopub.status.idle":"2022-07-22T17:25:51.812191Z","shell.execute_reply.started":"2022-07-22T17:25:51.338900Z","shell.execute_reply":"2022-07-22T17:25:51.810460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_epochs = 40\nBATCH_SIZE = 64\ncallback = tf.keras.callbacks.EarlyStopping(monitor='loss', patience=5)\n\nhistory = model.fit(train_padded, train_labels, epochs=num_epochs, validation_split=0.2,batch_size = BATCH_SIZE, verbose=2,callbacks=[callback])\nplot_graphs(history, \"accuracy\")\nplot_graphs(history, \"loss\")\n\nreport = evaluate(model,test_padded,test_labels)\nprint(report)","metadata":{"papermill":{"duration":122.42639,"end_time":"2022-07-16T15:20:36.068581","exception":false,"start_time":"2022-07-16T15:18:33.642191","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-07-22T17:25:51.813522Z","iopub.execute_input":"2022-07-22T17:25:51.813963Z","iopub.status.idle":"2022-07-22T17:27:19.293560Z","shell.execute_reply.started":"2022-07-22T17:25:51.813928Z","shell.execute_reply":"2022-07-22T17:27:19.292324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## GRUs\n<a href=\"#toc\" id=\"4d\">Table Of Contents</a>","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nimport keras\ntwitter_embedding_dim = 200\nmodel = tf.keras.Sequential([\n    tf.keras.layers.Embedding(len(tokenizer.word_index.items()) + 1, twitter_embedding_dim, weights = [twitter_embedding_matrix], input_length=max_length, trainable=True),\n    tf.keras.layers.Dropout(0.25),\n    tf.keras.layers.Bidirectional(tf.keras.layers.GRU(32, return_sequences=True,input_shape=(max_length, twitter_embedding_dim))),\n    tf.keras.layers.Dense(24, activation='relu'),\n    tf.keras.layers.Dense(1, activation='sigmoid')\n])\nmodel.compile(loss='binary_crossentropy',optimizer=optimizer,metrics=['accuracy'])\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:27:19.295148Z","iopub.execute_input":"2022-07-22T17:27:19.295487Z","iopub.status.idle":"2022-07-22T17:27:19.710750Z","shell.execute_reply.started":"2022-07-22T17:27:19.295453Z","shell.execute_reply":"2022-07-22T17:27:19.709754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_epochs = 40\nBATCH_SIZE = 64\ncallback = tf.keras.callbacks.EarlyStopping(monitor='loss', patience=5)\n\nhistory = model.fit(train_padded, train_labels, epochs=num_epochs, validation_split=0.2,batch_size = BATCH_SIZE, verbose=2,callbacks=[callback])\nplot_graphs(history, \"accuracy\")\nplot_graphs(history, \"loss\")\n\nreport = evaluate(model,test_padded,test_labels)\nprint(report)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:27:19.712301Z","iopub.execute_input":"2022-07-22T17:27:19.712629Z","iopub.status.idle":"2022-07-22T17:28:12.816357Z","shell.execute_reply.started":"2022-07-22T17:27:19.712595Z","shell.execute_reply":"2022-07-22T17:28:12.815093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Contextual Embeddings\n<a href=\"#toc\" id=\"5\">Table Of Contents</a>","metadata":{}},{"cell_type":"code","source":"from transformers import BertTokenizer, TFBertModel\nfrom tqdm.notebook import tqdm\n\nimport tensorflow_hub as hub\n\n# bert_model = TFBertForSequenceClassification.from_pretrained(\"bert-base-uncased\",num_labels=2)\nbert_tokenizer = BertTokenizer.from_pretrained(\"bert-base-uncased\")\nbert_model = TFBertModel.from_pretrained('bert-base-uncased')","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:29:29.171141Z","iopub.execute_input":"2022-07-22T17:29:29.171543Z","iopub.status.idle":"2022-07-22T17:29:39.054468Z","shell.execute_reply.started":"2022-07-22T17:29:29.171514Z","shell.execute_reply":"2022-07-22T17:29:39.053483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"max_length = 150\n\nsent= 'how to train the model, lets look at how a trained model calculates its prediction.'\ntokens=bert_tokenizer.tokenize(sent)\nprint(tokens)\n\ntokenized_sequence= bert_tokenizer.encode_plus(sent,add_special_tokens = True,truncation=True,max_length =max_length,\nreturn_attention_mask = True)\ntokenized_sequence","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:29:44.671210Z","iopub.execute_input":"2022-07-22T17:29:44.671569Z","iopub.status.idle":"2022-07-22T17:29:44.684222Z","shell.execute_reply.started":"2022-07-22T17:29:44.671538Z","shell.execute_reply":"2022-07-22T17:29:44.683159Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bert_tokenizer.decode(tokenized_sequence['input_ids'])","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:29:46.833729Z","iopub.execute_input":"2022-07-22T17:29:46.834331Z","iopub.status.idle":"2022-07-22T17:29:48.228563Z","shell.execute_reply.started":"2022-07-22T17:29:46.834291Z","shell.execute_reply":"2022-07-22T17:29:48.227629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nimport keras\nfrom tensorflow.keras.models import Model, Sequential\nfrom tensorflow.keras.optimizers import SGD, Adam\n\ndef bert_encode(data, max_length) :\n    input_ids = []\n    attention_masks = []\n\n    for text in data:\n        encoded = tokenizer.encode_plus(\n            text, \n            add_special_tokens=True,\n            max_length=max_length,\n            pad_to_max_length=True,\n\n            return_attention_mask=True,\n        )\n        input_ids.append(encoded['input_ids'])\n        attention_masks.append(encoded['attention_mask'])\n        \n    return np.array(input_ids),np.array(attention_masks)\n\ndef create_bert(bert_model,max_length):\n    \n    input_ids = tf.keras.Input(shape=(max_length,),dtype='int32')\n    attention_masks = tf.keras.Input(shape=(max_length,),dtype='int32')\n\n    output = bert_model([input_ids,attention_masks])\n    output = output[1]\n    output = tf.keras.layers.Dense(32,activation='relu')(output)\n    output = tf.keras.layers.Dropout(0.5)(output)\n    output = tf.keras.layers.Dense(1,activation='sigmoid')(output)\n    \n    model = tf.keras.models.Model(inputs = [input_ids,attention_masks],outputs = output)\n    model.compile(Adam(lr=1e-5), loss='binary_crossentropy', metrics=['accuracy'])\n    return model","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:29:49.196681Z","iopub.execute_input":"2022-07-22T17:29:49.198138Z","iopub.status.idle":"2022-07-22T17:29:49.210565Z","shell.execute_reply.started":"2022-07-22T17:29:49.198103Z","shell.execute_reply":"2022-07-22T17:29:49.209358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = create_bert(bert_model,max_length)\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:29:51.836618Z","iopub.execute_input":"2022-07-22T17:29:51.837794Z","iopub.status.idle":"2022-07-22T17:29:57.738145Z","shell.execute_reply.started":"2022-07-22T17:29:51.837742Z","shell.execute_reply":"2022-07-22T17:29:57.737102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"input_ids=[]\nattention_masks=[]\n\nfor sent in X:\n    bert_inp=bert_tokenizer.encode_plus(sent,add_special_tokens = True,truncation=True,max_length =max_length,pad_to_max_length = True,return_attention_mask = True)\n    input_ids.append(bert_inp['input_ids'])\n    attention_masks.append(bert_inp['attention_mask'])\n\ninput_ids=np.asarray(input_ids)\nattention_masks=np.array(attention_masks)\ny=np.array(y)\n\ntrain_inp,val_inp,train_label,val_label,train_mask,val_mask=train_test_split(input_ids,y,attention_masks,test_size=0.2)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:51:58.770402Z","iopub.execute_input":"2022-07-22T17:51:58.770777Z","iopub.status.idle":"2022-07-22T17:52:03.191070Z","shell.execute_reply.started":"2022-07-22T17:51:58.770747Z","shell.execute_reply":"2022-07-22T17:52:03.190048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tensorflow as tf\nimport keras\n\n# callbacks = [tf.keras.callbacks.ModelCheckpoint(filepath=model_save_path,save_weights_only=True,monitor='val_loss',mode='min',save_best_only=True),keras.callbacks.TensorBoard(log_dir=log_dir),tf.keras.callbacks.EarlyStopping(monitor='loss', patience=5)]\n# history=model.fit([train_inp,train_mask],train_label,batch_size=BATCH_SIZE,epochs=10,validation_data=([val_inp,val_mask],val_label),callbacks=callbacks,verbose=2)\n\nhistory = model.fit(\n    [train_inp, train_mask],\n    train_label,\n    validation_split=0.2, \n    epochs=10,\n    batch_size=BATCH_SIZE,\n    verbose=2\n)\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-22T17:52:11.256572Z","iopub.execute_input":"2022-07-22T17:52:11.257295Z","iopub.status.idle":"2022-07-22T18:07:33.261340Z","shell.execute_reply.started":"2022-07-22T17:52:11.257250Z","shell.execute_reply":"2022-07-22T18:07:33.260341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"report = evaluate(model,[val_inp,val_mask],val_label)\nprint(report)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T18:10:08.928115Z","iopub.execute_input":"2022-07-22T18:10:08.928490Z","iopub.status.idle":"2022-07-22T18:10:20.196857Z","shell.execute_reply.started":"2022-07-22T18:10:08.928460Z","shell.execute_reply":"2022-07-22T18:10:20.195856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nplot_graphs(history, \"accuracy\")\nplot_graphs(history, \"loss\")","metadata":{"execution":{"iopub.status.busy":"2022-07-22T18:10:53.793232Z","iopub.execute_input":"2022-07-22T18:10:53.793580Z","iopub.status.idle":"2022-07-22T18:10:54.142487Z","shell.execute_reply.started":"2022-07-22T18:10:53.793549Z","shell.execute_reply":"2022-07-22T18:10:54.141553Z"},"trusted":true},"execution_count":null,"outputs":[]}]}