{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<b><font size = 2><span style=\"color:#2a6592\">Transformer models comparison for disaster tweets 🌊 </span></font></b>  \n\n<b><font size = 2><span style=\"color:#2a6592\">Created By Burhanuddin Latsaheb </span></font> </b> \n\n\n<center><font size = 6><span style=\"color:#2F4F4F\"> Transformer models comparison for disaster tweets 🌊 </span></font></center>  \n\n## <center><font size =4><span style=\"color:#2F4F4F\"> If you find this notebook useful,support with an upvote👍👍 </span></font></center>\n","metadata":{}},{"cell_type":"markdown","source":" <b><center><font size = 6><span style=\"color:#2a6592\"> Introduction </span></font></center></b>\n<font size = 5><span style=\"color:#db9833\">Notebook Overview : </span></font>\n\n* <font size = 3><span style=\"color:#2F4F4F\"> This notebook contains training , evaluation and comparision of 5 transformer models on Disaster Tweets Dataset . </span></font>\n*  <font size = 3><span style=\"color:#2F4F4F\">5 of the pretrained transformer models and their respective tokenizers are taken from *[hugging face](https://huggingface.co)*</span></font>\n*  <font size = 3><span style=\"color:#2F4F4F\">The models are trained using the `Tensorflow` Library</span></font>\n","metadata":{}},{"cell_type":"markdown","source":"<a id = \"toc\"></a>\n<font size = 5><span style=\"color:#db9833\">Table of Contents : </span></font>\n\n- [1. Imports](#imports)\n- [2. Helper Functions](#helper_functions)\n- [3. Hyperparameters](#hyperparameters)\n- [4. EDA](#EDA)\n  * [4.1 Random examples](#Random_Examples)\n  * [4.2 Common Words](#Common_Words)\n- [5. Dataset](#dataset)\n- [6. Models](#Models)\n  * [6.1 Bert Base Uncased](#bert_base_uncased)\n    + [6.1.1 Tokenization](#bert_base_uncased_tokenization)\n    + [6.1.2 Model Loading](#bert_base_uncased_model_loading)\n    + [6.1.3 Model Building](#bert_base_uncased_model)\n    + [6.1.4 Callback](#bert_base_uncased_callback)\n    + [6.1.5 Compile](#bert_base_uncased_compile)\n    + [6.1.6 Training](#bert_base_uncased_training)\n    + [6.1.7 Graphs](#bert_base_uncased_graph)\n    + [6.1.8 Confusion Matrix](#bert_base_uncased_confusion_matrix)\n    + [6.1.9 Predictions](#bert_base_uncased_predictions)\n  * [6.2 Bert Large](#bert_large)\n    + [6.2.1 Tokenization](#bert_large_tokenization)\n    + [6.2.2 Model Loading](#bert_large_model_loading)\n    + [6.2.3 Model Building](#bert_large_model)\n    + [6.2.4 Callback](#bert_large_callback)\n    + [6.2.5 Compile](#bert_large_compile)\n    + [6.2.6 Training](#bert_large_training)\n    + [6.2.7 Graphs](#bert_large_graph)\n    + [6.2.8 Confusion Matrix](#bert_large_confusion_matrix)\n    + [6.2.9 Predictions](#bert_large_predictions)\n * [6.3 Distill Bert](#distill_bert)\n    + [6.3.1 Tokenization](#distill_bert_tokenization)\n    + [6.3.2 Model Loading](#distill_bert_model_loading)\n    + [6.3.3 Model Building](#distill_bert_model)\n    + [6.3.4 Callback](#distill_bert_callback)\n    + [6.3.5 Compile](#distill_bert_compile)\n    + [6.3.6 Training](#distill_bert_training)\n    + [6.3.7 Graphs](#distill_bert_graph)\n    + [6.3.8 Confusion Matrix](#distill_bert_confusion_matrix)\n    + [6.3.9 Predictions](#dsitill_bert_predictions)\n  * [6.4 Roberta](#roberta_base)\n    + [6.4.1 Tokenization](#roberta_tokenization)\n    + [6.4.2 Model Loading](#roberta_model_loading)\n    + [6.4.3 Model Building](#roberta_model)\n    + [6.4.4 Callback](#roberta_callback)\n    + [6.4.5 Compile](#roberta_compile)\n    + [6.4.6 Training](#roberta_training)\n    + [6.4.7 Graphs](#roberta_graph)\n    + [6.4.8 Confusion Matrix](#roberta_confusion_matrix)\n    + [6.4.9 Predictions](#roberta_predictions)\n  * [6.5 Roberta Large](#roberta_large)\n    + [6.5.1 Tokenization](#roberta_large_tokenization)\n    + [6.5.2 Model Loading](#roberta_large_loading)\n    + [6.5.3 Model Building](#roberta_large_model)\n    + [6.5.4 Callback](#roberta_large_callback)\n    + [6.5.5 Compile](#roberta_large_compile)\n    + [6.5.6 Training](#roberta_large_training)\n    + [6.5.7 Graphs](#roberta_large_graph)\n    + [6.5.8 Confusion Matrix](#roberta_large_confusion_matrix)\n    + [6.5.9 Predictions](#roberta_large_predictions)\n * [6.6 XLM Roberta](#xlm_roberta)\n- [7. Comparision](#Comaprision)","metadata":{}},{"cell_type":"markdown","source":"<a id = \"imports\"></a>\n **<center><font size = 6><span style=\"color:#2a6592\">1. Imports </span></font></center>**","metadata":{}},{"cell_type":"code","source":"import os\nimport re\nimport nltk\nimport keras.backend as K\nimport random\nimport string\nimport numpy as np \nimport pandas as pd \nimport tensorflow as tf\nimport plotly.graph_objects as go\nimport plotly.express as px\nimport matplotlib.pyplot as plt\nfrom nltk.probability import FreqDist\nfrom sklearn.model_selection import train_test_split\nfrom wordcloud import WordCloud, ImageColorGenerator ,STOPWORDS\nfrom transformers import AutoTokenizer , TFAutoModel\nfrom sklearn.metrics import confusion_matrix , classification_report\nos.environ[\"WANDB_DISABLED\"] = \"true\"\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-07-11T11:09:38.027871Z","iopub.execute_input":"2022-07-11T11:09:38.028333Z","iopub.status.idle":"2022-07-11T11:09:53.963876Z","shell.execute_reply.started":"2022-07-11T11:09:38.028236Z","shell.execute_reply":"2022-07-11T11:09:53.96276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"helper_functions\"></a>\n **<center><font size = 6><span style=\"color:#2a6592\">2. Helper Functions </span></font></center>**","metadata":{}},{"cell_type":"code","source":"def clean_dataset(text):\n    text = text.lower()\n    text = re.sub(r'https?://\\S+|www\\.\\S+', '',text) #Removes Websites\n    text  = re.sub(r'<.*?>' ,'', text) \n    text = re.sub(r'\\x89\\S+' , ' ', text) #Removes string starting from \\x89\n    text = re.sub('\\w*\\d\\w*', '', text)  # Removes numbers\n    text = re.sub(r'[^\\w\\s]','',text)   # Removes Punctuations\n    emoji_pattern = re.compile(\"[\"\n                               u\"\\U0001F600-\\U0001F64F\"  # emoticons\n                               u\"\\U0001F300-\\U0001F5FF\"  # symbols & pictographs\n                               u\"\\U0001F680-\\U0001F6FF\"  # transport & map symbols\n                               u\"\\U0001F1E0-\\U0001F1FF\"  # flags (iOS)\n                               u\"\\U00002500-\\U00002BEF\"  # chinese char\n                               u\"\\U00002702-\\U000027B0\"\n                               u\"\\U00002702-\\U000027B0\"\n                               u\"\\U000024C2-\\U0001F251\"\n                               u\"\\U0001f926-\\U0001f937\"\n                               u\"\\U00010000-\\U0010ffff\"\n                               u\"\\u2640-\\u2642\"\n                               u\"\\u2600-\\u2B55\"\n                               u\"\\u200d\"\n                               u\"\\u23cf\"\n                               u\"\\u23e9\"\n                               u\"\\u231a\"\n                               u\"\\ufe0f\"  # dingbats\n                               u\"\\u3030\"\n                               \"]+\", flags=re.UNICODE)\n    text = emoji_pattern.sub(r'', text)\n    return text\n\ndef Convert(string):\n    li = list(string.lower().split(\" \"))\n    return li\n    ","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-07-11T11:10:09.316066Z","iopub.execute_input":"2022-07-11T11:10:09.316542Z","iopub.status.idle":"2022-07-11T11:10:09.327276Z","shell.execute_reply.started":"2022-07-11T11:10:09.316501Z","shell.execute_reply":"2022-07-11T11:10:09.3256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"hyperparameters\"></a>\n **<center><font size = 6><span style=\"color:#2a6592\">3. Hyperparameters  </span></font></center>**","metadata":{}},{"cell_type":"code","source":"class config:\n    PATH = \"../input/nlp-getting-started/\"\n    TRAIN_PATH = \"../input/nlp-getting-started/train.csv\"\n    TEST_PATH = \"../input/nlp-getting-started/test.csv\"\n    MAX_LEN = 36\n    LOWER_CASE = True\n    RANDOM_STATE = 12\n    TEST_SIZE = 0.2\n    NUM_LABELS = 1\n    BATCH_SIZE = 128\n    LEARNING_RATE = 5e-5\n    EPOCHS = 10\n    WEIGTH_DECAY = 0.01\n    DEVICE = \"cuda\"","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-07-11T11:10:10.036043Z","iopub.execute_input":"2022-07-11T11:10:10.036445Z","iopub.status.idle":"2022-07-11T11:10:10.044318Z","shell.execute_reply.started":"2022-07-11T11:10:10.036412Z","shell.execute_reply":"2022-07-11T11:10:10.043074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"EDA\"></a>\n **<center><font size = 6><span style=\"color:#2a6592\">4. EDA </span></font></center>**","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(config.TRAIN_PATH)\ntest_df = pd.read_csv(config.TEST_PATH)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-11T11:10:10.927374Z","iopub.execute_input":"2022-07-11T11:10:10.92826Z","iopub.status.idle":"2022-07-11T11:10:10.999912Z","shell.execute_reply.started":"2022-07-11T11:10:10.928202Z","shell.execute_reply":"2022-07-11T11:10:10.998725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head(4)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-11T11:10:11.573005Z","iopub.execute_input":"2022-07-11T11:10:11.573701Z","iopub.status.idle":"2022-07-11T11:10:11.596974Z","shell.execute_reply.started":"2022-07-11T11:10:11.573664Z","shell.execute_reply":"2022-07-11T11:10:11.596073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"Random_Examples\"></a>\n<font size = 5><span style=\"color:#db9833\">4.1. Random examples :</span></font>","metadata":{}},{"cell_type":"code","source":"random_index = random.randint(0,len(train_df)-5)\nfor row in train_df[[\"text\",\"target\"]][random_index:random_index +5].itertuples():\n  _,text,target = row\n  print(f\"\\033[2mTarget :\" , f\"\\033[92m{target} (real disaster)\" if target > 0 else f\"\\033[91m {target} (not a real disaster)\")\n  print(f\"\\033[114mText :\\n{text}\\n\")\n  s = clean_dataset(text)\n  print(f\"\\033[30mCleaned Text: \\n{s}\\n\")\n  print(\"-------------------------------------------------------\\n\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-11T11:24:23.368757Z","iopub.execute_input":"2022-07-11T11:24:23.369161Z","iopub.status.idle":"2022-07-11T11:24:23.379733Z","shell.execute_reply.started":"2022-07-11T11:24:23.369128Z","shell.execute_reply":"2022-07-11T11:24:23.378091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[\"text\"] = train_df[\"text\"].map(clean_dataset)\ntrain_df[\"List of Words\"] = train_df[\"text\"].map(Convert)\ntest_df[\"text\"] = test_df[\"text\"].map(clean_dataset)\ntest_df[\"List of Words\"] = test_df[\"text\"].map(Convert)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T18:09:47.762715Z","iopub.execute_input":"2022-07-10T18:09:47.763638Z","iopub.status.idle":"2022-07-10T18:09:48.182463Z","shell.execute_reply.started":"2022-07-10T18:09:47.763597Z","shell.execute_reply":"2022-07-10T18:09:48.18153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"Common_Words\"></a>\n<font size = 5><span style=\"color:#db9833\">4.2. Common Words :</span></font>","metadata":{}},{"cell_type":"code","source":"stop_words = nltk.corpus.stopwords.words(\"english\")\ndisaster_tweets = train_df[train_df[\"target\"]==1][\"text\"].tolist()\nnon_disaster_tweets = train_df[train_df[\"target\"]==0][\"text\"].to_list()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T18:09:48.184194Z","iopub.execute_input":"2022-07-10T18:09:48.184533Z","iopub.status.idle":"2022-07-10T18:09:48.195508Z","shell.execute_reply.started":"2022-07-10T18:09:48.184499Z","shell.execute_reply":"2022-07-10T18:09:48.19454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"disaster_tweets_df = pd.DataFrame(disaster_tweets , columns = [\"text\"])\ndisaster_tweets_df[\"List of Words\"] = disaster_tweets_df[\"text\"].map(Convert)\n\nnon_disaster_tweets_df = pd.DataFrame(non_disaster_tweets , columns = [\"text\"])\nnon_disaster_tweets_df[\"List of Words\"] = non_disaster_tweets_df[\"text\"].map(Convert)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T18:09:48.27156Z","iopub.execute_input":"2022-07-10T18:09:48.271868Z","iopub.status.idle":"2022-07-10T18:09:48.300975Z","shell.execute_reply.started":"2022-07-10T18:09:48.271816Z","shell.execute_reply":"2022-07-10T18:09:48.300019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"disaster_words = disaster_tweets_df[\"List of Words\"]\ndisaster_allwords = []\nfor wordlist in disaster_words:\n    for disaster_word in wordlist:\n        if disaster_word not in stop_words:\n            disaster_allwords.append(disaster_word)\n\n\nnon_disaster_words = non_disaster_tweets_df[\"List of Words\"]\nnon_disaster_allwords = []\nfor wordlist in non_disaster_words:\n    for non_disaster_word in wordlist:\n        if non_disaster_word not in stop_words:\n            non_disaster_allwords.append(non_disaster_word)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T18:09:48.609696Z","iopub.execute_input":"2022-07-10T18:09:48.611154Z","iopub.status.idle":"2022-07-10T18:09:48.809942Z","shell.execute_reply.started":"2022-07-10T18:09:48.611105Z","shell.execute_reply":"2022-07-10T18:09:48.808943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mostcommon = FreqDist(disaster_allwords).most_common(2000)\nwordcloud = WordCloud(width=1800, height=1000, background_color='white').generate(str(mostcommon))\nfig = plt.figure(figsize=(30,10), facecolor='white')\nplt.imshow(wordcloud, interpolation=\"bilinear\")\nplt.axis('off')\nplt.title('Disaster Tweets common words', fontsize=80)\nplt.tight_layout(pad=0)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T18:09:48.997473Z","iopub.execute_input":"2022-07-10T18:09:48.997818Z","iopub.status.idle":"2022-07-10T18:09:52.721648Z","shell.execute_reply.started":"2022-07-10T18:09:48.997789Z","shell.execute_reply":"2022-07-10T18:09:52.720803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mostcommon = FreqDist(non_disaster_allwords).most_common(2000)\nwordcloud = WordCloud(width=1800, height=1000, background_color='white').generate(str(mostcommon))\nfig = plt.figure(figsize=(30,10), facecolor='white')\nplt.imshow(wordcloud, interpolation=\"bilinear\")\nplt.axis('off')\nplt.title('Non Disaster Tweets commmon words', fontsize=80)\nplt.tight_layout(pad=0)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T18:09:52.723652Z","iopub.execute_input":"2022-07-10T18:09:52.724282Z","iopub.status.idle":"2022-07-10T18:09:56.563716Z","shell.execute_reply.started":"2022-07-10T18:09:52.724244Z","shell.execute_reply":"2022-07-10T18:09:56.562821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"dataset\"></a>\n# <b><center><font size = 6><span style=\"color:#2a6592\">5. Dataset </span></font></center></b>","metadata":{}},{"cell_type":"code","source":"test_ids = test_df[\"id\"]\ntrain_df = train_df.drop([\"id\" , \"keyword\" , \"location\" , \"List of Words\"] , axis = 1)\ntest_df = test_df.drop([\"id\" , \"keyword\" , \"location\" , \"List of Words\"], axis =1 )","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-10T18:09:56.565381Z","iopub.execute_input":"2022-07-10T18:09:56.566123Z","iopub.status.idle":"2022-07-10T18:09:56.582121Z","shell.execute_reply.started":"2022-07-10T18:09:56.566083Z","shell.execute_reply":"2022-07-10T18:09:56.58123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"Models\"></a>\n**<center><font size = 6><span style=\"color:#2a6592\">6. Models </span></font></center>**\n\n## <font size = 3><span style=\"color:#2F4F4F\">We are going to train 5 transformer models in this section : </span></font>\n*  <font size = 3><span style=\"color:#2F4F4F\">1. Bert Base Uncased</span></font>\n*  <font size = 3><span style=\"color:#2F4F4F\">2. Bert Large </span></font>\n*  <font size = 3><span style=\"color:#2F4F4F\">3. Distill Bert </span></font>\n*  <font size = 3><span style=\"color:#2F4F4F\">4. Roberta </span></font>\n*  <font size = 3><span style=\"color:#2F4F4F\">5. Roberta Large </span></font>\n","metadata":{}},{"cell_type":"markdown","source":"<a id = \"bert_base_uncased\"></a>\n**<center><font size = 5><span style=\"color:#db9833\">6.1 Bert Base Uncased </span></font></center>**\n* <font size = 3><span style=\"color:#2F4F4F\">The model weights can be found at [Bert Base uncased](https://huggingface.co/bert-base-uncased) </span></font>\n* <font size = 3><span style=\"color:#2F4F4F\"> There are a total of 110M parameters. </span></font>","metadata":{}},{"cell_type":"markdown","source":"<a id = \"bert_base_uncased_tokenization\"></a>\n<font size = 5><span style=\"color:#db9833\">6.1.1 Tokenization: </span></font>","metadata":{}},{"cell_type":"code","source":"K.clear_session()\nMODEL_1 = \"bert-base-uncased\"\ntokenizer = AutoTokenizer.from_pretrained(MODEL_1 , do_lower_case = config.LOWER_CASE , max_length = config.MAX_LEN )\nx_train = tokenizer(\n        text = train_df[\"text\"].tolist(),\n        add_special_tokens = True,\n        max_length = config.MAX_LEN,\n        truncation = True,\n        padding = True,\n        return_tensors = \"tf\",\n        return_token_type_ids = False,\n        return_attention_mask = True,\n        verbose = True\n        )\n\nx_test = tokenizer(\n        text = test_df[\"text\"].tolist(),\n        add_special_tokens = True,\n        max_length = config.MAX_LEN,\n        truncation = True,\n        padding = True,\n        return_tensors = \"tf\",\n        return_token_type_ids = False,\n        return_attention_mask = True,\n        verbose = True\n        )\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:15:45.287903Z","iopub.execute_input":"2022-07-09T15:15:45.28883Z","iopub.status.idle":"2022-07-09T15:15:58.228407Z","shell.execute_reply.started":"2022-07-09T15:15:45.288794Z","shell.execute_reply":"2022-07-09T15:15:58.227385Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_base_uncased_model_loading\"></a>\n<font size = 5><span style=\"color:#db9833\">6.1.2 Loading the pretrained transformer model : </span></font>","metadata":{}},{"cell_type":"code","source":"bert_based_uncased = TFAutoModel.from_pretrained(MODEL_1)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:15:58.229891Z","iopub.execute_input":"2022-07-09T15:15:58.230555Z","iopub.status.idle":"2022-07-09T15:16:14.433372Z","shell.execute_reply.started":"2022-07-09T15:15:58.230516Z","shell.execute_reply":"2022-07-09T15:16:14.432471Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_base_uncased_model\"></a>\n<font size = 5><span style=\"color:#db9833\">6.1.3 Model Building: </span></font>","metadata":{}},{"cell_type":"code","source":"input_ids = tf.keras.layers.Input(shape = (config.MAX_LEN,) , dtype = tf.int32 , name = \"input_ids\")\ninput_mask = tf.keras.layers.Input(shape = (config.MAX_LEN,) , dtype = tf.int32 , name = \"attention_mask\")\nembeddings = bert_based_uncased(input_ids , attention_mask = input_mask)[1]\nx = tf.keras.layers.Dropout(0.3)(embeddings)\nx = tf.keras.layers.Dense(128 , activation = \"relu\")(x)\nx = tf.keras.layers.Dropout(0.2)(x)\nx = tf.keras.layers.Dense(32 , activation = \"relu\")(x)\noutput = tf.keras.layers.Dense(config.NUM_LABELS , activation = \"sigmoid\")(x)\n\nmodel_1 = tf.keras.Model(inputs = [input_ids , input_mask] , outputs = output)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:16:14.434905Z","iopub.execute_input":"2022-07-09T15:16:14.435848Z","iopub.status.idle":"2022-07-09T15:16:21.930913Z","shell.execute_reply.started":"2022-07-09T15:16:14.435807Z","shell.execute_reply":"2022-07-09T15:16:21.929946Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Transformer Layer Unfreezed!!\")\nmodel_1.layers[2].trainable = True\nmodel_1.summary()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:16:21.932409Z","iopub.execute_input":"2022-07-09T15:16:21.933186Z","iopub.status.idle":"2022-07-09T15:16:21.96132Z","shell.execute_reply.started":"2022-07-09T15:16:21.933149Z","shell.execute_reply":"2022-07-09T15:16:21.960355Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_base_uncased_callback\"></a>\n<font size = 5><span style=\"color:#db9833\">6.1.4 Callback: </span></font>","metadata":{}},{"cell_type":"code","source":"if  os.path.isdir(\"./weights/bert_base_uncased_weights\") is None:\n          os.makedirs(\"./weights/bert_base_uncased_weights\")\ncheckpoint_filepath_bert_base_uncased  = \"./weights/bert_base_uncased_weights\"\ncheckpoint_callback_bert_base_uncased = tf.keras.callbacks.ModelCheckpoint(\n    checkpoint_filepath_bert_base_uncased,\n    save_weights_only=True,\n    monitor='val_accuracy',\n    mode='auto',\n    save_best_only=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:16:21.962832Z","iopub.execute_input":"2022-07-09T15:16:21.963836Z","iopub.status.idle":"2022-07-09T15:16:22.335362Z","shell.execute_reply.started":"2022-07-09T15:16:21.963796Z","shell.execute_reply":"2022-07-09T15:16:22.334138Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_base_uncased_compile\"></a>\n<font size = 5><span style=\"color:#db9833\">6.1.5 Compile: </span></font>","metadata":{}},{"cell_type":"code","source":"model_1.compile(loss = tf.keras.losses.BinaryCrossentropy(from_logits = True), \n             optimizer = tf.keras.optimizers.Adam(lr = config.LEARNING_RATE , epsilon = 1e-8 , decay  =config.WEIGTH_DECAY , clipnorm = 1.0),\n             metrics = [\"accuracy\"])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:16:22.337124Z","iopub.execute_input":"2022-07-09T15:16:22.337872Z","iopub.status.idle":"2022-07-09T15:16:22.361824Z","shell.execute_reply.started":"2022-07-09T15:16:22.337835Z","shell.execute_reply":"2022-07-09T15:16:22.360938Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_base_uncased_training\"></a>\n<font size = 5><span style=\"color:#db9833\">6.1.6 Training: </span></font>","metadata":{}},{"cell_type":"code","source":"bert_based_uncased_history  = model_1.fit(x = {\"input_ids\": x_train[\"input_ids\"] , \"attention_mask\" : x_train[\"attention_mask\"]},\n                y = train_df[\"target\"] , \n                epochs = config.EPOCHS , \n                validation_split = 0.2,\n                batch_size = 256 , callbacks = [checkpoint_callback_bert_base_uncased])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:16:22.363723Z","iopub.execute_input":"2022-07-09T15:16:22.364365Z","iopub.status.idle":"2022-07-09T15:21:03.789924Z","shell.execute_reply.started":"2022-07-09T15:16:22.36433Z","shell.execute_reply":"2022-07-09T15:21:03.788911Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_1.load_weights(checkpoint_filepath_bert_base_uncased)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:21:03.792068Z","iopub.execute_input":"2022-07-09T15:21:03.792917Z","iopub.status.idle":"2022-07-09T15:21:05.16632Z","shell.execute_reply.started":"2022-07-09T15:21:03.792875Z","shell.execute_reply":"2022-07-09T15:21:05.165337Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bert_based_uncased_hist_df = pd.DataFrame(bert_based_uncased_history.history , columns = ['loss', 'accuracy', 'val_loss', 'val_accuracy'])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:21:05.167565Z","iopub.execute_input":"2022-07-09T15:21:05.168533Z","iopub.status.idle":"2022-07-09T15:21:05.175534Z","shell.execute_reply.started":"2022-07-09T15:21:05.168494Z","shell.execute_reply":"2022-07-09T15:21:05.17454Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_base_uncased_graph\"></a>\n<font size = 5><span style=\"color:#db9833\">6.1.7 Graphs: </span></font>","metadata":{}},{"cell_type":"code","source":"fig = px.line(bert_based_uncased_hist_df, y=[\"accuracy\" , \"val_accuracy\"], title=\"Accuracy\") \nfig.update_xaxes(title=\"Epochs\")\nfig.update_yaxes(title = \"Accuracy\")\nfig.update_layout(showlegend = True,\n        title = {\n            'text': \"Bert Base uncased Accuracy\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'})\nfig.show()\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:21:05.176948Z","iopub.execute_input":"2022-07-09T15:21:05.177363Z","iopub.status.idle":"2022-07-09T15:21:05.937623Z","shell.execute_reply.started":"2022-07-09T15:21:05.177328Z","shell.execute_reply":"2022-07-09T15:21:05.936723Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.line(bert_based_uncased_hist_df, y=[\"loss\" , \"val_loss\"], title=\"Loss\") \nfig.update_xaxes(title=\"Epochs\")\nfig.update_yaxes(title = \"Loss\")\nfig.update_layout(showlegend = True,\n        title = {\n            'text': \"Bert Base uncased Loss\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'})\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:21:05.945395Z","iopub.execute_input":"2022-07-09T15:21:05.945694Z","iopub.status.idle":"2022-07-09T15:21:06.019495Z","shell.execute_reply.started":"2022-07-09T15:21:05.945668Z","shell.execute_reply":"2022-07-09T15:21:06.018449Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_base_uncased_confusion_matrix\"></a>\n<font size = 5><span style=\"color:#db9833\">6.1.8 Confusion Matrix: </span></font>","metadata":{}},{"cell_type":"code","source":"y_pred = model_1.predict({\"input_ids\" : x_train[\"input_ids\"] ,\"attention_mask\" : x_train[\"attention_mask\"]})\ny_pred = np.where(y_pred > 0.5 , 1,0)\ny_test = train_df[\"target\"]\nCLASS_LABELS = [\"Disaster Tweet\" , \"Non Disaster Tweet\"]\ncm_data = confusion_matrix(y_test , y_pred)\ncm = pd.DataFrame(cm_data , columns = CLASS_LABELS , index = CLASS_LABELS)\nfig = px.imshow(img = cm_data ,\n                x = CLASS_LABELS, \n                y = CLASS_LABELS,\n                aspect=\"auto\" , \n                color_continuous_scale = \"mint\")\nfig.update_xaxes(title=\"Predicted\")\nfig.update_yaxes(title = \"Actual\")\nfig.update_layout(title = \"Confusion Matrix\",\n                  template = \"plotly_white\",\n                  title_x = 0.5)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:21:06.02121Z","iopub.execute_input":"2022-07-09T15:21:06.021586Z","iopub.status.idle":"2022-07-09T15:21:22.486929Z","shell.execute_reply.started":"2022-07-09T15:21:06.021549Z","shell.execute_reply":"2022-07-09T15:21:22.486Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(classification_report(y_test , y_pred))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:21:22.491306Z","iopub.execute_input":"2022-07-09T15:21:22.493868Z","iopub.status.idle":"2022-07-09T15:21:22.526681Z","shell.execute_reply.started":"2022-07-09T15:21:22.493831Z","shell.execute_reply":"2022-07-09T15:21:22.525731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_base_uncased_predictions\"></a>\n<font size = 5><span style=\"color:#db9833\">6.1.9 Predictions: </span></font>","metadata":{}},{"cell_type":"code","source":"model_1_pred_probs = model_1.predict({\"input_ids\" : x_test[\"input_ids\"] ,\"attention_mask\" : x_test[\"attention_mask\"]})\ny_pred_1 = np.where(model_1_pred_probs > 0.5 , 1,0)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:21:22.528559Z","iopub.execute_input":"2022-07-09T15:21:22.529182Z","iopub.status.idle":"2022-07-09T15:21:28.508708Z","shell.execute_reply.started":"2022-07-09T15:21:22.529125Z","shell.execute_reply":"2022-07-09T15:21:28.507585Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bert_base_uncased_df=pd.DataFrame()\nbert_base_uncased_df['id'] = test_ids\nbert_base_uncased_df['target'] = y_pred_1","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:21:28.5108Z","iopub.execute_input":"2022-07-09T15:21:28.511144Z","iopub.status.idle":"2022-07-09T15:21:28.518391Z","shell.execute_reply.started":"2022-07-09T15:21:28.511108Z","shell.execute_reply":"2022-07-09T15:21:28.517488Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bert_base_uncased_df.to_csv('bert_base_uncased.csv',index = False)\nbert_base_uncased_df[\"target\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:21:28.52Z","iopub.execute_input":"2022-07-09T15:21:28.520622Z","iopub.status.idle":"2022-07-09T15:21:28.545713Z","shell.execute_reply.started":"2022-07-09T15:21:28.520585Z","shell.execute_reply":"2022-07-09T15:21:28.544749Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_large\"></a>\n**<center><font size = 5><span style=\"color:#db9833\">6.2 Bert Large </span></font></center>**\n* <font size = 3><span style=\"color:#2F4F4F\">The model weights can be found at [Bert Large uncased](https://huggingface.co/bert-large-uncased) </span></font>\n* <font size = 3><span style=\"color:#2F4F4F\"> There are a total of 340M parameters. </span></font>","metadata":{}},{"cell_type":"markdown","source":"<a id = \"bert_large_tokenization\"></a>\n<font size = 5><span style=\"color:#db9833\">6.2.1 Tokenization: </span></font>","metadata":{}},{"cell_type":"code","source":"K.clear_session()\nMODEL_2 = \"bert-large-uncased\"\ntokenizer = AutoTokenizer.from_pretrained(MODEL_2 , do_lower_case = config.LOWER_CASE , max_length = config.MAX_LEN )\nx_train = tokenizer(\n        text = train_df[\"text\"].tolist(),\n        add_special_tokens = True,\n        max_length = config.MAX_LEN,\n        truncation = True,\n        padding = True,\n        return_tensors = \"tf\",\n        return_token_type_ids = False,\n        return_attention_mask = True,\n        verbose = True\n        )\n\nx_test = tokenizer(\n        text = test_df[\"text\"].tolist(),\n        add_special_tokens = True,\n        max_length = config.MAX_LEN,\n        truncation = True,\n        padding = True,\n        return_tensors = \"tf\",\n        return_token_type_ids = False,\n        return_attention_mask = True,\n        verbose = True\n        )\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:21:28.547097Z","iopub.execute_input":"2022-07-09T15:21:28.547649Z","iopub.status.idle":"2022-07-09T15:21:34.549382Z","shell.execute_reply.started":"2022-07-09T15:21:28.547614Z","shell.execute_reply":"2022-07-09T15:21:34.548297Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_large_model_loading\"></a>\n<font size = 5><span style=\"color:#db9833\">6.2.2 Loading the pretrained transformer model: </span></font>","metadata":{}},{"cell_type":"code","source":"bert_large = TFAutoModel.from_pretrained(MODEL_2)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:21:34.55327Z","iopub.execute_input":"2022-07-09T15:21:34.554771Z","iopub.status.idle":"2022-07-09T15:22:13.62501Z","shell.execute_reply.started":"2022-07-09T15:21:34.554729Z","shell.execute_reply":"2022-07-09T15:22:13.623984Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_large_model\"></a>\n<font size = 5><span style=\"color:#db9833\">6.2.3 Model Building: </span></font>","metadata":{}},{"cell_type":"code","source":"input_ids = tf.keras.layers.Input(shape = (config.MAX_LEN,) , dtype = tf.int32 , name = \"input_ids\")\ninput_mask = tf.keras.layers.Input(shape = (config.MAX_LEN,) , dtype = tf.int32 , name = \"attention_mask\")\nembeddings = bert_large(input_ids , attention_mask = input_mask)[1]\nx = tf.keras.layers.Dropout(0.3)(embeddings)\nx = tf.keras.layers.Dense(128 , activation = \"relu\")(x)\nx = tf.keras.layers.Dropout(0.2)(x)\nx = tf.keras.layers.Dense(32 , activation = \"relu\")(x)\noutput = tf.keras.layers.Dense(config.NUM_LABELS , activation = \"sigmoid\")(x)\n\nmodel_2 = tf.keras.Model(inputs = [input_ids , input_mask] , outputs = output)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:22:13.626643Z","iopub.execute_input":"2022-07-09T15:22:13.626997Z","iopub.status.idle":"2022-07-09T15:22:16.401045Z","shell.execute_reply.started":"2022-07-09T15:22:13.626959Z","shell.execute_reply":"2022-07-09T15:22:16.400115Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Transformer Layer Unfreezed!!\")\nmodel_2.layers[2].trainable = True\nmodel_2.summary()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:22:16.402286Z","iopub.execute_input":"2022-07-09T15:22:16.402631Z","iopub.status.idle":"2022-07-09T15:22:16.447675Z","shell.execute_reply.started":"2022-07-09T15:22:16.402598Z","shell.execute_reply":"2022-07-09T15:22:16.446717Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_large_callback\"></a>\n<font size = 5><span style=\"color:#db9833\">6.2.4 Callback: </span></font>","metadata":{}},{"cell_type":"code","source":"if  os.path.isdir(\"./weights/bert_large_weights\") is None:\n          os.makedirs(\"./weights/bert_large_weights\")\ncheckpoint_filepath_bert_large  = \"./weights/bert_large_weights\"\ncheckpoint_callback_bert_large = tf.keras.callbacks.ModelCheckpoint(\n    checkpoint_filepath_bert_large,\n    save_weights_only=True,\n    monitor='val_accuracy',\n    mode='auto',\n    save_best_only=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:22:16.448926Z","iopub.execute_input":"2022-07-09T15:22:16.449357Z","iopub.status.idle":"2022-07-09T15:22:16.455553Z","shell.execute_reply.started":"2022-07-09T15:22:16.449321Z","shell.execute_reply":"2022-07-09T15:22:16.454359Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_large_compile\"></a>\n<font size = 5><span style=\"color:#db9833\">6.2.5 Compile: </span></font>","metadata":{}},{"cell_type":"code","source":"model_2.compile(loss = tf.keras.losses.BinaryCrossentropy(from_logits = True), \n             optimizer = tf.keras.optimizers.Adam(lr = config.LEARNING_RATE , epsilon = 1e-8 , decay  =config.WEIGTH_DECAY , clipnorm = 1.0),\n             metrics = [\"accuracy\"])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:22:16.456754Z","iopub.execute_input":"2022-07-09T15:22:16.457719Z","iopub.status.idle":"2022-07-09T15:22:16.478852Z","shell.execute_reply.started":"2022-07-09T15:22:16.457677Z","shell.execute_reply":"2022-07-09T15:22:16.477943Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_large_training\"></a>\n<font size = 5><span style=\"color:#db9833\">6.2.6 Training: </span></font>","metadata":{}},{"cell_type":"code","source":"bert_large_history  = model_2.fit(x = {\"input_ids\": x_train[\"input_ids\"] , \"attention_mask\" : x_train[\"attention_mask\"]},\n                y = train_df[\"target\"] , \n                epochs = config.EPOCHS , \n                validation_split = 0.2,\n                batch_size = 32 , callbacks = [checkpoint_callback_bert_large])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:22:16.482606Z","iopub.execute_input":"2022-07-09T15:22:16.484626Z","iopub.status.idle":"2022-07-09T15:40:43.692679Z","shell.execute_reply.started":"2022-07-09T15:22:16.4846Z","shell.execute_reply":"2022-07-09T15:40:43.691752Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_2.load_weights(checkpoint_filepath_bert_large)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:40:43.694576Z","iopub.execute_input":"2022-07-09T15:40:43.694935Z","iopub.status.idle":"2022-07-09T15:40:55.69822Z","shell.execute_reply.started":"2022-07-09T15:40:43.694885Z","shell.execute_reply":"2022-07-09T15:40:55.697282Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bert_large_hist_df = pd.DataFrame(bert_large_history.history , columns = ['loss', 'accuracy', 'val_loss', 'val_accuracy'])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:40:55.699678Z","iopub.execute_input":"2022-07-09T15:40:55.700299Z","iopub.status.idle":"2022-07-09T15:40:55.707126Z","shell.execute_reply.started":"2022-07-09T15:40:55.700255Z","shell.execute_reply":"2022-07-09T15:40:55.706142Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_large_graph\"></a>\n<font size = 5><span style=\"color:#db9833\">6.2.7 Graphs: </span></font>","metadata":{}},{"cell_type":"code","source":"fig = px.line(bert_large_hist_df, y=[\"accuracy\" , \"val_accuracy\"], title=\"Accuracy\") \nfig.update_xaxes(title=\"Epochs\")\nfig.update_yaxes(title = \"Accuracy\")\nfig.update_layout(showlegend = True,\n        title = {\n            'text': \"Bert Large Accuracy\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'})\nfig.show()\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:40:55.708659Z","iopub.execute_input":"2022-07-09T15:40:55.709277Z","iopub.status.idle":"2022-07-09T15:40:55.793075Z","shell.execute_reply.started":"2022-07-09T15:40:55.709241Z","shell.execute_reply":"2022-07-09T15:40:55.791965Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.line(bert_large_hist_df, y=[\"loss\" , \"val_loss\"], title=\"Loss\") \nfig.update_xaxes(title=\"Epochs\")\nfig.update_yaxes(title = \"Loss\")\nfig.update_layout(showlegend = True,\n        title = {\n            'text': \"Bert Large Loss\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'})\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:40:55.794755Z","iopub.execute_input":"2022-07-09T15:40:55.79511Z","iopub.status.idle":"2022-07-09T15:40:55.868767Z","shell.execute_reply.started":"2022-07-09T15:40:55.795077Z","shell.execute_reply":"2022-07-09T15:40:55.867731Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_large_confusion_matrix\"></a>\n<font size = 5><span style=\"color:#db9833\">6.2.8 Confusion Matrix: </span></font>","metadata":{}},{"cell_type":"code","source":"y_pred = model_2.predict({\"input_ids\" : x_train[\"input_ids\"] ,\"attention_mask\" : x_train[\"attention_mask\"]})\ny_pred = np.where(y_pred > 0.5 , 1,0)\ny_test = train_df[\"target\"]\nCLASS_LABELS = [\"Disaster Tweet\" , \"Non Disaster Tweet\"]\ncm_data = confusion_matrix(y_test , y_pred)\ncm = pd.DataFrame(cm_data , columns = CLASS_LABELS , index = CLASS_LABELS)\nfig = px.imshow(img = cm_data ,\n                x = CLASS_LABELS, \n                y = CLASS_LABELS,\n                aspect=\"auto\" , \n                color_continuous_scale = \"mint\")\nfig.update_xaxes(title=\"Predicted\")\nfig.update_yaxes(title = \"Actual\")\nfig.update_layout(title = \"Confusion Matrix\",\n                  template = \"plotly_white\",\n                  title_x = 0.5)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:40:55.870478Z","iopub.execute_input":"2022-07-09T15:40:55.870814Z","iopub.status.idle":"2022-07-09T15:41:42.730367Z","shell.execute_reply.started":"2022-07-09T15:40:55.870778Z","shell.execute_reply":"2022-07-09T15:41:42.729413Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(classification_report(y_test , y_pred))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:41:42.731845Z","iopub.execute_input":"2022-07-09T15:41:42.732191Z","iopub.status.idle":"2022-07-09T15:41:42.755187Z","shell.execute_reply.started":"2022-07-09T15:41:42.732157Z","shell.execute_reply":"2022-07-09T15:41:42.754098Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"bert_large_predictions\"></a>\n<font size = 5><span style=\"color:#db9833\">6.2.9 Predictions: </span></font>","metadata":{}},{"cell_type":"code","source":"model_2_pred_probs = model_2.predict({\"input_ids\" : x_test[\"input_ids\"] ,\"attention_mask\" : x_test[\"attention_mask\"]})\ny_pred_2 = np.where(model_2_pred_probs > 0.5 , 1,0)\ny_pred_2","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:41:42.756772Z","iopub.execute_input":"2022-07-09T15:41:42.75774Z","iopub.status.idle":"2022-07-09T15:42:00.468714Z","shell.execute_reply.started":"2022-07-09T15:41:42.757703Z","shell.execute_reply":"2022-07-09T15:42:00.467797Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bert_large_df=pd.DataFrame()\nbert_large_df['id'] = test_ids\nbert_large_df['target'] = y_pred_2","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:42:00.470233Z","iopub.execute_input":"2022-07-09T15:42:00.470602Z","iopub.status.idle":"2022-07-09T15:42:00.477712Z","shell.execute_reply.started":"2022-07-09T15:42:00.470567Z","shell.execute_reply":"2022-07-09T15:42:00.476758Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bert_large_df.to_csv('bert_large.csv',index = False)\nbert_large_df[\"target\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:42:00.479647Z","iopub.execute_input":"2022-07-09T15:42:00.480382Z","iopub.status.idle":"2022-07-09T15:42:00.498217Z","shell.execute_reply.started":"2022-07-09T15:42:00.480344Z","shell.execute_reply":"2022-07-09T15:42:00.497395Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"distill_bert\"></a>\n**<center><font size = 5><span style=\"color:#db9833\">6.3 Distill Bert Uncased </span></font></center>**\n* <font size = 3><span style=\"color:#2F4F4F\">The model weights can be found at [Distill Bert uncased](https://huggingface.co/distilbert-base-uncased) </span></font>\n* <font size = 3><span style=\"color:#2F4F4F\"> There are a total of 66M parameters. </span></font>","metadata":{}},{"cell_type":"markdown","source":"<a id = \"dsitll_bert_tokenization\"></a>\n<font size = 5><span style=\"color:#db9833\">6.3.1 Tokenization: </span></font>","metadata":{}},{"cell_type":"code","source":"K.clear_session()\nMODEL_3 = \"distilbert-base-uncased\"\ntokenizer = AutoTokenizer.from_pretrained(MODEL_3, do_lower_case = config.LOWER_CASE , max_length = config.MAX_LEN )\nx_train = tokenizer(\n        text = train_df[\"text\"].tolist(),\n        add_special_tokens = True,\n        max_length = config.MAX_LEN,\n        truncation = True,\n        padding = True,\n        return_tensors = \"tf\",\n        return_token_type_ids = False,\n        return_attention_mask = True,\n        verbose = True\n        )\n\nx_test = tokenizer(\n        text = test_df[\"text\"].tolist(),\n        add_special_tokens = True,\n        max_length = config.MAX_LEN,\n        truncation = True,\n        padding = True,\n        return_tensors = \"tf\",\n        return_token_type_ids = False,\n        return_attention_mask = True,\n        verbose = True\n        )","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:42:00.499421Z","iopub.execute_input":"2022-07-09T15:42:00.500233Z","iopub.status.idle":"2022-07-09T15:42:06.934124Z","shell.execute_reply.started":"2022-07-09T15:42:00.500193Z","shell.execute_reply":"2022-07-09T15:42:06.933139Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"dsitill_bert_model_loading\"></a>\n<font size = 5><span style=\"color:#db9833\">6.3.2 Loading the pretrained transformer model : </span></font>","metadata":{}},{"cell_type":"code","source":"distill_bert_uncased = TFAutoModel.from_pretrained(MODEL_3)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:42:06.935958Z","iopub.execute_input":"2022-07-09T15:42:06.936291Z","iopub.status.idle":"2022-07-09T15:42:16.271411Z","shell.execute_reply.started":"2022-07-09T15:42:06.936257Z","shell.execute_reply":"2022-07-09T15:42:16.270339Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"dsitill_bert_model\"></a>\n<font size = 5><span style=\"color:#db9833\">6.3.3 Model Building: </span></font>","metadata":{}},{"cell_type":"code","source":"input_ids = tf.keras.layers.Input(shape = (config.MAX_LEN,) , dtype = tf.int32 , name = \"input_ids\")\ninput_mask = tf.keras.layers.Input(shape = (config.MAX_LEN,) , dtype = tf.int32 , name = \"attention_mask\")\nembeddings = distill_bert_uncased(input_ids , attention_mask = input_mask)[0]\nx = tf.keras.layers.GlobalAveragePooling1D()(embeddings)\nx = tf.keras.layers.Dropout(0.3)(x)\nx = tf.keras.layers.Dense(128 , activation = \"relu\")(x)\nx = tf.keras.layers.Dropout(0.2)(x)\nx = tf.keras.layers.Dense(32 , activation = \"relu\")(x)\noutput = tf.keras.layers.Dense(config.NUM_LABELS , activation = \"sigmoid\")(x)\n\nmodel_3 = tf.keras.Model(inputs = [input_ids , input_mask] , outputs = output)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:42:16.275773Z","iopub.execute_input":"2022-07-09T15:42:16.276113Z","iopub.status.idle":"2022-07-09T15:42:18.395936Z","shell.execute_reply.started":"2022-07-09T15:42:16.276079Z","shell.execute_reply":"2022-07-09T15:42:18.394939Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Transformer Layer Unfreezed!!\")\nmodel_3.layers[2].trainable = True\nmodel_3.summary()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:42:18.397532Z","iopub.execute_input":"2022-07-09T15:42:18.398112Z","iopub.status.idle":"2022-07-09T15:42:18.415015Z","shell.execute_reply.started":"2022-07-09T15:42:18.39807Z","shell.execute_reply":"2022-07-09T15:42:18.414008Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"distill_bert_callback\"></a>\n<font size = 5><span style=\"color:#db9833\">6.3.4 Callback: </span></font>","metadata":{}},{"cell_type":"code","source":"if  os.path.isdir(\"./weights/distill_bert_weights\") is None:\n          os.makedirs(\"./weights/distill_bert_weights\")\ncheckpoint_filepath_distill_bert  = \"./weights/distill_bert_weights\"\ncheckpoint_callback_distill_bert = tf.keras.callbacks.ModelCheckpoint(\n    checkpoint_filepath_distill_bert,\n    save_weights_only=True,\n    monitor='val_accuracy',\n    mode='auto',\n    save_best_only=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:42:18.416663Z","iopub.execute_input":"2022-07-09T15:42:18.417464Z","iopub.status.idle":"2022-07-09T15:42:18.680881Z","shell.execute_reply.started":"2022-07-09T15:42:18.417403Z","shell.execute_reply":"2022-07-09T15:42:18.679502Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"dsitill_bert_compile\"></a>\n<font size = 5><span style=\"color:#db9833\">6.3.5 Compile: </span></font>","metadata":{}},{"cell_type":"code","source":"model_3.compile(loss = tf.keras.losses.BinaryCrossentropy(from_logits = True), \n             optimizer = tf.keras.optimizers.Adam(lr = config.LEARNING_RATE , epsilon = 1e-8 , decay  =config.WEIGTH_DECAY , clipnorm = 1.0),\n             metrics = [\"accuracy\"])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:42:18.682399Z","iopub.execute_input":"2022-07-09T15:42:18.682877Z","iopub.status.idle":"2022-07-09T15:42:18.69936Z","shell.execute_reply.started":"2022-07-09T15:42:18.682838Z","shell.execute_reply":"2022-07-09T15:42:18.698372Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"dsitill_bert_training\"></a>\n<font size = 5><span style=\"color:#db9833\">6.3.6 Training: </span></font>","metadata":{}},{"cell_type":"code","source":"distill_bert_uncased_history  = model_3.fit(x = {\"input_ids\": x_train[\"input_ids\"] , \"attention_mask\" : x_train[\"attention_mask\"]},\n                y = train_df[\"target\"] , \n                epochs = config.EPOCHS , \n                validation_split = 0.2,\n                batch_size = 256 , callbacks = [checkpoint_callback_distill_bert])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:42:18.701181Z","iopub.execute_input":"2022-07-09T15:42:18.701635Z","iopub.status.idle":"2022-07-09T15:44:49.776505Z","shell.execute_reply.started":"2022-07-09T15:42:18.701599Z","shell.execute_reply":"2022-07-09T15:44:49.775481Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_3.load_weights(checkpoint_filepath_distill_bert)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:44:49.777875Z","iopub.execute_input":"2022-07-09T15:44:49.778504Z","iopub.status.idle":"2022-07-09T15:44:51.140296Z","shell.execute_reply.started":"2022-07-09T15:44:49.778465Z","shell.execute_reply":"2022-07-09T15:44:51.139121Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"distill_bert_uncased_hist_df = pd.DataFrame(distill_bert_uncased_history.history , columns = ['loss', 'accuracy', 'val_loss', 'val_accuracy'])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:44:51.145647Z","iopub.execute_input":"2022-07-09T15:44:51.146016Z","iopub.status.idle":"2022-07-09T15:44:51.160881Z","shell.execute_reply.started":"2022-07-09T15:44:51.145983Z","shell.execute_reply":"2022-07-09T15:44:51.159286Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"distill_bert_graph\"></a>\n<font size = 5><span style=\"color:#db9833\">6.3.7 Graphs: </span></font>","metadata":{}},{"cell_type":"code","source":"fig = px.line(distill_bert_uncased_hist_df, y=[\"accuracy\" , \"val_accuracy\"], title=\"Accuracy\") \nfig.update_xaxes(title=\"Epochs\")\nfig.update_yaxes(title = \"Accuracy\")\nfig.update_layout(showlegend = True,\n        title = {\n            'text': \"Distill Bert uncased Accuracy\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'})\nfig.show()\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:44:51.166886Z","iopub.execute_input":"2022-07-09T15:44:51.169794Z","iopub.status.idle":"2022-07-09T15:44:51.315066Z","shell.execute_reply.started":"2022-07-09T15:44:51.169755Z","shell.execute_reply":"2022-07-09T15:44:51.314053Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.line(distill_bert_uncased_hist_df, y=[\"loss\" , \"val_loss\"], title=\"Loss\") \nfig.update_xaxes(title=\"Epochs\")\nfig.update_yaxes(title = \"Loss\")\nfig.update_layout(showlegend = True,\n        title = {\n            'text': \"Distill Bert uncased Loss\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'})\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:44:51.320128Z","iopub.execute_input":"2022-07-09T15:44:51.322745Z","iopub.status.idle":"2022-07-09T15:44:51.457885Z","shell.execute_reply.started":"2022-07-09T15:44:51.322706Z","shell.execute_reply":"2022-07-09T15:44:51.456251Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"distill_bert_confusion_matrix\"></a>\n<font size = 5><span style=\"color:#db9833\">6.3.8 Confusion Matrix: </span></font>","metadata":{}},{"cell_type":"code","source":"y_pred = model_3.predict({\"input_ids\" : x_train[\"input_ids\"] ,\"attention_mask\" : x_train[\"attention_mask\"]})\ny_pred = np.where(y_pred > 0.5 , 1,0)\ny_test = train_df[\"target\"]\nCLASS_LABELS = [\"Disaster Tweet\" , \"Non Disaster Tweet\"]\ncm_data = confusion_matrix(y_test , y_pred)\ncm = pd.DataFrame(cm_data , columns = CLASS_LABELS , index = CLASS_LABELS)\nfig = px.imshow(img = cm_data ,\n                x = CLASS_LABELS, \n                y = CLASS_LABELS,\n                aspect=\"auto\" , \n                color_continuous_scale = \"mint\")\nfig.update_xaxes(title=\"Predicted\")\nfig.update_yaxes(title = \"Actual\")\nfig.update_layout(title = \"Confusion Matrix\",\n                  template = \"plotly_white\",\n                  title_x = 0.5)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:44:51.463382Z","iopub.execute_input":"2022-07-09T15:44:51.466191Z","iopub.status.idle":"2022-07-09T15:45:00.453397Z","shell.execute_reply.started":"2022-07-09T15:44:51.46614Z","shell.execute_reply":"2022-07-09T15:45:00.452318Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(classification_report(y_test , y_pred))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:45:00.45514Z","iopub.execute_input":"2022-07-09T15:45:00.455832Z","iopub.status.idle":"2022-07-09T15:45:00.481646Z","shell.execute_reply.started":"2022-07-09T15:45:00.455789Z","shell.execute_reply":"2022-07-09T15:45:00.480416Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"dsitill_bert_predictions\"></a>\n<font size = 5><span style=\"color:#db9833\">6.3.9 Predictions: </span></font>","metadata":{}},{"cell_type":"code","source":"model_3_pred_probs = model_3.predict({\"input_ids\" : x_test[\"input_ids\"] ,\"attention_mask\" : x_test[\"attention_mask\"]})\ny_pred_3 = np.where(model_3_pred_probs > 0.5 , 1,0)\ny_pred_3","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:45:00.483208Z","iopub.execute_input":"2022-07-09T15:45:00.484206Z","iopub.status.idle":"2022-07-09T15:45:03.59733Z","shell.execute_reply.started":"2022-07-09T15:45:00.484174Z","shell.execute_reply":"2022-07-09T15:45:03.596464Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"distill_bert_uncased_df=pd.DataFrame()\ndistill_bert_uncased_df['id'] = test_ids\ndistill_bert_uncased_df['target'] = y_pred_2","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:45:03.598876Z","iopub.execute_input":"2022-07-09T15:45:03.599233Z","iopub.status.idle":"2022-07-09T15:45:03.608663Z","shell.execute_reply.started":"2022-07-09T15:45:03.599198Z","shell.execute_reply":"2022-07-09T15:45:03.607775Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"distill_bert_uncased_df.to_csv('distill_bert_uncased.csv',index = False)\ndistill_bert_uncased_df[\"target\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:45:03.610307Z","iopub.execute_input":"2022-07-09T15:45:03.610662Z","iopub.status.idle":"2022-07-09T15:45:03.626468Z","shell.execute_reply.started":"2022-07-09T15:45:03.610629Z","shell.execute_reply":"2022-07-09T15:45:03.625483Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_base\"></a>\n**<center><font size = 5><span style=\"color:#db9833\">6.4 Roberta Base </span></font></center>**\n* <font size = 3><span style=\"color:#2F4F4F\">The model weights can be found at [Roberta Base](https://huggingface.co/roberta-base) </span></font>\n* <font size = 3><span style=\"color:#2F4F4F\"> There are a total of 125M parameters. </span></font>","metadata":{}},{"cell_type":"markdown","source":"<a id = \"roberta_tokenization\"></a>\n<font size = 5><span style=\"color:#db9833\">6.4.1 Tokenization: </span></font>","metadata":{}},{"cell_type":"code","source":"K.clear_session()\nMODEL_4 = \"roberta-base\"\ntokenizer = AutoTokenizer.from_pretrained(MODEL_4 , do_lower_case = config.LOWER_CASE , max_length = config.MAX_LEN )\nx_train = tokenizer(\n        text = train_df[\"text\"].tolist(),\n        add_special_tokens = True,\n        max_length = config.MAX_LEN,\n        truncation = True,\n        padding = True,\n        return_tensors = \"tf\",\n        return_token_type_ids = False,\n        return_attention_mask = True,\n        verbose = True\n        )\n\nx_test = tokenizer(\n        text = test_df[\"text\"].tolist(),\n        add_special_tokens = True,\n        max_length = config.MAX_LEN,\n        truncation = True,\n        padding = True,\n        return_tensors = \"tf\",\n        return_token_type_ids = False,\n        return_attention_mask = True,\n        verbose = True\n        )\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:45:03.628116Z","iopub.execute_input":"2022-07-09T15:45:03.628551Z","iopub.status.idle":"2022-07-09T15:45:10.420973Z","shell.execute_reply.started":"2022-07-09T15:45:03.628517Z","shell.execute_reply":"2022-07-09T15:45:10.41989Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_model_loading\"></a>\n<font size = 5><span style=\"color:#db9833\">6.4.2 Loading the pretrained transformer model : </span></font>","metadata":{}},{"cell_type":"code","source":"roberta_base = TFAutoModel.from_pretrained(MODEL_4)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:45:10.422839Z","iopub.execute_input":"2022-07-09T15:45:10.423516Z","iopub.status.idle":"2022-07-09T15:45:28.358472Z","shell.execute_reply.started":"2022-07-09T15:45:10.423477Z","shell.execute_reply":"2022-07-09T15:45:28.357498Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_model\"></a>\n<font size = 5><span style=\"color:#db9833\">6.4.3 Model Building: </span></font>","metadata":{}},{"cell_type":"code","source":"input_ids = tf.keras.layers.Input(shape = (config.MAX_LEN,) , dtype = tf.int32 , name = \"input_ids\")\ninput_mask = tf.keras.layers.Input(shape = (config.MAX_LEN,) , dtype = tf.int32 , name = \"attention_mask\")\nembeddings = roberta_base(input_ids , attention_mask = input_mask)[1]\n# x = tf.keras.layers.GlobalAveragePooling1D()(embeddings)\nx = tf.keras.layers.Dropout(0.3)(embeddings)\nx = tf.keras.layers.Dense(128 , activation = \"relu\")(x)\nx = tf.keras.layers.Dropout(0.2)(x)\nx = tf.keras.layers.Dense(32 , activation = \"relu\")(x)\noutput = tf.keras.layers.Dense(config.NUM_LABELS , activation = \"sigmoid\")(x)\n\nmodel_4 = tf.keras.Model(inputs = [input_ids , input_mask] , outputs = output)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:45:28.359985Z","iopub.execute_input":"2022-07-09T15:45:28.360802Z","iopub.status.idle":"2022-07-09T15:45:32.055048Z","shell.execute_reply.started":"2022-07-09T15:45:28.360762Z","shell.execute_reply":"2022-07-09T15:45:32.054079Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Transformer Layer freezed!!\")\nmodel_4.layers[2].trainable = True\nmodel_4.summary()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:45:32.056499Z","iopub.execute_input":"2022-07-09T15:45:32.056841Z","iopub.status.idle":"2022-07-09T15:45:32.083496Z","shell.execute_reply.started":"2022-07-09T15:45:32.056807Z","shell.execute_reply":"2022-07-09T15:45:32.082491Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_callback\"></a>\n<font size = 5><span style=\"color:#db9833\">6.4.4 Callback: </span></font>","metadata":{}},{"cell_type":"code","source":"if  os.path.isdir(\"./weights/roberta_base_weights\") is None:\n          os.makedirs(\"./weights/roberta_base_weights\")\ncheckpoint_filepath_roberta_base  = \"./weights/roberta_base_weights\"\ncheckpoint_callback_roberta_base = tf.keras.callbacks.ModelCheckpoint(\n    checkpoint_filepath_roberta_base,\n    save_weights_only=True,\n    monitor='val_accuracy',\n    mode='auto',\n    save_best_only=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:45:32.084909Z","iopub.execute_input":"2022-07-09T15:45:32.085511Z","iopub.status.idle":"2022-07-09T15:45:32.095225Z","shell.execute_reply.started":"2022-07-09T15:45:32.085466Z","shell.execute_reply":"2022-07-09T15:45:32.09418Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_compile\"></a>\n<font size = 5><span style=\"color:#db9833\">6.4.5 Compile: </span></font>","metadata":{}},{"cell_type":"code","source":"model_4.compile(loss = tf.keras.losses.BinaryCrossentropy(from_logits = True), \n             optimizer = tf.keras.optimizers.Adam(lr = config.LEARNING_RATE , epsilon = 1e-8 , decay  =config.WEIGTH_DECAY , clipnorm = 1.0),\n             metrics = [\"accuracy\"])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:45:32.097398Z","iopub.execute_input":"2022-07-09T15:45:32.097791Z","iopub.status.idle":"2022-07-09T15:45:32.119088Z","shell.execute_reply.started":"2022-07-09T15:45:32.097754Z","shell.execute_reply":"2022-07-09T15:45:32.118174Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_training\"></a>\n<font size = 5><span style=\"color:#db9833\">6.4.6 Training: </span></font>","metadata":{}},{"cell_type":"code","source":"roberta_base_history  = model_4.fit(x = {\"input_ids\": x_train[\"input_ids\"] , \"attention_mask\" : x_train[\"attention_mask\"]},\n                y = train_df[\"target\"] , \n                epochs = config.EPOCHS , \n                validation_split = 0.2,\n                batch_size = 128 , callbacks = [checkpoint_callback_roberta_base])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:45:32.120726Z","iopub.execute_input":"2022-07-09T15:45:32.121797Z","iopub.status.idle":"2022-07-09T15:50:27.022137Z","shell.execute_reply.started":"2022-07-09T15:45:32.12176Z","shell.execute_reply":"2022-07-09T15:50:27.021214Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_4.load_weights(checkpoint_filepath_roberta_base)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:50:27.023947Z","iopub.execute_input":"2022-07-09T15:50:27.02439Z","iopub.status.idle":"2022-07-09T15:50:28.545866Z","shell.execute_reply.started":"2022-07-09T15:50:27.024353Z","shell.execute_reply":"2022-07-09T15:50:28.544985Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"roberta_base_hist_df = pd.DataFrame(roberta_base_history.history , columns = ['loss', 'accuracy', 'val_loss', 'val_accuracy'])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:50:28.547145Z","iopub.execute_input":"2022-07-09T15:50:28.547509Z","iopub.status.idle":"2022-07-09T15:50:28.554689Z","shell.execute_reply.started":"2022-07-09T15:50:28.547474Z","shell.execute_reply":"2022-07-09T15:50:28.553685Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_graph\"></a>\n<font size = 5><span style=\"color:#db9833\">6.4.7 Graphs: </span></font>","metadata":{}},{"cell_type":"code","source":"fig = px.line(roberta_base_hist_df, y=[\"accuracy\" , \"val_accuracy\"], title=\"Accuracy\") \nfig.update_xaxes(title=\"Epochs\")\nfig.update_yaxes(title = \"Accuracy\")\nfig.update_layout(showlegend = True,\n        title = {\n            'text': \"Roberta Base Accuracy\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'})\nfig.show()\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:50:28.556096Z","iopub.execute_input":"2022-07-09T15:50:28.557791Z","iopub.status.idle":"2022-07-09T15:50:28.630182Z","shell.execute_reply.started":"2022-07-09T15:50:28.55773Z","shell.execute_reply":"2022-07-09T15:50:28.629305Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.line(roberta_base_hist_df, y=[\"loss\" , \"val_loss\"], title=\"Loss\") \nfig.update_xaxes(title=\"Epochs\")\nfig.update_yaxes(title = \"Loss\")\nfig.update_layout(showlegend = True,\n        title = {\n            'text': \"Bert Base uncased Loss\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'})\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:50:28.63173Z","iopub.execute_input":"2022-07-09T15:50:28.632359Z","iopub.status.idle":"2022-07-09T15:50:28.703438Z","shell.execute_reply.started":"2022-07-09T15:50:28.632324Z","shell.execute_reply":"2022-07-09T15:50:28.702448Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_confusion_matrix\"></a>\n<font size = 5><span style=\"color:#db9833\">6.4.8 Confusion Matrix: </span></font>","metadata":{}},{"cell_type":"code","source":"y_pred = model_4.predict({\"input_ids\" : x_train[\"input_ids\"] ,\"attention_mask\" : x_train[\"attention_mask\"]})\ny_pred = np.where(y_pred > 0.5 , 1,0)\ny_test = train_df[\"target\"]\nCLASS_LABELS = [\"Disaster Tweet\" , \"Non Disaster Tweet\"]\ncm_data = confusion_matrix(y_test , y_pred)\ncm = pd.DataFrame(cm_data , columns = CLASS_LABELS , index = CLASS_LABELS)\nfig = px.imshow(img = cm_data ,\n                x = CLASS_LABELS, \n                y = CLASS_LABELS,\n                aspect=\"auto\" , \n                color_continuous_scale = \"mint\")\nfig.update_xaxes(title=\"Predicted\")\nfig.update_yaxes(title = \"Actual\")\nfig.update_layout(title = \"Confusion Matrix\",\n                  template = \"plotly_white\",\n                  title_x = 0.5)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:50:28.704768Z","iopub.execute_input":"2022-07-09T15:50:28.705177Z","iopub.status.idle":"2022-07-09T15:50:45.00273Z","shell.execute_reply.started":"2022-07-09T15:50:28.705142Z","shell.execute_reply":"2022-07-09T15:50:45.001797Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(classification_report(y_test , y_pred))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:50:45.004128Z","iopub.execute_input":"2022-07-09T15:50:45.005715Z","iopub.status.idle":"2022-07-09T15:50:45.027411Z","shell.execute_reply.started":"2022-07-09T15:50:45.005671Z","shell.execute_reply":"2022-07-09T15:50:45.026467Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_predictions\"></a>\n<font size = 5><span style=\"color:#db9833\">6.4.9 Predictions: </span></font>","metadata":{}},{"cell_type":"code","source":"model_4_pred_probs = model_4.predict({\"input_ids\" : x_test[\"input_ids\"] ,\"attention_mask\" : x_test[\"attention_mask\"]})\ny_pred_4 = np.where(model_4_pred_probs > 0.5 , 1,0)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:50:45.028955Z","iopub.execute_input":"2022-07-09T15:50:45.029329Z","iopub.status.idle":"2022-07-09T15:50:50.896119Z","shell.execute_reply.started":"2022-07-09T15:50:45.029295Z","shell.execute_reply":"2022-07-09T15:50:50.89524Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"roberta_base_df=pd.DataFrame()\nroberta_base_df['id'] = test_ids\nroberta_base_df['target'] = y_pred_4","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:50:50.897445Z","iopub.execute_input":"2022-07-09T15:50:50.898135Z","iopub.status.idle":"2022-07-09T15:50:50.906443Z","shell.execute_reply.started":"2022-07-09T15:50:50.898097Z","shell.execute_reply":"2022-07-09T15:50:50.905461Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"roberta_base_df.to_csv('roberta_base.csv',index = False)\nroberta_base_df[\"target\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:50:50.907995Z","iopub.execute_input":"2022-07-09T15:50:50.909012Z","iopub.status.idle":"2022-07-09T15:50:50.927056Z","shell.execute_reply.started":"2022-07-09T15:50:50.908985Z","shell.execute_reply":"2022-07-09T15:50:50.926228Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_large\"></a>\n**<center><font size = 5><span style=\"color:#db9833\">6.5 Roberta large </span></font></center>**\n* <font size = 3><span style=\"color:#2F4F4F\">The model weights can be found at [Roberta Large](https://huggingface.co/roberta-large) </span></font>\n* <font size = 3><span style=\"color:#2F4F4F\"> There are a total of 355M parameters. </span></font>","metadata":{}},{"cell_type":"markdown","source":"<a id = \"roberta_large_tokenization\"></a>\n<font size = 5><span style=\"color:#db9833\">6.5.1 Tokenization: </span></font>","metadata":{}},{"cell_type":"code","source":"K.clear_session()\nMODEL_5 = \"roberta-large\"\ntokenizer = AutoTokenizer.from_pretrained(MODEL_5 , do_lower_case = config.LOWER_CASE , max_length = config.MAX_LEN )\nx_train = tokenizer(\n        text = train_df[\"text\"].tolist(),\n        add_special_tokens = True,\n        max_length = config.MAX_LEN,\n        truncation = True,\n        padding = True,\n        return_tensors = \"tf\",\n        return_token_type_ids = False,\n        return_attention_mask = True,\n        verbose = True\n        )\n\nx_test = tokenizer(\n        text = test_df[\"text\"].tolist(),\n        add_special_tokens = True,\n        max_length = config.MAX_LEN,\n        truncation = True,\n        padding = True,\n        return_tensors = \"tf\",\n        return_token_type_ids = False,\n        return_attention_mask = True,\n        verbose = True\n        )\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:50:50.940544Z","iopub.execute_input":"2022-07-09T15:50:50.940795Z","iopub.status.idle":"2022-07-09T15:50:57.796754Z","shell.execute_reply.started":"2022-07-09T15:50:50.940771Z","shell.execute_reply":"2022-07-09T15:50:57.795711Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_model_loading\"></a>\n<font size = 5><span style=\"color:#db9833\">6.5.2 Loading the pretrained transformer model : </span></font>","metadata":{}},{"cell_type":"code","source":"roberta_large = TFAutoModel.from_pretrained(MODEL_5)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:50:57.798429Z","iopub.execute_input":"2022-07-09T15:50:57.799206Z","iopub.status.idle":"2022-07-09T15:51:38.983381Z","shell.execute_reply.started":"2022-07-09T15:50:57.799163Z","shell.execute_reply":"2022-07-09T15:51:38.982457Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_model\"></a>\n<font size = 5><span style=\"color:#db9833\">6.5.3 Model Building: </span></font>","metadata":{}},{"cell_type":"code","source":"input_ids = tf.keras.layers.Input(shape = (config.MAX_LEN,) , dtype = tf.int32 , name = \"input_ids\")\ninput_mask = tf.keras.layers.Input(shape = (config.MAX_LEN,) , dtype = tf.int32 , name = \"attention_mask\")\nembeddings = roberta_large(input_ids , attention_mask = input_mask)[1]\n# x = tf.keras.layers.GlobalAveragePooling1D()(embeddings)\nx = tf.keras.layers.Dropout(0.3)(embeddings)\nx = tf.keras.layers.Dense(128 , activation = \"relu\")(x)\nx = tf.keras.layers.Dropout(0.2)(x)\nx = tf.keras.layers.Dense(32 , activation = \"relu\")(x)\noutput = tf.keras.layers.Dense(config.NUM_LABELS , activation = \"sigmoid\")(x)\n\nmodel_5 = tf.keras.Model(inputs = [input_ids , input_mask] , outputs = output)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:51:38.984893Z","iopub.execute_input":"2022-07-09T15:51:38.985899Z","iopub.status.idle":"2022-07-09T15:51:41.866064Z","shell.execute_reply.started":"2022-07-09T15:51:38.985861Z","shell.execute_reply":"2022-07-09T15:51:41.865123Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Transformer Layer freezed!!\")\nmodel_5.layers[2].trainable = True\nmodel_5.summary()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:51:41.867852Z","iopub.execute_input":"2022-07-09T15:51:41.868212Z","iopub.status.idle":"2022-07-09T15:51:41.913234Z","shell.execute_reply.started":"2022-07-09T15:51:41.868176Z","shell.execute_reply":"2022-07-09T15:51:41.912279Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_large_callback\"></a>\n<font size = 5><span style=\"color:#db9833\">6.5.4 Callback: </span></font>","metadata":{}},{"cell_type":"code","source":"if  os.path.isdir(\"./weights/roberta_large_weights\") is None:\n          os.makedirs(\"./weights/roberta_large_weights\")\ncheckpoint_filepath_roberta_large  = \"./weights/roberta_large_weights\"\ncheckpoint_callback_roberta_large = tf.keras.callbacks.ModelCheckpoint(\n    checkpoint_filepath_roberta_large,\n    save_weights_only=True,\n    monitor='val_accuracy',\n    mode='auto',\n    save_best_only=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:51:41.915526Z","iopub.execute_input":"2022-07-09T15:51:41.916201Z","iopub.status.idle":"2022-07-09T15:51:41.92295Z","shell.execute_reply.started":"2022-07-09T15:51:41.916163Z","shell.execute_reply":"2022-07-09T15:51:41.922036Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_large_compile\"></a>\n<font size = 5><span style=\"color:#db9833\">6.5.5 Compile: </span></font>","metadata":{}},{"cell_type":"code","source":"model_5.compile(loss = tf.keras.losses.BinaryCrossentropy(from_logits = True), \n             optimizer = tf.keras.optimizers.Adam(lr = config.LEARNING_RATE , epsilon = 1e-8 , decay  =config.WEIGTH_DECAY , clipnorm = 1.0),\n             metrics = [\"accuracy\"])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:51:41.924761Z","iopub.execute_input":"2022-07-09T15:51:41.925513Z","iopub.status.idle":"2022-07-09T15:51:41.947753Z","shell.execute_reply.started":"2022-07-09T15:51:41.925477Z","shell.execute_reply":"2022-07-09T15:51:41.946709Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_large_training\"></a>\n<font size = 5><span style=\"color:#db9833\">6.5.6 Training: </span></font>","metadata":{}},{"cell_type":"code","source":"roberta_large_history  = model_5.fit(x = {\"input_ids\": x_train[\"input_ids\"] , \"attention_mask\" : x_train[\"attention_mask\"]},\n                y = train_df[\"target\"] , \n                epochs = config.EPOCHS , \n                validation_split = 0.2,\n                batch_size = 16 , callbacks = [checkpoint_callback_roberta_large])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:51:41.949518Z","iopub.execute_input":"2022-07-09T15:51:41.949861Z","iopub.status.idle":"2022-07-09T16:14:41.468356Z","shell.execute_reply.started":"2022-07-09T15:51:41.949805Z","shell.execute_reply":"2022-07-09T16:14:41.467479Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_5.load_weights(checkpoint_filepath_roberta_large)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:14:41.47025Z","iopub.execute_input":"2022-07-09T16:14:41.470685Z","iopub.status.idle":"2022-07-09T16:14:58.24286Z","shell.execute_reply.started":"2022-07-09T16:14:41.470658Z","shell.execute_reply":"2022-07-09T16:14:58.241941Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"roberta_large_hist_df = pd.DataFrame(roberta_large_history.history , columns = ['loss', 'accuracy', 'val_loss', 'val_accuracy'])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:14:58.244339Z","iopub.execute_input":"2022-07-09T16:14:58.244717Z","iopub.status.idle":"2022-07-09T16:14:58.254828Z","shell.execute_reply.started":"2022-07-09T16:14:58.244681Z","shell.execute_reply":"2022-07-09T16:14:58.253791Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_large_graph\"></a>\n<font size = 5><span style=\"color:#db9833\">6.5.7 Graphs: </span></font>","metadata":{}},{"cell_type":"code","source":"fig = px.line(roberta_large_hist_df, y=[\"accuracy\" , \"val_accuracy\"], title=\"Accuracy\") \nfig.update_xaxes(title=\"Epochs\")\nfig.update_yaxes(title = \"Accuracy\")\nfig.update_layout(showlegend = True,\n        title = {\n            'text': \"Roberta Large Accuracy\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'})\nfig.show()\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:14:58.256765Z","iopub.execute_input":"2022-07-09T16:14:58.257604Z","iopub.status.idle":"2022-07-09T16:14:58.34218Z","shell.execute_reply.started":"2022-07-09T16:14:58.257563Z","shell.execute_reply":"2022-07-09T16:14:58.341286Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.line(roberta_large_hist_df, y=[\"loss\" , \"val_loss\"], title=\"Loss\") \nfig.update_xaxes(title=\"Epochs\")\nfig.update_yaxes(title = \"Loss\")\nfig.update_layout(showlegend = True,\n        title = {\n            'text': \"Roberta Large Loss\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'})\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:14:58.343387Z","iopub.execute_input":"2022-07-09T16:14:58.34412Z","iopub.status.idle":"2022-07-09T16:14:58.416745Z","shell.execute_reply.started":"2022-07-09T16:14:58.344083Z","shell.execute_reply":"2022-07-09T16:14:58.41582Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_large_confusion_matrix\"></a>\n<font size = 5><span style=\"color:#db9833\">6.5.8 Confusion Matrix: </span></font>","metadata":{}},{"cell_type":"code","source":"y_pred = model_5.predict({\"input_ids\" : x_train[\"input_ids\"] ,\"attention_mask\" : x_train[\"attention_mask\"]})\ny_pred = np.where(y_pred > 0.5 , 1,0)\ny_test = train_df[\"target\"]\nCLASS_LABELS = [\"Disaster Tweet\" , \"Non Disaster Tweet\"]\ncm_data = confusion_matrix(y_test , y_pred)\ncm = pd.DataFrame(cm_data , columns = CLASS_LABELS , index = CLASS_LABELS)\nfig = px.imshow(img = cm_data ,\n                x = CLASS_LABELS, \n                y = CLASS_LABELS,\n                aspect=\"auto\" , \n                color_continuous_scale = \"mint\")\nfig.update_xaxes(title=\"Predicted\")\nfig.update_yaxes(title = \"Actual\")\nfig.update_layout(title = \"Confusion Matrix\",\n                  template = \"plotly_white\",\n                  title_x = 0.5)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:14:58.417944Z","iopub.execute_input":"2022-07-09T16:14:58.418666Z","iopub.status.idle":"2022-07-09T16:15:45.044935Z","shell.execute_reply.started":"2022-07-09T16:14:58.41863Z","shell.execute_reply":"2022-07-09T16:15:45.044009Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(classification_report(y_test , y_pred))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:15:45.046148Z","iopub.execute_input":"2022-07-09T16:15:45.048713Z","iopub.status.idle":"2022-07-09T16:15:45.070073Z","shell.execute_reply.started":"2022-07-09T16:15:45.048672Z","shell.execute_reply":"2022-07-09T16:15:45.069118Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"roberta_large_predictions\"></a>\n<font size = 5><span style=\"color:#db9833\">6.5.9 Predictions: </span></font>","metadata":{}},{"cell_type":"code","source":"model_5_pred_probs = model_5.predict({\"input_ids\" : x_test[\"input_ids\"] ,\"attention_mask\" : x_test[\"attention_mask\"]})\ny_pred_5 = np.where(model_5_pred_probs > 0.5 , 1,0)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:15:45.071567Z","iopub.execute_input":"2022-07-09T16:15:45.071889Z","iopub.status.idle":"2022-07-09T16:16:02.981583Z","shell.execute_reply.started":"2022-07-09T16:15:45.071857Z","shell.execute_reply":"2022-07-09T16:16:02.98068Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"roberta_large_df=pd.DataFrame()\nroberta_large_df['id'] = test_ids\nroberta_large_df['target'] = y_pred_5","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:16:02.983105Z","iopub.execute_input":"2022-07-09T16:16:02.983472Z","iopub.status.idle":"2022-07-09T16:16:02.991761Z","shell.execute_reply.started":"2022-07-09T16:16:02.983414Z","shell.execute_reply":"2022-07-09T16:16:02.990733Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"roberta_large_df.to_csv('roberta_large.csv',index = False)\nroberta_large_df[\"target\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:16:02.993377Z","iopub.execute_input":"2022-07-09T16:16:02.994011Z","iopub.status.idle":"2022-07-09T16:16:03.011969Z","shell.execute_reply.started":"2022-07-09T16:16:02.993976Z","shell.execute_reply":"2022-07-09T16:16:03.010926Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"xlm_roberta\"></a>\n<center><font size = 5><span style=\"color:#db9833\">6.5 XLM Roberta </span></font></center>\n\n* <font size = 3><span style=\"color:#2F4F4F\">The model weights can be found at [XLM Roberta base](https://huggingface.co/xlm-roberta-base) </span></font>\n* <font size = 3><span style=\"color:#2F4F4F\"> There are a total of 125M parameters. </span></font>\n\nNote : To train this model just uncomment the cells and comment the cells of any one of the transformer models","metadata":{}},{"cell_type":"code","source":"# K.clear_session()\n# MODEL_6 = \"cardiffnlp/twitter-xlm-roberta-base-sentiment\"\n# tokenizer = AutoTokenizer.from_pretrained(MODEL_6 , do_lower_case = config.LOWER_CASE , max_length = config.MAX_LEN )\n# x_train = tokenizer(\n#         text = train_df[\"text\"].tolist(),\n#         add_special_tokens = True,\n#         max_length = config.MAX_LEN,\n#         truncation = True,\n#         padding = True,\n#         return_tensors = \"tf\",\n#         return_token_type_ids = False,\n#         return_attention_mask = True,\n#         verbose = True\n#         )\n\n# x_test = tokenizer(\n#         text = test_df[\"text\"].tolist(),\n#         add_special_tokens = True,\n#         max_length = config.MAX_LEN,\n#         truncation = True,\n#         padding = True,\n#         return_tensors = \"tf\",\n#         return_token_type_ids = False,\n#         return_attention_mask = True,\n#         verbose = True\n#         )\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:16:03.013285Z","iopub.execute_input":"2022-07-09T16:16:03.014109Z","iopub.status.idle":"2022-07-09T16:16:03.019341Z","shell.execute_reply.started":"2022-07-09T16:16:03.014073Z","shell.execute_reply":"2022-07-09T16:16:03.018213Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# xlm_roberta = TFAutoModel.from_pretrained(MODEL_6)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:16:03.020614Z","iopub.execute_input":"2022-07-09T16:16:03.021718Z","iopub.status.idle":"2022-07-09T16:16:03.030359Z","shell.execute_reply.started":"2022-07-09T16:16:03.021682Z","shell.execute_reply":"2022-07-09T16:16:03.029394Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# input_ids = tf.keras.layers.Input(shape = (config.MAX_LEN,) , dtype = tf.int32 , name = \"input_ids\")\n# input_mask = tf.keras.layers.Input(shape = (config.MAX_LEN,) , dtype = tf.int32 , name = \"attention_mask\")\n# embeddings = xlm_roberta(input_ids , attention_mask = input_mask)[1]\n# x = tf.keras.layers.GlobalAveragePooling1D()(embeddings)\n# x = tf.keras.layers.Dropout(0.3)(embeddings)\n# x = tf.keras.layers.Dense(128 , activation = \"relu\")(x)\n# x = tf.keras.layers.Dropout(0.2)(x)\n# x = tf.keras.layers.Dense(32 , activation = \"relu\")(x)\n# output = tf.keras.layers.Dense(config.NUM_LABELS , activation = \"sigmoid\")(x)\n\n# model_6 = tf.keras.Model(inputs = [input_ids , input_mask] , outputs = output)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:16:03.032785Z","iopub.execute_input":"2022-07-09T16:16:03.034646Z","iopub.status.idle":"2022-07-09T16:16:03.041919Z","shell.execute_reply.started":"2022-07-09T16:16:03.03461Z","shell.execute_reply":"2022-07-09T16:16:03.040935Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print(\"Transformer Layer freezed!!\")\n# model_6.layers[2].trainable = True\n# model_6.summary()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:16:03.045194Z","iopub.execute_input":"2022-07-09T16:16:03.045602Z","iopub.status.idle":"2022-07-09T16:16:03.052052Z","shell.execute_reply.started":"2022-07-09T16:16:03.045574Z","shell.execute_reply":"2022-07-09T16:16:03.051052Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# if  os.path.isdir(\"./weights/xlm_roberta_weights\") is None:\n#           os.makedirs(\"./weights/xlm_roberta_weights\")\n# checkpoint_filepath_xlm_roberta  = \"./weights/xlm_roberta_weights\"\n# checkpoint_callback_xlm_roberta = tf.keras.callbacks.ModelCheckpoint(\n#     checkpoint_filepath_xlm_roberta,\n#     save_weights_only=True,\n#     monitor='val_accuracy',\n#     mode='auto',\n#     save_best_only=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:16:03.053155Z","iopub.execute_input":"2022-07-09T16:16:03.053857Z","iopub.status.idle":"2022-07-09T16:16:03.062482Z","shell.execute_reply.started":"2022-07-09T16:16:03.053821Z","shell.execute_reply":"2022-07-09T16:16:03.06155Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# model_6.compile(loss = tf.keras.losses.BinaryCrossentropy(from_logits = True), \n#              optimizer = tf.keras.optimizers.Adam(lr = config.LEARNING_RATE , epsilon = 1e-8 , decay  =config.WEIGTH_DECAY , clipnorm = 1.0),\n#              metrics = [\"accuracy\"])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:16:03.064494Z","iopub.execute_input":"2022-07-09T16:16:03.065135Z","iopub.status.idle":"2022-07-09T16:16:03.075959Z","shell.execute_reply.started":"2022-07-09T16:16:03.065098Z","shell.execute_reply":"2022-07-09T16:16:03.074986Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# xlm_roberta_history  = model_6.fit(x = {\"input_ids\": x_train[\"input_ids\"] , \"attention_mask\" : x_train[\"attention_mask\"]},\n#                 y = train_df[\"target\"] , \n#                 epochs = config.EPOCHS , \n#                 validation_split = 0.2,\n#                 batch_size = 1 , callbacks=[checkpoint_callback_xlm_roberta])\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:16:03.078315Z","iopub.execute_input":"2022-07-09T16:16:03.079007Z","iopub.status.idle":"2022-07-09T16:16:03.086224Z","shell.execute_reply.started":"2022-07-09T16:16:03.078973Z","shell.execute_reply":"2022-07-09T16:16:03.085289Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# model_6.load_weights(checkpoint_filepath_xlm_roberta)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:16:03.088318Z","iopub.execute_input":"2022-07-09T16:16:03.088986Z","iopub.status.idle":"2022-07-09T16:16:03.095821Z","shell.execute_reply.started":"2022-07-09T16:16:03.088952Z","shell.execute_reply":"2022-07-09T16:16:03.094906Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# model_6_pred_probs = model_6.predict({\"input_ids\" : x_test[\"input_ids\"] ,\"attention_mask\" : x_test[\"attention_mask\"]})\n# y_pred_6 = np.where(model_6_pred_probs > 0.5 , 1,0)\n# y_pred_6","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:16:03.097245Z","iopub.execute_input":"2022-07-09T16:16:03.098086Z","iopub.status.idle":"2022-07-09T16:16:03.105965Z","shell.execute_reply.started":"2022-07-09T16:16:03.098049Z","shell.execute_reply":"2022-07-09T16:16:03.104938Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# xlm_roberta_df=pd.DataFrame()\n# xlm_roberta_df['id'] = test_ids\n# xlm_roberta_df['target'] = y_pred_6","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:16:03.10807Z","iopub.execute_input":"2022-07-09T16:16:03.109079Z","iopub.status.idle":"2022-07-09T16:16:03.11668Z","shell.execute_reply.started":"2022-07-09T16:16:03.109045Z","shell.execute_reply":"2022-07-09T16:16:03.11556Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# xlm_roberta_df.to_csv('xlm_roberta.csv',index = False)\n# xlm_roberta_df[\"target\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:16:03.11967Z","iopub.execute_input":"2022-07-09T16:16:03.120781Z","iopub.status.idle":"2022-07-09T16:16:03.126825Z","shell.execute_reply.started":"2022-07-09T16:16:03.120744Z","shell.execute_reply":"2022-07-09T16:16:03.125894Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id = \"Comaprision\"></a>\n**<center><font size = 6><span style=\"color:#2a6592\">7. Comparision </span></font></center>**","metadata":{}},{"cell_type":"code","source":"fig = go.Figure(data=[\n    go.Bar(x=bert_based_uncased_hist_df.index, y=bert_based_uncased_hist_df[\"val_accuracy\"] , name = \"Bert Base Uncased\"),\n    go.Bar(x=bert_large_hist_df.index, y=bert_large_hist_df[\"val_accuracy\"] , name = \"Bert Large\"),\n    go.Bar(x=distill_bert_uncased_hist_df.index, y=distill_bert_uncased_hist_df[\"val_accuracy\"] , name = \"Distill Bert Uncased\"),\n    go.Bar(x=roberta_base_hist_df.index, y=roberta_base_hist_df[\"val_accuracy\"] , name = \"Roberta\"),\n    go.Bar(x=roberta_large_hist_df.index, y=roberta_large_hist_df[\"val_accuracy\"] , name = \"Roberta Large\")\n    \n])\n\nfig.update_xaxes(title=\"Epochs\")\nfig.update_yaxes(title = \"Validation Accuracy\")\n\nfig.update_layout(showlegend = True,\n        title = {\n            'text': \"Validation Accuracy Per Epoch\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'} \n        ,xaxis = dict(\n        tickmode = 'linear',\n        tick0 = 1,\n    \n    ))\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:46:44.706166Z","iopub.execute_input":"2022-07-09T16:46:44.706636Z","iopub.status.idle":"2022-07-09T16:46:44.744176Z","shell.execute_reply.started":"2022-07-09T16:46:44.706595Z","shell.execute_reply":"2022-07-09T16:46:44.743302Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = go.Figure(data=[\n    go.Bar(x=bert_based_uncased_hist_df.index, y=bert_based_uncased_hist_df[\"val_loss\"] , name = \"Bert Base Uncased\"),\n    go.Bar(x=bert_large_hist_df.index, y=bert_large_hist_df[\"val_loss\"] , name = \"Bert Large\"),\n    go.Bar(x=distill_bert_uncased_hist_df.index, y=distill_bert_uncased_hist_df[\"val_loss\"] , name = \"Distill Bert Uncased\"),\n    go.Bar(x=roberta_base_hist_df.index, y=roberta_base_hist_df[\"val_loss\"] , name = \"Roberta\"),\n    go.Bar(x=roberta_large_hist_df.index, y=roberta_large_hist_df[\"val_loss\"] , name = \"Roberta Large\")\n    \n])\n\nfig.update_xaxes(title=\"Epochs\")\nfig.update_yaxes(title = \"Validation Loss\")\n\nfig.update_layout(showlegend = True,\n        title = {\n            'text': \"Validation Loss Per Epoch\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'} \n        ,xaxis = dict(\n        tickmode = 'linear',\n        tick0 = 1,\n    \n    ))\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T16:48:31.982971Z","iopub.execute_input":"2022-07-09T16:48:31.983318Z","iopub.status.idle":"2022-07-09T16:48:32.009347Z","shell.execute_reply.started":"2022-07-09T16:48:31.98329Z","shell.execute_reply":"2022-07-09T16:48:32.00815Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = [bert_based_uncased_hist_df[\"val_accuracy\"].max() , bert_large_hist_df[\"val_accuracy\"].max() , distill_bert_uncased_hist_df[\"val_accuracy\"].max() , roberta_base_hist_df[\"val_accuracy\"].max() , roberta_large_hist_df[\"val_accuracy\"].max()]\nx = [\"Bert Base Uncased\" , \"Bert Large\" , \"Distill Bert Uncase\" , \"Roberta Base\" , \"Roberta Large\" ]\n\nfig = px.bar(x=x, y=y)\nfig.update_xaxes(title=\"Epochs\")\nfig.update_yaxes(title = \"Validation Accuracy\")\nfig.update_layout(showlegend = True,\n        title = {\n            'text': \"Validation Accuracy Comparision\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'} \n        )\n\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-09T18:10:02.358999Z","iopub.execute_input":"2022-07-09T18:10:02.359693Z","iopub.status.idle":"2022-07-09T18:10:02.562296Z","shell.execute_reply.started":"2022-07-09T18:10:02.359649Z","shell.execute_reply":"2022-07-09T18:10:02.561117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = [bert_based_uncased_hist_df[\"val_loss\"].max() , bert_large_hist_df[\"val_loss\"].max() , distill_bert_uncased_hist_df[\"val_loss\"].max() , roberta_base_hist_df[\"val_loss\"].max() , roberta_large_hist_df[\"val_loss\"].max()]\nx = [\"Bert Base Uncased\" , \"Bert Large\" , \"Distill Bert Uncase\" , \"Roberta Base\" , \"Roberta Large\" ]\n\nfig = px.bar(x=x, y=y)\nfig.update_xaxes(title=\"Epochs\")\nfig.update_yaxes(title = \"Validation Loss\")\nfig.update_layout(showlegend = True,\n        title = {\n            'text': \"Validation Loss Comparision\",\n            'y':0.95,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'} \n        )\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T17:01:27.683217Z","iopub.execute_input":"2022-07-09T17:01:27.683586Z","iopub.status.idle":"2022-07-09T17:01:27.749247Z","shell.execute_reply.started":"2022-07-09T17:01:27.683556Z","shell.execute_reply":"2022-07-09T17:01:27.748351Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><b><font size = 3><span style=\"color:#2F4F4F\"> Thank You for reading 😊</span></font></b></center>\n\n<center><b><font size = 3><span style=\"color:#2F4F4F\"> If you have any suggestions or feeback, please let me know</span></font></b></center>","metadata":{}}]}