{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px\"> Objective </h2>\n\nIf we have two sentences, there are three ways they could be related: one could entail the other, one could contradict the other, or they could be unrelated. Natural Language Inferencing (NLI) is a popular NLP problem that involves determining how pairs of sentences (consisting of a premise and a hypothesis) are related.\n\nOur task is to create an NLI model that assigns labels of 0, 1, or 2 (corresponding to entailment, neutral, and contradiction) to pairs of premises and hypotheses. ","metadata":{}},{"cell_type":"markdown","source":"\n<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px\"> Load data and libraries </h2>\n","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2022-07-14T06:15:37.430226Z","iopub.execute_input":"2022-07-14T06:15:37.430544Z","iopub.status.idle":"2022-07-14T06:15:37.460365Z","shell.execute_reply.started":"2022-07-14T06:15:37.430466Z","shell.execute_reply":"2022-07-14T06:15:37.459591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from transformers import BertTokenizer, TFBertModel\nimport matplotlib.pyplot as plt\nimport tensorflow as tf","metadata":{"execution":{"iopub.status.busy":"2022-07-14T06:15:37.461896Z","iopub.execute_input":"2022-07-14T06:15:37.462351Z","iopub.status.idle":"2022-07-14T06:15:44.258174Z","shell.execute_reply.started":"2022-07-14T06:15:37.462315Z","shell.execute_reply":"2022-07-14T06:15:44.257377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/contradictory-my-dear-watson/train.csv')\ntest = pd.read_csv('/kaggle/input/contradictory-my-dear-watson/test.csv')\nsubmission = pd.read_csv('/kaggle/input/contradictory-my-dear-watson/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-14T06:15:44.259222Z","iopub.execute_input":"2022-07-14T06:15:44.259431Z","iopub.status.idle":"2022-07-14T06:15:44.442365Z","shell.execute_reply.started":"2022-07-14T06:15:44.259407Z","shell.execute_reply":"2022-07-14T06:15:44.441721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T06:15:44.443847Z","iopub.execute_input":"2022-07-14T06:15:44.444252Z","iopub.status.idle":"2022-07-14T06:15:44.466322Z","shell.execute_reply.started":"2022-07-14T06:15:44.444209Z","shell.execute_reply":"2022-07-14T06:15:44.465120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px\"> Data Exploration </h2>\n","metadata":{}},{"cell_type":"code","source":"train.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:45:13.847869Z","iopub.execute_input":"2022-07-14T02:45:13.848921Z","iopub.status.idle":"2022-07-14T02:45:13.854860Z","shell.execute_reply.started":"2022-07-14T02:45:13.848841Z","shell.execute_reply":"2022-07-14T02:45:13.854019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Majority of the training set in English, and the rest of the languages are equally distributed","metadata":{}},{"cell_type":"code","source":"import plotly.graph_objects as go\nfrom plotly.offline import init_notebook_mode, iplot\n\ncol = \"language\"\ngrouped = train[col].value_counts().reset_index()\ngrouped = grouped.rename(columns = {col : \"count\", \"index\" : col})\n\n## plot\ntrace = go.Pie(labels=grouped[col], values=grouped['count'], pull=[0.05, 0], marker=dict(colors=[\"#6ad49b\", \"#a678de\"]))\nlayout = go.Layout(title=\"\", height=500, legend=dict(x=0.1, y=1.1))\nfig = go.Figure(data = [trace], layout = layout)\niplot(fig)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-14T02:45:13.856330Z","iopub.execute_input":"2022-07-14T02:45:13.857416Z","iopub.status.idle":"2022-07-14T02:45:15.072488Z","shell.execute_reply.started":"2022-07-14T02:45:13.857354Z","shell.execute_reply":"2022-07-14T02:45:15.071799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px\"> Modelling using TPU </h2>\n","metadata":{}},{"cell_type":"code","source":"os.environ[\"WANDB_API_KEY\"] = \"0\" ## to silence warning","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:45:15.073933Z","iopub.execute_input":"2022-07-14T02:45:15.074487Z","iopub.status.idle":"2022-07-14T02:45:15.079034Z","shell.execute_reply.started":"2022-07-14T02:45:15.074448Z","shell.execute_reply":"2022-07-14T02:45:15.078113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"try:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nexcept ValueError:\n    strategy = tf.distribute.get_strategy() # for CPU and single GPU\n\nprint('Number of replicas:', strategy.num_replicas_in_sync)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:45:15.080260Z","iopub.execute_input":"2022-07-14T02:45:15.080564Z","iopub.status.idle":"2022-07-14T02:45:21.279850Z","shell.execute_reply.started":"2022-07-14T02:45:15.080524Z","shell.execute_reply":"2022-07-14T02:45:21.278831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## BERT model from huggingface ","metadata":{}},{"cell_type":"markdown","source":"### Load Model","metadata":{}},{"cell_type":"code","source":"#https://huggingface.co/transformers/model_doc/bert.html#tfbertmodel\nmodel_name = 'bert-base-multilingual-cased'\ntokenizer = BertTokenizer.from_pretrained(model_name)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:45:21.281523Z","iopub.execute_input":"2022-07-14T02:45:21.281854Z","iopub.status.idle":"2022-07-14T02:45:23.645374Z","shell.execute_reply.started":"2022-07-14T02:45:21.281812Z","shell.execute_reply":"2022-07-14T02:45:23.644389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Tokenize","metadata":{}},{"cell_type":"code","source":"def encode_sentence(s):\n    tokens = list(tokenizer.tokenize(s))\n    tokens.append('[SEP]')\n    return tokenizer.convert_tokens_to_ids(tokens)\n\ns = \"Carpe Diem\"\nprint(list(tokenizer.tokenize(s)))\nprint(encode_sentence(s))","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:45:23.646609Z","iopub.execute_input":"2022-07-14T02:45:23.646847Z","iopub.status.idle":"2022-07-14T02:45:23.656406Z","shell.execute_reply.started":"2022-07-14T02:45:23.646820Z","shell.execute_reply":"2022-07-14T02:45:23.655675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"max_len = 220","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:45:23.658963Z","iopub.execute_input":"2022-07-14T02:45:23.659714Z","iopub.status.idle":"2022-07-14T02:45:23.672124Z","shell.execute_reply.started":"2022-07-14T02:45:23.659676Z","shell.execute_reply":"2022-07-14T02:45:23.670962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def bert_encode(hypotheses, premises, tokenizer):\n    num_examples = len(hypotheses)\n    sentence1 = tf.ragged.constant([\n      encode_sentence(s)\n      for s in np.array(hypotheses)])\n    sentence2 = tf.ragged.constant([\n      encode_sentence(s)\n       for s in np.array(premises)])\n    \n    #input word ids\n    cls = [tokenizer.convert_tokens_to_ids(['[CLS]'])]*sentence1.shape[0]\n    input_word_ids = tf.concat([cls, sentence1, sentence2], axis=-1)\n    \n    #input mask\n    input_mask = tf.ones_like(input_word_ids).to_tensor(shape=[num_examples, max_len])\n    \n    #input type ids\n    type_cls = tf.zeros_like(cls)\n    type_s1 = tf.zeros_like(sentence1)\n    type_s2 = tf.ones_like(sentence2)\n    input_type_ids = tf.concat(\n    [type_cls, type_s1, type_s2], axis=-1).to_tensor(shape=[num_examples, max_len])\n    \n    inputs = {\n      'input_word_ids': input_word_ids.to_tensor(shape=[num_examples, max_len]),\n      'input_mask': input_mask,\n      'input_type_ids': input_type_ids}\n\n    return inputs","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:45:23.673097Z","iopub.execute_input":"2022-07-14T02:45:23.673367Z","iopub.status.idle":"2022-07-14T02:45:23.686698Z","shell.execute_reply.started":"2022-07-14T02:45:23.673335Z","shell.execute_reply":"2022-07-14T02:45:23.685466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_input = bert_encode(train.hypothesis.values, train.premise.values, tokenizer)\ntest_input = bert_encode(test.hypothesis.values, test.premise.values, tokenizer)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:45:23.689233Z","iopub.execute_input":"2022-07-14T02:45:23.689780Z","iopub.status.idle":"2022-07-14T02:45:45.506852Z","shell.execute_reply.started":"2022-07-14T02:45:23.689732Z","shell.execute_reply":"2022-07-14T02:45:45.505894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Train using Keras Functional API \nhttps://www.tensorflow.org/guide/keras/functional","metadata":{}},{"cell_type":"code","source":"def build_model():\n    bert_encoder = TFBertModel.from_pretrained(model_name)\n    input_word_ids = tf.keras.Input(shape=(max_len,), dtype=tf.int32, name=\"input_word_ids\")\n    input_mask = tf.keras.Input(shape=(max_len,), dtype=tf.int32, name=\"input_mask\")\n    input_type_ids = tf.keras.Input(shape=(max_len,), dtype=tf.int32, name=\"input_type_ids\")\n    \n    embedding = bert_encoder([input_word_ids, input_mask, input_type_ids])[0]\n    output = tf.keras.layers.Dense(3, activation='softmax')(embedding[:,0,:])\n    \n    model = tf.keras.Model(inputs=[input_word_ids, input_mask, input_type_ids], outputs=output)\n    model.compile(tf.keras.optimizers.Adam(lr=1e-5), loss='sparse_categorical_crossentropy', metrics=['accuracy'])\n    \n    return model","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:45:45.554986Z","iopub.execute_input":"2022-07-14T02:45:45.555488Z","iopub.status.idle":"2022-07-14T02:45:45.564260Z","shell.execute_reply.started":"2022-07-14T02:45:45.555451Z","shell.execute_reply":"2022-07-14T02:45:45.563330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with strategy.scope():\n    model = build_model()\n    model.summary()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:45:45.565536Z","iopub.execute_input":"2022-07-14T02:45:45.566435Z","iopub.status.idle":"2022-07-14T02:46:47.314673Z","shell.execute_reply.started":"2022-07-14T02:45:45.566398Z","shell.execute_reply":"2022-07-14T02:46:47.313603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.fit(train_input, train.label.values, epochs = 2, verbose = 1, batch_size = 64, validation_split = 0.2)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:46:47.317446Z","iopub.execute_input":"2022-07-14T02:46:47.317846Z","iopub.status.idle":"2022-07-14T02:49:35.714629Z","shell.execute_reply.started":"2022-07-14T02:46:47.317801Z","shell.execute_reply":"2022-07-14T02:49:35.713226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Generate Predictions","metadata":{}},{"cell_type":"code","source":"predictions = [np.argmax(i) for i in model.predict(test_input)]\nsubmission = test.id.copy().to_frame()\nsubmission['prediction'] = predictions","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:49:35.716201Z","iopub.execute_input":"2022-07-14T02:49:35.716547Z","iopub.status.idle":"2022-07-14T02:49:57.820135Z","shell.execute_reply.started":"2022-07-14T02:49:35.716517Z","shell.execute_reply":"2022-07-14T02:49:57.819037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv(\"submission.csv\", index = False)\nsubmission","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:49:57.821455Z","iopub.execute_input":"2022-07-14T02:49:57.821715Z","iopub.status.idle":"2022-07-14T02:49:57.852761Z","shell.execute_reply.started":"2022-07-14T02:49:57.821684Z","shell.execute_reply":"2022-07-14T02:49:57.851221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<h2 id=\"adda\" style=\"color:black;background:#F4B400;padding:10px;border-radius:8px\"> Acknowledgements </h2>\n\n1. https://www.kaggle.com/code/anasofiauzsoy/tutorial-notebook/notebook\n2. https://www.kaggle.com/tanulsingh077/deep-learning-for-nlp-zero-to-transformers-bert#BERT-and-Its-Implementation-on-this-Competition","metadata":{}}]}