{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Baseline with HuggingFace + Submission for Beginners ","metadata":{}},{"cell_type":"markdown","source":"The notebook contains code to create a baseline by just using `huggingface`. \n* K-fold is not implemnented, but it can easily be done so, by instantiating the trainer for each fold\n* No complex pre-processing or post-processing. Just a simple guide to use huggingface to get a initial baseline\n\n### Salient features\n* **Simple** huggingface API used\n* A simple text normalising pipeline\n* Tries to **account for imbalance** by using class weights\n* **Optimising transformers** with. [(refer this link for more information)](https://huggingface.co/docs/transformers/perf_train_gpu_one#efficient-training-on-a-single-gpu)\n    * Gradient accumulation\n    * Mixed precision training\n    * Auto-choosing batch size (so it fits into memory)\n    * Dynamic padding\n    \n    \nPlease **upvote** if the notebook was helpful :) and feel free to suggest changes or provide feedback!","metadata":{"tags":[]}},{"cell_type":"markdown","source":"**Table of Contents**\n* [Import required libraries](#1)\n* [Define configuration](#2)\n* [Prepare Data](#3)\n    * [Get Data - And apply simple normalization ](#3.1)\n    * [Create a label column and input column](#3.2)\n    * [Split into train-valid](#3.3)\n* [Create tokenized dataset](#4)\n* [Define Dynamic padding](#5)\n* [Define model](#6)\n* [Define Training Arguments](#7)\n    * [Combating class imbalance with class weights](#7.1)\n* [Define Trainer](#8)\n* [Train model](#9)    ","metadata":{}},{"cell_type":"markdown","source":"## Import required libraries <a id=1> </a>","metadata":{"jp-MarkdownHeadingCollapsed":true,"tags":[]}},{"cell_type":"code","source":"import os\nimport re\nimport pandas as pd\nimport numpy as np\nfrom sklearn.model_selection import train_test_split\nfrom sklearn import preprocessing\n\nfrom transformers import AutoTokenizer\nfrom datasets import Dataset\nfrom transformers import DataCollatorWithPadding\nfrom transformers import AutoModelForSequenceClassification, TrainingArguments, Trainer\n\nimport torch\nfrom torch.utils.checkpoint import checkpoint\nimport torch.nn as nn\n\n# You can change this if you want hugginface to automatically log to wandb\nos.environ[\"WANDB_DISABLED\"] = \"true\"\nos.environ[\"TOKENIZERS_PARALLELISM\"] = \"false\"\n\n# Suppress warnings\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:59:59.207054Z","iopub.execute_input":"2022-07-29T07:59:59.207923Z","iopub.status.idle":"2022-07-29T07:59:59.214897Z","shell.execute_reply.started":"2022-07-29T07:59:59.207882Z","shell.execute_reply":"2022-07-29T07:59:59.214019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Define configuration <a id=2> </a>","metadata":{"tags":[]}},{"cell_type":"code","source":"# You can change the model name here, look up model names from huggingface docs\nmodel_name = \"microsoft/deberta-v3-base\"","metadata":{"execution":{"iopub.status.busy":"2022-07-29T07:59:59.748828Z","iopub.execute_input":"2022-07-29T07:59:59.749492Z","iopub.status.idle":"2022-07-29T07:59:59.753575Z","shell.execute_reply.started":"2022-07-29T07:59:59.749457Z","shell.execute_reply":"2022-07-29T07:59:59.752628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prepare Data <a id=3> </a>","metadata":{"tags":[]}},{"cell_type":"markdown","source":"### Get Data - And apply simple normalization <a id=3.1> </a>","metadata":{}},{"cell_type":"code","source":"path = \"train\"\ndef get_essay(essay_id):\n    essay_path = os.path.join(\"../input/feedback-prize-effectiveness/\"+path+\"/\", f\"{essay_id}.txt\")\n    essay_text = open(essay_path, 'r').read()\n    return essay_text\n\ndf = pd.read_csv(\"../input/feedback-prize-effectiveness/train.csv\")\ndf['essay_text'] = df['essay_id'].apply(get_essay)\n\n# This function helps strip extra whitespaces from the text, \n# convert it all to lower case for uniformity, and remove end of line characters\ndef normalise(text):\n    text = text.lower()\n    text = text.strip()\n    text = re.sub(\"\\n\", \" \", text)\n    return text\n\ndf['discourse_text'] = df['discourse_text'].apply(normalise)\ndf['discourse_type'] = df['discourse_type'].apply(normalise)\ndf['essay_text'] = df['essay_text'].apply(normalise)\ndf['discourse_effectiveness'] = df['discourse_effectiveness'].apply(normalise)\n\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:00:00.272285Z","iopub.execute_input":"2022-07-29T08:00:00.273154Z","iopub.status.idle":"2022-07-29T08:00:15.410180Z","shell.execute_reply.started":"2022-07-29T08:00:00.273110Z","shell.execute_reply":"2022-07-29T08:00:15.409103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get tokenizer\ntokenizer = AutoTokenizer.from_pretrained(model_name)\ntokenizer.model_max_length = 512","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:00:15.412409Z","iopub.execute_input":"2022-07-29T08:00:15.413081Z","iopub.status.idle":"2022-07-29T08:00:20.965311Z","shell.execute_reply.started":"2022-07-29T08:00:15.413043Z","shell.execute_reply":"2022-07-29T08:00:20.964227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Create a label column and input column <a id=3.2> </a>\n* Input is going to be [discourse_text + essay_text]\n* Label is going to be numericalised version of the three classes","metadata":{}},{"cell_type":"code","source":"df['text'] = df['discourse_text']+tokenizer.sep_token+df['essay_text']\n\nclasses_to_labels = {\n    \"adequate\":0,\n    \"effective\":1,\n    \"ineffective\":2,\n}\ndf['labels'] = df['discourse_effectiveness'].replace(classes_to_labels)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:00:20.966601Z","iopub.execute_input":"2022-07-29T08:00:20.966960Z","iopub.status.idle":"2022-07-29T08:00:21.066723Z","shell.execute_reply.started":"2022-07-29T08:00:20.966922Z","shell.execute_reply":"2022-07-29T08:00:21.065772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Split into train-valid <a id=3.3> </a>","metadata":{}},{"cell_type":"code","source":"# Stratify by labels makes sure that the distribution of \n# the classes in both train and test remains the same\ntrain_df, valid_df = train_test_split(df, test_size=0.2, random_state=42, stratify=df[\"labels\"])","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:00:21.069450Z","iopub.execute_input":"2022-07-29T08:00:21.069915Z","iopub.status.idle":"2022-07-29T08:00:21.110700Z","shell.execute_reply.started":"2022-07-29T08:00:21.069877Z","shell.execute_reply":"2022-07-29T08:00:21.109822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create tokenized dataset <a id=4> </a>","metadata":{"jp-MarkdownHeadingCollapsed":true,"tags":[]}},{"cell_type":"code","source":"train_dataset = Dataset.from_pandas(train_df)\nvalid_dataset = Dataset.from_pandas(valid_df)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:00:21.112060Z","iopub.execute_input":"2022-07-29T08:00:21.112400Z","iopub.status.idle":"2022-07-29T08:00:21.580258Z","shell.execute_reply.started":"2022-07-29T08:00:21.112364Z","shell.execute_reply":"2022-07-29T08:00:21.579251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def tokenize_function(examples):\n    return tokenizer(examples[\"text\"], padding=\"max_length\", truncation=True)\n\nto_remove = ['discourse_text','discourse_type','text','discourse_id','essay_id', 'essay_text']\n\ntokenized_train_dataset = train_dataset.shuffle(seed=42).map(tokenize_function, batched=True, remove_columns=to_remove)\ntokenized_test_dataset = valid_dataset.shuffle(seed=42).map(tokenize_function, batched=True, remove_columns=to_remove)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:00:21.581761Z","iopub.execute_input":"2022-07-29T08:00:21.582133Z","iopub.status.idle":"2022-07-29T08:01:35.610880Z","shell.execute_reply.started":"2022-07-29T08:00:21.582098Z","shell.execute_reply":"2022-07-29T08:01:35.609915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tokenized_train_dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:01:35.612581Z","iopub.execute_input":"2022-07-29T08:01:35.614441Z","iopub.status.idle":"2022-07-29T08:01:35.621104Z","shell.execute_reply.started":"2022-07-29T08:01:35.614402Z","shell.execute_reply":"2022-07-29T08:01:35.620119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Define Dynamic padding <a id=5> </a>","metadata":{"jp-MarkdownHeadingCollapsed":true,"tags":[]}},{"cell_type":"code","source":"data_collator = DataCollatorWithPadding(tokenizer=tokenizer)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:01:35.622721Z","iopub.execute_input":"2022-07-29T08:01:35.623432Z","iopub.status.idle":"2022-07-29T08:01:35.630813Z","shell.execute_reply.started":"2022-07-29T08:01:35.623395Z","shell.execute_reply":"2022-07-29T08:01:35.629784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Define model <a id=6> </a>","metadata":{"jp-MarkdownHeadingCollapsed":true,"tags":[]}},{"cell_type":"code","source":"model = AutoModelForSequenceClassification.from_pretrained(model_name, num_labels=3)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:01:35.632716Z","iopub.execute_input":"2022-07-29T08:01:35.633918Z","iopub.status.idle":"2022-07-29T08:01:46.476798Z","shell.execute_reply.started":"2022-07-29T08:01:35.633889Z","shell.execute_reply":"2022-07-29T08:01:46.475775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Define Training Arguments <a id=7> </a>","metadata":{"tags":[]}},{"cell_type":"code","source":"training_args = TrainingArguments(\n    output_dir=\"./results\",\n    num_train_epochs=2,\n    evaluation_strategy=\"epoch\",\n    save_strategy=\"epoch\",\n    warmup_ratio=0.1, \n    lr_scheduler_type='cosine',\n    # Optimising\n    auto_find_batch_size=True,\n    # The num of workers may vary for different machines, if you are not sure, just comment this line out\n    dataloader_num_workers=2,\n    gradient_accumulation_steps=4,\n    fp16=True,\n)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:01:46.480463Z","iopub.execute_input":"2022-07-29T08:01:46.480935Z","iopub.status.idle":"2022-07-29T08:01:46.583284Z","shell.execute_reply.started":"2022-07-29T08:01:46.480858Z","shell.execute_reply":"2022-07-29T08:01:46.582386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Combating class imbalance with class weights <a id=7.1> </a>\n\n* We give less weightage to classes with more data samples\n* The idea is that due to high number of samples, a model might inherit bias during training\n* So, we penalise the loss function for that particular class to reduce the chances of bias","metadata":{"tags":[]}},{"cell_type":"code","source":"# Calculating the weights\n# Weightage = 1 - (num_of_samples_of_class)/(total_num_of_samples)\n# less samples, more weightage\n\nw_adequate = 1-len(df[df['discourse_effectiveness'] == 'adequate'])/len(train_df)\nw_effective = 1-len(df[df['discourse_effectiveness'] == 'effective'])/len(train_df)\nw_ineffective = 1-len(df[df['discourse_effectiveness'] == 'ineffective'])/len(train_df)\n\nclass_weights = torch.tensor(\n    [w_adequate, w_effective, w_ineffective]\n).cuda()\n\nclass_weights","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:01:46.584754Z","iopub.execute_input":"2022-07-29T08:01:46.585501Z","iopub.status.idle":"2022-07-29T08:01:51.718102Z","shell.execute_reply.started":"2022-07-29T08:01:46.585464Z","shell.execute_reply":"2022-07-29T08:01:51.717157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# huggingface has no straightforward way to incorparate class_weights as far as I know, \n# Hence we override the compute_loss function of the Trainer and introduce our class weighgts\nclass CustomTrainer(Trainer):\n    def compute_loss(self, model, inputs, return_outputs=False):\n        labels = inputs.get(\"labels\")\n        # forward pass\n        outputs = model(**inputs)\n        logits = outputs.get('logits')\n        # compute custom loss\n        # Class weighting\n        loss_fct = nn.CrossEntropyLoss(weight=class_weights)\n        loss = loss_fct(logits.view(-1, self.model.config.num_labels), labels.view(-1))\n        return (loss, outputs) if return_outputs else loss","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:01:51.719746Z","iopub.execute_input":"2022-07-29T08:01:51.720430Z","iopub.status.idle":"2022-07-29T08:01:51.729661Z","shell.execute_reply.started":"2022-07-29T08:01:51.720381Z","shell.execute_reply":"2022-07-29T08:01:51.727203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Define Trainer <a id=8> </a>","metadata":{"jp-MarkdownHeadingCollapsed":true,"tags":[]}},{"cell_type":"code","source":"trainer = CustomTrainer(\n    model=model,\n    args=training_args,\n    train_dataset=tokenized_train_dataset,\n    eval_dataset=tokenized_test_dataset,\n    tokenizer=tokenizer,\n    data_collator=data_collator,\n)","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:01:51.731048Z","iopub.execute_input":"2022-07-29T08:01:51.731413Z","iopub.status.idle":"2022-07-29T08:01:53.129107Z","shell.execute_reply.started":"2022-07-29T08:01:51.731370Z","shell.execute_reply":"2022-07-29T08:01:53.128098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train the model <a id=9> </a>","metadata":{"jp-MarkdownHeadingCollapsed":true,"tags":[]}},{"cell_type":"code","source":"trainer.train()","metadata":{"execution":{"iopub.status.busy":"2022-07-29T08:01:53.130444Z","iopub.execute_input":"2022-07-29T08:01:53.131564Z"},"trusted":true},"execution_count":null,"outputs":[]}]}