{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<img src=\"https://i.imgur.com/HRzzyIp.png\">\n\n<center><h1> - Training & Tuning using Transformers - </h1></center>\n\n> 📜 **Goal**: Predict the correct ordering of the **cells** within a Jupyter Notebook.\n\n**What is a 🦠 cell** - notice here that a **cell** can be both:\n* a `coding` cell - where you write code\n* a `markdown` cell - where you can write text, add images etc.\n\n**❗ Important** - my initial analysis, understanding and baseline can be found in [📜 AI4Code - Language Detection and Model Tuning](https://www.kaggle.com/code/andradaolteanu/ai4code-language-detection-and-model-tuning). I have took inspiration from [Ahmet Erdem](https://www.kaggle.com/aerdem4)'s work in his notebook [AI4Code Pytorch DistilBert Baseline](https://www.kaggle.com/code/aerdem4/ai4code-pytorch-distilbert-baseline).\n\n### ⬇ Libraries","metadata":{}},{"cell_type":"code","source":"# Libraries\nimport os\nimport gc\nimport wandb\nfrom time import time\nimport random\nimport math\nimport glob\nimport json\nfrom bisect import bisect\nfrom scipy.sparse import vstack\nfrom tqdm import tqdm\nimport warnings\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib as mpl\nfrom matplotlib import cm\nimport matplotlib.patches as patches\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nfrom matplotlib.offsetbox import AnnotationBbox, OffsetImage\nfrom matplotlib.colors import ListedColormap, LinearSegmentedColormap\nfrom matplotlib.patches import Rectangle\nfrom IPython.display import display_html\nplt.rcParams.update({'font.size': 16})\n\n# Sklearn\nfrom sklearn.model_selection import GroupShuffleSplit\nfrom sklearn.metrics import mean_squared_error\n\n# PyTorch\nimport torch\nfrom torch.utils.data import DataLoader, Dataset\nfrom torch.optim import Adam\nimport torch.nn.functional as F\nimport torch.nn as nn\n\n# BERT\nfrom transformers import BertTokenizer, BertModel, DistilBertTokenizer, DistilBertModel\nimport transformers\ntransformers.logging.set_verbosity_error()\n\n\n# Environment check\nwarnings.filterwarnings(\"ignore\")\nos.environ[\"WANDB_SILENT\"] = \"true\"\nCONFIG = {'competition': 'AI4Code', '_wandb_kernel': 'aot'}\n\n# Custom colors\nclass clr:\n    S = '\\033[1m' + '\\033[93m'\n    E = '\\033[0m'\n    \nmy_colors = [\"#CDFC74\", \"#F3EA56\", \"#EBB43D\", \n             \"#DF7D27\", \"#D14417\", \"#B80A0A\", \"#9C0042\"]\nmy_pastels = [\"#A5EC9B\", \"#B4E185\", \"#C3D973\", \n             \"#CDCD61\", \"#CCB049\", \"#CB812D\", \"#B93221\"]\nmy_darks = [\"#FCF238\", \"#F19321\", \"#E54F14\", \n             \"#C22318\", \"#B01028\", \"#9D0642\", \"#85006C\"]\n\ngradient1 = [\"#a5ec9b\", \"#abe890\", \"#b0e485\", \"#b7e07b\", \"#bddb71\", \n             \"#c3d667\", \"#cad15e\", \"#d1cc55\", \"#d7c64e\", \"#dec147\", \"#e5ba41\", \"#ebb43d\"]\nCMAP1 = ListedColormap(my_colors)\nCMAP2 = ListedColormap(my_colors[-1])\n\nprint(clr.S+\"Notebook Color Schemes:\"+clr.E)\nsns.palplot(sns.color_palette(my_pastels))\nsns.palplot(sns.color_palette(my_colors))\nsns.palplot(sns.color_palette(my_darks))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:05:27.877238Z","iopub.execute_input":"2022-07-20T13:05:27.878284Z","iopub.status.idle":"2022-07-20T13:05:28.081651Z","shell.execute_reply.started":"2022-07-20T13:05:27.878241Z","shell.execute_reply":"2022-07-20T13:05:28.080605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 🐝 W&B Fork & Run\n\nIn order to run this notebook you will need to input your own **secret API key** within the `! wandb login $secret_value_0` line. \n\n🐝**How do you get your own API key?**\n\nSuper simple! Go to **https://wandb.ai/site** -> Login -> Click on your profile in the top right corner -> Settings -> Scroll down to API keys -> copy your very own key (for more info check [this amazing notebook for ML Experiment Tracking on Kaggle](https://www.kaggle.com/ayuraj/experiment-tracking-with-weights-and-biases)).\n\n<center><img src=\"https://i.imgur.com/fFccmoS.png\" width=500></center>","metadata":{}},{"cell_type":"code","source":"# # 🐝 Secrets\n# from kaggle_secrets import UserSecretsClient\n# user_secrets = UserSecretsClient()\n# secret_value_0 = user_secrets.get_secret(\"wandb\")\n\n# ! wandb login $secret_value_0","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:05:28.083768Z","iopub.execute_input":"2022-07-20T13:05:28.084334Z","iopub.status.idle":"2022-07-20T13:05:28.088411Z","shell.execute_reply.started":"2022-07-20T13:05:28.084278Z","shell.execute_reply":"2022-07-20T13:05:28.087453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### ⬇ Helper Functions","metadata":{}},{"cell_type":"code","source":"def count_inversions(a):\n    '''src: https://www.kaggle.com/code/ryanholbrook/getting-started-with-ai4code'''\n    inversions = 0\n    sorted_so_far = []\n    for i, u in enumerate(a):\n        j = bisect(sorted_so_far, u)\n        inversions += i - j\n        sorted_so_far.insert(j, u)\n    return inversions\n\n\ndef kendall_tau(ground_truth, predictions):\n    '''src: https://www.kaggle.com/code/ryanholbrook/getting-started-with-ai4code'''\n    total_inversions = 0\n    total_2max = 0  # twice the maximum possible inversions across all instances\n    for gt, pred in zip(ground_truth, predictions):\n        ranks = [gt.index(x) for x in pred]  # rank predicted order in terms of ground truth\n        total_inversions += count_inversions(ranks)\n        n = len(gt)\n        total_2max += n * (n - 1)\n    return 1 - 4 * total_inversions / total_2max\n\n\ndef show_values_on_bars(axs, h_v=\"v\", space=0.4):\n    '''Plots the value at the end of the a seaborn barplot.\n    axs: the ax of the plot\n    h_v: weather or not the barplot is vertical/ horizontal'''\n    \n    def _show_on_single_plot(ax):\n        if h_v == \"v\":\n            for p in ax.patches:\n                _x = p.get_x() + p.get_width() / 2\n                _y = p.get_y() + p.get_height()\n                value = int(p.get_height())\n                ax.text(_x, _y, format(value, ','), ha=\"center\") \n        elif h_v == \"h\":\n            for p in ax.patches:\n                _x = p.get_x() + p.get_width() + float(space)\n                _y = p.get_y() + p.get_height()\n                value = int(p.get_width())\n                ax.text(_x, _y, format(value, ','), ha=\"left\")\n\n    if isinstance(axs, np.ndarray):\n        for idx, ax in np.ndenumerate(axs):\n            _show_on_single_plot(ax)\n    else:\n        _show_on_single_plot(axs)\n        \n        \n# === 🐝 W&B ===\ndef save_dataset_artifact(run_name, artifact_name, path):\n    '''Saves dataset to W&B Artifactory.\n    run_name: name of the experiment\n    artifact_name: under what name should the dataset be stored\n    path: path to the dataset'''\n    \n    run = wandb.init(project='AI4Code', \n                     name=run_name, \n                     config=CONFIG)\n    artifact = wandb.Artifact(name=artifact_name, \n                              type='dataset')\n    artifact.add_file(path)\n\n    wandb.log_artifact(artifact)\n    wandb.finish()\n    print(\"Artifact has been saved successfully.\")\n    \n    \ndef create_wandb_plot(x_data=None, y_data=None, x_name=None, y_name=None, title=None, log=None, plot=\"line\"):\n    '''Create and save lineplot/barplot in W&B Environment.\n    x_data & y_data: Pandas Series containing x & y data\n    x_name & y_name: strings containing axis names\n    title: title of the graph\n    log: string containing name of log'''\n    \n    data = [[label, val] for (label, val) in zip(x_data, y_data)]\n    table = wandb.Table(data=data, columns = [x_name, y_name])\n    \n    if plot == \"line\":\n        wandb.log({log : wandb.plot.line(table, x_name, y_name, title=title)})\n    elif plot == \"bar\":\n        wandb.log({log : wandb.plot.bar(table, x_name, y_name, title=title)})\n    elif plot == \"scatter\":\n        wandb.log({log : wandb.plot.scatter(table, x_name, y_name, title=title)})\n        \n        \ndef create_wandb_hist(x_data=None, x_name=None, title=None, log=None):\n    '''Create and save histogram in W&B Environment.\n    x_data: Pandas Series containing x values\n    x_name: strings containing axis name\n    title: title of the graph\n    log: string containing name of log'''\n    \n    data = [[x] for x in x_data]\n    table = wandb.Table(data=data, columns=[x_name])\n    wandb.log({log : wandb.plot.histogram(table, x_name, title=title)})\n    \n    \n# 🐝 Log Cover Photo\n# run = wandb.init(project='AI4Code', name='CoverPhoto', config=CONFIG)\n# cover = plt.imread(\"../input/ai4code-processed-data/AI4Code Cover.png\")\n# wandb.log({\"example\": wandb.Image(cover)})\n# wandb.finish()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-07-20T13:05:28.090167Z","iopub.execute_input":"2022-07-20T13:05:28.090736Z","iopub.status.idle":"2022-07-20T13:05:28.123646Z","shell.execute_reply.started":"2022-07-20T13:05:28.090701Z","shell.execute_reply":"2022-07-20T13:05:28.122595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 🌍 Global Parameters\n\nThis is the area where I will be setting the *GLOBAL* parameters for this notebook.\n\n📜 **Things to note**:\n* at first I tried using `bert_base_uncased`, however the model is too large for the notebook and the memory runs out during training.\n* hence I settled for `distilbert-base-uncased`, which is *smaller and faster*.","metadata":{}},{"cell_type":"code","source":"# ~~~~~~~~~~~~~ PARAMS ~~~~~~~~~~~~~\nSEED = 24\nDF_SIZE = 0.002       # set less than 1 to work faster with smaller dataset\nMAX_LEN = 512         # length of the tokenizer\nLAYER_SIZE = 768\nNVALID = 0.1          # size of validation set\nLR = 0.0005\nBATCH_SIZE = 16       # if set bigger -> memory runs out\nDEVICE = torch.device('cuda:0')\nEPOCHS = 1\n\nDISTIL = True         # whether we wanna use distil or simple bert\n# BERT_MODEL = 'bert-base-uncased'\nBERT_MODEL = 'distilbert-base-uncased'\n# ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:05:28.125556Z","iopub.execute_input":"2022-07-20T13:05:28.126107Z","iopub.status.idle":"2022-07-20T13:05:28.136455Z","shell.execute_reply.started":"2022-07-20T13:05:28.126068Z","shell.execute_reply":"2022-07-20T13:05:28.135441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Dataset Import\n\n> 📌 **Note**: The dataframe below [was taken from my dataset](https://www.kaggle.com/datasets/andradaolteanu/ai4code-processed-data) on this competition. For more details on how it was made you can also check out [my previous noteook](https://www.kaggle.com/code/andradaolteanu/ai4code-language-detection-and-model-tuning).\n\n📜 **The dataset contains the following columns**:\n* `cell_id` - unique id number of each cell in a notebook\n* `cell_type` - the cell type - either **code** or **markdown**\n* `source` - the text within each cell\n* `id` - unique number for each notebook\n* `rank` - the **order** of the cells within the notebook\n* `percent_rank` - the order divided by the total number of cells in each notebook\n* `ancestor_id` - to be used when doing data validation\n* `parent_id` - not always available\n\n*! `.parquet` format is faster and smaller, so I used it instead of the usual `.csv` format.*\n\n*! Also keep in mind that the original `dataframe` has a total number of rows of: 5,785,610.*","metadata":{}},{"cell_type":"code","source":"random.seed(SEED)\n\n# Read in the training dataset\ndf = pd.read_parquet(\"../input/ai4code-processed-data/train.parquet\")\n\n# Get all unique ids\nunique_ids = df[\"id\"].unique().tolist()\n\n# Sample down\nunique_ids = random.sample(unique_ids, k=int(len(unique_ids)*DF_SIZE))\ndf = df[df[\"id\"].isin(unique_ids)].reset_index(drop=True)\n\n# Shuffle\ndf = df.reindex(np.random.permutation(df.index))\n\nprint(clr.S+\"Sampled dataframe Shape:\"+clr.E, df.shape)\n\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:05:28.140694Z","iopub.execute_input":"2022-07-20T13:05:28.141569Z","iopub.status.idle":"2022-07-20T13:05:41.778457Z","shell.execute_reply.started":"2022-07-20T13:05:28.141533Z","shell.execute_reply":"2022-07-20T13:05:41.777504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del unique_ids\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:05:41.779803Z","iopub.execute_input":"2022-07-20T13:05:41.780476Z","iopub.status.idle":"2022-07-20T13:05:41.973791Z","shell.execute_reply.started":"2022-07-20T13:05:41.780438Z","shell.execute_reply":"2022-07-20T13:05:41.972862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. PyTorch Dataset\n\n> 📌 **Note**: This is just a custom function for us to be able to read the information in the correct format and then to pass it to the model.","metadata":{}},{"cell_type":"code","source":"class AI4CodeDataset(Dataset):\n    \n    def __init__(self, data, max_len):\n        super().__init__()\n        \n        # Use the correct tokenizer for each option\n        if DISTIL:\n            self.tokenizer = DistilBertTokenizer.from_pretrained(BERT_MODEL)\n        else:\n            self.tokenizer = BertTokenizer.from_pretrained(BERT_MODEL)\n    \n        # The maximum length in number of tokens for the inputs to the BERT model\n        self.max_len = max_len\n        self.data = data\n        \n    def __len__(self):\n        return self.data.shape[0]\n\n    def __getitem__(self, index):\n        # Get one row at a time\n        row = self.data.iloc[index]\n        text = row.source\n        \n        # Tokenize the text\n        inputs = self.tokenizer(text,\n                                max_length=self.max_len,\n                                padding=\"max_length\",\n                                truncation=True)\n        \n        ids = torch.tensor(inputs['input_ids'], dtype=torch.long)\n        mask = torch.tensor(inputs['attention_mask'], dtype=torch.long)\n        target = torch.tensor(row.percent_rank, dtype=torch.float)\n        \n        # Return the encoded info & variable to predict\n        return {\"input_ids\" : ids,\n                \"attention_mask\" : mask,\n                \"target\" : target}","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:05:41.975542Z","iopub.execute_input":"2022-07-20T13:05:41.975945Z","iopub.status.idle":"2022-07-20T13:05:41.987413Z","shell.execute_reply.started":"2022-07-20T13:05:41.975909Z","shell.execute_reply":"2022-07-20T13:05:41.986540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Sanity Check\n\n📜 Below is a schema that explains how this function works:\n\n<center><img src=\"https://i.imgur.com/RvPSmXo.png\" width=1000></center>","metadata":{}},{"cell_type":"code","source":"# Sample data\nsample_df = df.head(6)\n\n# Instantiate Dataset object\ndataset = AI4CodeDataset(data=sample_df, max_len=MAX_LEN)\n# The Dataloader\ndataloader = DataLoader(dataset, batch_size=3, shuffle=False)\n\n# Output of the Dataloader\nfor k, data in enumerate(dataloader):\n    ids, mask, target = data.values()\n    print(clr.S + f\"Batch: {k}\" + clr.E, \"\\n\" +\n          clr.S + \"Ids:\" + clr.E, ids, \"\\n\" +\n          clr.S + \"Mask:\" + clr.E, mask, \"\\n\" +\n          clr.S + \"Target:\" + clr.E, target, \"\\n\" +\n          \"=\"*50)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:05:41.988846Z","iopub.execute_input":"2022-07-20T13:05:41.989276Z","iopub.status.idle":"2022-07-20T13:05:45.162292Z","shell.execute_reply.started":"2022-07-20T13:05:41.989240Z","shell.execute_reply":"2022-07-20T13:05:45.161047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Bonus\n\nI will also create a function that automatically takes all the info passed out by `AI4CodeDataset()` and adds it to cuda.","metadata":{}},{"cell_type":"code","source":"def data_to_device(data):\n    \n    ids, mask, target = data.values()\n    return ids.to(DEVICE), mask.to(DEVICE), target.to(DEVICE)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:05:45.167087Z","iopub.execute_input":"2022-07-20T13:05:45.169668Z","iopub.status.idle":"2022-07-20T13:05:45.176615Z","shell.execute_reply.started":"2022-07-20T13:05:45.169629Z","shell.execute_reply":"2022-07-20T13:05:45.175477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Model\n\n🤗 [bert-base-uncased](https://huggingface.co/bert-base-uncased) - BERT is a **transformers** model pretrained on a large corpus of English data in a self-supervised fashion. This means it was pretrained on the raw texts only, with no humans labelling them in any way (which is why it can use lots of publicly available data) with an automatic process to generate inputs and labels from those texts.\n\n*!uncased - it doesn't make the difference between upper and lower case.*\n\n🤗 [distilbert-base-uncased](https://huggingface.co/distilbert-base-uncased) - (distilled) version of the normal BERT - is faster and smaller, so it's much easier to be used within this environment.","metadata":{}},{"cell_type":"code","source":"class TransformersModel(nn.Module):\n    \n    def __init__(self, bert_model, layer_size):\n        super(TransformersModel, self).__init__()\n        \n        # Initiate the model accordingly\n        if DISTIL:\n            self.Model = DistilBertModel.from_pretrained(bert_model)\n        else:\n            self.Model = BertModel.from_pretrained(bert_model)\n        \n        # Create a final Linear Layer\n        self.Linear = nn.Linear(layer_size, 1)\n        \n    def forward(self, ids, mask, prints=False):\n        '''A forward pass of this network.\n        Use `prints=True` if you want to see output shape at each pass.'''\n            \n        if DISTIL:\n            text = self.Model(ids, mask)[0]\n            out = self.Linear(text)\n            out = out[:, 0, :]\n        else:\n            _, text = self.Model(ids, mask, return_dict=False)\n            out = self.Linear(text)\n            out = out.view(-1)\n        \n        \n        if prints:\n            print(\"===============\")\n            print(clr.S+\"Text Out Shape:\"+clr.E, text.shape)\n            print(clr.S+\"After FNN Shape:\"+clr.E, out.shape)\n            print(clr.S+\"Output Shape:\"+clr.E, out.shape)\n            \n        return out","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:05:45.181257Z","iopub.execute_input":"2022-07-20T13:05:45.183745Z","iopub.status.idle":"2022-07-20T13:05:45.197740Z","shell.execute_reply.started":"2022-07-20T13:05:45.183706Z","shell.execute_reply":"2022-07-20T13:05:45.196544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Sanity Check\n\n📜 Let's do the same for the `TransformersModel()` class and see if this one works fine before going further into creating the *training* functions.\n\n<center><img src=\"https://i.imgur.com/jrJG7l1.png\" width=1000></center>","metadata":{}},{"cell_type":"code","source":"# Initiate the model\nmodel_example = TransformersModel(BERT_MODEL, layer_size=LAYER_SIZE)\nmodel_example.train()  ### training mode: ON\n\n# We'll use the dataset & dataloader from previous example\nfor k, data in enumerate(dataloader):\n    ids, mask, target = data.values()\n    break\n    \nprint(clr.S+\"Input data shape:\"+clr.E, len(ids), \"paragraphs.\", \"\\n\")\n\n# Make a prediction\nout = model_example(ids, mask, prints=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:05:45.203476Z","iopub.execute_input":"2022-07-20T13:05:45.206176Z","iopub.status.idle":"2022-07-20T13:06:01.320030Z","shell.execute_reply.started":"2022-07-20T13:05:45.206137Z","shell.execute_reply":"2022-07-20T13:06:01.318976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Optimizer\n\n📜 Create a customized `optimizer` (like I did in [📖 II.CommonLit: BERT vs RoBERTa + W&B testing](https://www.kaggle.com/code/andradaolteanu/ii-commonlit-bert-vs-roberta-w-b-testing)).","metadata":{}},{"cell_type":"code","source":"def custom_optimizer(model, LR, prints=False):\n    '''A custom optimizer for the model parameters.\n    model: the initialized model class\n    lr: learning rate'''\n    \n    # Get model parameters\n    parameters = list(model.named_parameters())\n    if prints:\n        print(clr.S+\"Number of Parameters to Optimize:\"+clr.E, len(parameters))\n        \n    no_decay = [\"bias\", \"LayerNorm.bias\", \"LayerNorm.weight\"]\n    \n    # Set weight_decay to start at 0.0 for the no_decay parameters\n    ### and the rest to start from 0.003\n    optimizer_parameters = [\n            {\"params\" : [p for n, p in parameters if not any(nd in n for nd in no_decay)],\n             \"weight_decay\" : 0.003}, \n            {\"params\" : [p for n, p in parameters if any(nd in n for nd in no_decay)],\n             \"weight_decay\" : 0.0}]\n    \n    optimizer = Adam(optimizer_parameters, lr=LR)\n    if prints: print(clr.S+\"Optimizer:\"+clr.E, optimizer)\n    \n    return optimizer","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:06:01.321403Z","iopub.execute_input":"2022-07-20T13:06:01.322565Z","iopub.status.idle":"2022-07-20T13:06:01.331664Z","shell.execute_reply.started":"2022-07-20T13:06:01.322523Z","shell.execute_reply":"2022-07-20T13:06:01.330807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Optimizer Example\noptimizer_example = custom_optimizer(model_example, LR=0.0005, prints=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:06:01.333104Z","iopub.execute_input":"2022-07-20T13:06:01.333465Z","iopub.status.idle":"2022-07-20T13:06:01.761034Z","shell.execute_reply.started":"2022-07-20T13:06:01.333431Z","shell.execute_reply":"2022-07-20T13:06:01.759710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Delete redundant variables\ndel model_example, dataloader, data, ids, mask, target\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:06:01.765701Z","iopub.execute_input":"2022-07-20T13:06:01.765981Z","iopub.status.idle":"2022-07-20T13:06:01.958943Z","shell.execute_reply.started":"2022-07-20T13:06:01.765956Z","shell.execute_reply":"2022-07-20T13:06:01.957974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Data Validation\n\n📜 **Group Shuffle Split Method**: [Provides randomized train/test indices to split data according to a third-party provided group.](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupShuffleSplit.html).\n\n<center><img src=\"https://miro.medium.com/max/1200/0*gWUAJ45fyETmwgl4.png\" width=800></center>","metadata":{}},{"cell_type":"code","source":"# Split based on acestor_id\nsplitter = GroupShuffleSplit(n_splits=1, test_size=NVALID, \n                             random_state=0)\n\nids_train, ids_valid = next(splitter.split(df, groups=df[\"ancestor_id\"]))\n\n# Separate into training and validation data\ntrain = df.loc[ids_train].reset_index(drop=True)\nvalid = df.loc[ids_valid].reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:06:01.960418Z","iopub.execute_input":"2022-07-20T13:06:01.960861Z","iopub.status.idle":"2022-07-20T13:06:01.986764Z","shell.execute_reply.started":"2022-07-20T13:06:01.960822Z","shell.execute_reply":"2022-07-20T13:06:01.985925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Train Function\n\n### Dataloaders Helper\n\n📜 This one is to shorten our code within the training function. It creates the dataset and dataloader instances from the `train` and `valid` datasets.","metadata":{}},{"cell_type":"code","source":"def get_loaders(train, valid):\n    '''Function to create the dataset and dataloaders from dataframes.\n    train: training data\n    valid: validation data\n    return :: train and valid loaders (in this order)'''\n\n    # Dataset Objects\n    train_dataset = AI4CodeDataset(data=train, max_len=MAX_LEN)\n    valid_dataset = AI4CodeDataset(data=valid, max_len=MAX_LEN)\n\n    # Dataloaders\n    train_loader = DataLoader(train_dataset, batch_size=BATCH_SIZE, \n                              num_workers=8, shuffle=True)\n    valid_loader = DataLoader(valid_dataset, batch_size=BATCH_SIZE, \n                              num_workers=8, shuffle=False)\n    \n    return train_loader, valid_loader","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:06:01.987989Z","iopub.execute_input":"2022-07-20T13:06:01.988414Z","iopub.status.idle":"2022-07-20T13:06:01.994813Z","shell.execute_reply.started":"2022-07-20T13:06:01.988379Z","shell.execute_reply":"2022-07-20T13:06:01.993828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training Function\n\n📜 The schema below shows a simplified explanation of the `model_trainer()` function:\n\n<center><img src=\"https://i.imgur.com/xSaoPNR.png\" width=1000></center>","metadata":{}},{"cell_type":"code","source":"def model_trainer():\n    '''Function to train the model.'''\n    \n    # 🐝 W&B Experiment\n    config_defaults = {\"layer_size\" : 768,\n                       \"lr\" : 0.0005,\n                       \"epochs\" : 3}\n    config_defaults.update(CONFIG)\n    \n    with wandb.init(project='AI4Code', name='distilBert', config=config_defaults):\n        config = wandb.config\n    \n        # Get loaders\n        train_loader, valid_loader = get_loaders(train, valid)\n\n        # Get Model\n        model = TransformersModel(BERT_MODEL, layer_size=config.layer_size).to(DEVICE)\n        # Optimizer & Criterion\n        optimizer = custom_optimizer(model, config.lr)\n        criterion = nn.MSELoss()\n\n\n        for epoch in range(config.epochs):\n            print(clr.S + f\"--- Epoch {epoch} ---\" + clr.E)\n\n            # -Train the Model-\n            model.train()   ### Training Mode: ON\n\n            losses = []\n            for k, data in tqdm(enumerate(train_loader)):\n                ids, mask, target = data_to_device(data)\n\n                optimizer.zero_grad()\n                out = model(ids, mask)\n                loss = criterion(out, target)\n                loss.backward()\n                optimizer.step()\n\n                losses.append(loss.cpu().detach().numpy().tolist())\n\n            # Log Training Voss into the experiment\n            train_loss = np.mean(losses)\n            wandb.log({\"epoch_mean_loss\": np.float(train_loss)})\n            print(\"Epoch Mean Loss:\", train_loss)\n\n\n            # -Evaluate the Model-\n            model.eval()   ### Evaluation Mode: ON\n\n            valid_preds, valid_targets = [], []\n            with torch.no_grad():\n                for k, data in tqdm(enumerate(valid_loader)):\n                    ids, mask, target = data_to_device(data)\n\n                    out = model(ids, mask)\n\n                    valid_preds.append(out.detach().cpu().numpy().ravel())\n                    valid_targets.append(target.detach().cpu().numpy().ravel())\n\n\n            # -Final Results-\n            valid_preds = np.concatenate(valid_preds)\n            valid_targets = np.concatenate(valid_targets)\n            rmse = mean_squared_error(valid_targets, valid_preds)\n            # TODO: Implement Kendall too\n    #         kendall = kendall_tau(valid_targets, valid_preds)\n\n            wandb.log({\"epoch_RMSE\": np.float(train_loss), \"epoch\": epoch})\n    #         wandb.log({\"epoch_kendall\": np.float(kendall)})\n            print(f\"epoch_RMSE: {rmse}\")","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:06:01.996362Z","iopub.execute_input":"2022-07-20T13:06:01.996993Z","iopub.status.idle":"2022-07-20T13:06:02.013142Z","shell.execute_reply.started":"2022-07-20T13:06:01.996956Z","shell.execute_reply":"2022-07-20T13:06:02.012218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### ❚·══·❚ Initial Training\n\nFirst model results:","metadata":{}},{"cell_type":"code","source":"model_trainer()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T13:06:02.014665Z","iopub.execute_input":"2022-07-20T13:06:02.015022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 7. Hyperparameter Tuning [🐝 W&B Sweeps]\n\nAs done in my [first notebook](https://www.kaggle.com/code/andradaolteanu/ai4code-language-detection-and-model-tuning), I will continue to be using the W&B integrated [Sweeps for XGBoost](https://docs.wandb.ai/guides/integrations/xgboost) method to log all my experiments.\n\n*🙏 The tutorial I am following is [Using_W&B_Sweeps_with_XGBoost](https://colab.research.google.com/github/wandb/examples/blob/master/colabs/boosting/Using_W%26B_Sweeps_with_XGBoost.ipynb#scrollTo=VCRlDRL6_5aA).*\n\n<center><img src=\"https://i.imgur.com/v949kdT.png\"></center>\n\n> ❗ **Note:** for the good functioning of Sweeps it is very important that the training function aka `train_XGBRanker()` does NOT have any **arguments** passed. Hence, the format of the `wandb.agent()` should always be `wandb.agent(sweep_id, train_XGBRanker, count=20)` and **NOT** `wandb.agent(sweep_id, train_XGBRanker(data, model, config), count=20)`.","metadata":{}},{"cell_type":"code","source":"# Sweep Config\nsweep_config = {\n    \"method\": \"random\", # grid for all\n    \"metric\": {\n      \"name\": \"epoch_RMSE\",\n      \"goal\": \"minimize\"   \n    },\n    \"parameters\": {\n        \"layer_size\": {\n            \"values\": [300, 500, 768]\n        },\n        \"lr\": {\n            'distribution': 'uniform',\n                            'max': 0.1,\n                            'min': 0\n        },\n        \"epochs\": {\n            \"values\": [3]\n        }\n    }\n}\n\n# Sweep ID\nsweep_id = wandb.sweep(sweep_config, project=\"AI4Code\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 RUN SWEEPS\nstart = time()\n\n# count = the number of trials/experiments to run\nwandb.agent(sweep_id, model_trainer, count=3)\nprint(\"Sweeping took:\", round((time()-start)/60, 1), \"mins\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# <center><video src=\"mp4\" width=800 controls></center>","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center><img src=\"https://i.imgur.com/0cx4xXI.png\"></center>\n\n### 🐝 W&B Dashboard\n\n> My [W&B Dashboard](https://wandb.ai/andrada/AI4Code?workspace=user-andrada).\n\n<center><img src=\"https://i.imgur.com/ZRwRJcw.png\"></center>\n\n<center><img src=\"https://i.imgur.com/knxTRkO.png\"></center>\n\n### My Specs\n\n* 🖥 Z8 G4 Workstation\n* 💾 2 CPUs & 96GB Memory\n* 🎮 2x NVIDIA A6000\n* 💻 Zbook Studio G7 on the go","metadata":{}}]}