{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# DeBLURRTa - A baseline model using fast.ai and blurr\n\nThis is the second shared notebook in my **blurrified series** where I show folks how they can use [blurr](https://github.com/ohmeow/blurr) for NLP tasks such as the one included in the [Feedback Prize - Effective Arguments](https://www.kaggle.com/competitions/feedback-prize-effectiveness) competition. For those not in-the-know, **BLURR** is a library I created for developers who want to train [Hugging Face transformers](https://huggingface.co/docs/transformers/main/en/index) using [fast.ai](https://docs.fast.ai/).\n\nIn my first notebook, [\"Iterate like a grandmaster ... blurrified!\"](https://www.kaggle.com/code/ohmeow/iterate-like-a-grandmaster-blurrified), I showed how blurr can be used for a NLP regression task.  Here, we'll turn our attention to setting things up for a **multiclass classification** task along with the postprocessing bits required to make a submission.  As you'll see, with blurr and fast.ai, only minimal changes will be required.\n\nAs before, I won't be focusing much on EDA and will instead link to some notebooks I'm finding helpful in this competition below.\n\nSo without further ado, lets go!!!","metadata":{}},{"cell_type":"code","source":"import os, sys","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:11.669025Z","iopub.execute_input":"2022-07-06T22:07:11.669432Z","iopub.status.idle":"2022-07-06T22:07:11.694017Z","shell.execute_reply.started":"2022-07-06T22:07:11.669342Z","shell.execute_reply":"2022-07-06T22:07:11.693041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Note**: I've created a couple Kaggle datasets from notebooks so that everything works offline.  Care to know more, you can check them out here ... [ohmeow-blurr-pip-installs](https://www.kaggle.com/ohmeow/ohmeow-blurr-pip-installs) and [ohmeow-transformers](https://www.kaggle.com/ohmeow/ohmeow-transformers). \n\nAnd yah, if they're helpful for ya ... feel free to give em' a proper upvote 👍","metadata":{}},{"cell_type":"code","source":"os.system('python -m pip install --no-index --find-links=../input/ohmeow-blurr-pip-installs transformers ohmeow-blurr')","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:11.763915Z","iopub.execute_input":"2022-07-06T22:07:11.764472Z","iopub.status.idle":"2022-07-06T22:07:22.298312Z","shell.execute_reply.started":"2022-07-06T22:07:11.764434Z","shell.execute_reply":"2022-07-06T22:07:22.297338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 1: Install and imports","metadata":{}},{"cell_type":"code","source":"import gc\n\nfrom fastai.callback.all import *\nfrom fastai.data.block import CategoryBlock, ColReader, DataBlock, IndexSplitter, RegressionBlock\nfrom fastai.imports import *\nfrom fastai.learner import *\nfrom fastai.optimizer import Adam\nfrom fastai.metrics import *\nfrom fastai.torch_core import *\nfrom fastai.torch_imports import *\nfrom transformers import AutoModelForSequenceClassification, logging\n\nfrom blurr.text.data.core import TextBlock\nfrom blurr.text.modeling.core import BaseModelWrapper, BaseModelCallback, blurr_splitter\nfrom blurr.text.utils import get_hf_objects\nfrom blurr.utils import PreCalculatedCrossEntropyLoss, print_versions, set_seed\n\n# silence all the HF warnings\nwarnings.simplefilter(\"ignore\")\nlogging.set_verbosity_error()\nos.environ[\"TOKENIZERS_PARALLELISM\"] = \"false\"","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:22.300536Z","iopub.execute_input":"2022-07-06T22:07:22.301253Z","iopub.status.idle":"2022-07-06T22:07:30.718218Z","shell.execute_reply.started":"2022-07-06T22:07:22.301213Z","shell.execute_reply":"2022-07-06T22:07:30.717135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"device = \"cuda:0\"\nrandom_seed = 9\n\ntorch.cuda.set_device(device)\nprint(f\"Using GPU #{torch.cuda.current_device()}: {torch.cuda.get_device_name()}\")","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:30.720128Z","iopub.execute_input":"2022-07-06T22:07:30.721121Z","iopub.status.idle":"2022-07-06T22:07:30.727760Z","shell.execute_reply.started":"2022-07-06T22:07:30.721082Z","shell.execute_reply":"2022-07-06T22:07:30.726584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**TIP**: Set the `random_seed` = None\n\nHere's a few quotes from Jeremy Howard made during the 2022 fastai course:\n\n> I almost never use a random seed ... to get an intuitive sense [of the stability of the model]\"\n\n> I would never set a random seed; I want to be able to run things multiple times and see how much it changes each time because that will give me a sense of [whether] the modifications I'm making are changing it because its improving it or is it just random variation.  You won't be able to see that if you always set a random seed.\"\n\n\nI'm using it here simpley because reproducibility helps to clarify the prose :)","metadata":{}},{"cell_type":"markdown","source":"As before, we'll define a variable, `is_kaggle`, to tell us whether we are running locally or on kaggle and get a `Path` reference to where the competition data lives.","metadata":{}},{"cell_type":"code","source":"is_kaggle = os.environ.get(\"KAGGLE_KERNEL_RUN_TYPE\", \"\")\n\nis_kaggle","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:30.731148Z","iopub.execute_input":"2022-07-06T22:07:30.731642Z","iopub.status.idle":"2022-07-06T22:07:30.740365Z","shell.execute_reply.started":"2022-07-06T22:07:30.731603Z","shell.execute_reply":"2022-07-06T22:07:30.739225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**TIP**: To maximize my use of Kaggle's ~30hrs of GPU, I tend to try and develop locally as much as possible before moving my work to the cloud. And just to show you that you don't need expensive GPUs to do that, everything I'm doing here was initially trained on a lowly 1080Ti (yes, I said 1080ti 😂)","metadata":{}},{"cell_type":"code","source":"if is_kaggle:\n    comp_data_path = Path(\"../input/feedback-prize-effectiveness\")\nelse:\n    comp_data_path = Path(\"../data/comp\")\n\ncomp_data_path.ls()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:30.742172Z","iopub.execute_input":"2022-07-06T22:07:30.742597Z","iopub.status.idle":"2022-07-06T22:07:30.754274Z","shell.execute_reply.started":"2022-07-06T22:07:30.742561Z","shell.execute_reply":"2022-07-06T22:07:30.752972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**TIP**: I tend to end up referencing serveral `Path`s so I recommend naming them so it's clear as to which is which. \n\n**TIP**: Locally, I like to have the competition data sit in a `../data/comp` folder and any cleaned data I'm using for EDA and feature engineering consideration in `../data/clean`.","metadata":{}},{"cell_type":"markdown","source":"## Step 2: Data and EDA\n\nAs mentioned above, instead of essentially duplicating the EDA others have already shared, I'll link to some of this competition's EDA notebooks I found helpful below. Here we'll only take a cursory look at the dataset to get an initial intuition for setting up our inputs.\n\n**TIP**: Training a baseline model ***should be considered part of your EDA***! You'll learn a lot about it but training a baseline model early in your project.","metadata":{}},{"cell_type":"markdown","source":"Let's look at the training set:","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(comp_data_path/\"train.csv\")\n\nprint(\"Length or TRAINING dataset:\", len(train_df))\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:30.755660Z","iopub.execute_input":"2022-07-06T22:07:30.756637Z","iopub.status.idle":"2022-07-06T22:07:31.091283Z","shell.execute_reply.started":"2022-07-06T22:07:30.756597Z","shell.execute_reply":"2022-07-06T22:07:31.090146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.describe(include=\"all\").T","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:31.093179Z","iopub.execute_input":"2022-07-06T22:07:31.093861Z","iopub.status.idle":"2022-07-06T22:07:31.185359Z","shell.execute_reply.started":"2022-07-06T22:07:31.093819Z","shell.execute_reply":"2022-07-06T22:07:31.184375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:31.187002Z","iopub.execute_input":"2022-07-06T22:07:31.187619Z","iopub.status.idle":"2022-07-06T22:07:31.215791Z","shell.execute_reply.started":"2022-07-06T22:07:31.187578Z","shell.execute_reply":"2022-07-06T22:07:31.214828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"... and the test set","metadata":{}},{"cell_type":"code","source":"test_df = pd.read_csv(comp_data_path/\"test.csv\")\n\nprint(\"Length or TEST dataset:\", len(train_df))\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:31.219115Z","iopub.execute_input":"2022-07-06T22:07:31.219412Z","iopub.status.idle":"2022-07-06T22:07:31.239087Z","shell.execute_reply.started":"2022-07-06T22:07:31.219383Z","shell.execute_reply":"2022-07-06T22:07:31.238009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some bits that may be helpful in your EDA ...","metadata":{}},{"cell_type":"markdown","source":"Min/Avg/Max `discourse_text` lengths","metadata":{}},{"cell_type":"code","source":"text_lengths = [len(x) for x in train_df[\"discourse_text\"]]\n\nmin(text_lengths), np.mean(text_lengths), max(text_lengths)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:31.243508Z","iopub.execute_input":"2022-07-06T22:07:31.243867Z","iopub.status.idle":"2022-07-06T22:07:31.269464Z","shell.execute_reply.started":"2022-07-06T22:07:31.243839Z","shell.execute_reply":"2022-07-06T22:07:31.268568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The number of discourses per essay\n","metadata":{}},{"cell_type":"code","source":"train_df[\"essay_id\"].value_counts()[-10:]","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:31.272488Z","iopub.execute_input":"2022-07-06T22:07:31.272771Z","iopub.status.idle":"2022-07-06T22:07:31.291074Z","shell.execute_reply.started":"2022-07-06T22:07:31.272745Z","shell.execute_reply":"2022-07-06T22:07:31.290247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Min/Avg/Max essay lengths","metadata":{}},{"cell_type":"code","source":"essay_lengths = []\n\nfor file in os.listdir(comp_data_path/'train'):\n    with open(comp_data_path/'train'/file) as f:\n        essay_lengths.append(len(f.read()))\n\nmin(essay_lengths), np.mean(essay_lengths), max(essay_lengths)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:31.293226Z","iopub.execute_input":"2022-07-06T22:07:31.293982Z","iopub.status.idle":"2022-07-06T22:07:43.657737Z","shell.execute_reply.started":"2022-07-06T22:07:31.293943Z","shell.execute_reply":"2022-07-06T22:07:43.656856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `discourse_type` distribution","metadata":{}},{"cell_type":"code","source":"train_df[\"discourse_type\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:43.659324Z","iopub.execute_input":"2022-07-06T22:07:43.659915Z","iopub.status.idle":"2022-07-06T22:07:43.671770Z","shell.execute_reply.started":"2022-07-06T22:07:43.659874Z","shell.execute_reply":"2022-07-06T22:07:43.670817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The target distribution","metadata":{}},{"cell_type":"code","source":"train_df[\"discourse_effectiveness\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:43.673012Z","iopub.execute_input":"2022-07-06T22:07:43.674018Z","iopub.status.idle":"2022-07-06T22:07:43.696100Z","shell.execute_reply.started":"2022-07-06T22:07:43.673979Z","shell.execute_reply":"2022-07-06T22:07:43.695057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The target distribution by `discourse_type`","metadata":{}},{"cell_type":"code","source":"train_df[[\"discourse_type\", \"discourse_effectiveness\"]].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:43.699348Z","iopub.execute_input":"2022-07-06T22:07:43.700058Z","iopub.status.idle":"2022-07-06T22:07:43.725514Z","shell.execute_reply.started":"2022-07-06T22:07:43.700025Z","shell.execute_reply":"2022-07-06T22:07:43.724499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 3: Get your Hugging Face objects\n\nWe'll use blurr's `get_hf_objects` method to get all the Hugging Face objects needed using the \"classification go to\" deberta-v3-small pretrained checkpoint.  \n\nThe objective of this competition is to predict the probabilities a given `discourse_id` should be labeled as either Ineffective, Adequate, or Effective. Since the default number of labels for `AutoModelForSequenceClassification` models is 2, we'll need to it to output 3 activations via the `config_kwargs` parameter.","metadata":{}},{"cell_type":"code","source":"model_checkpoint = \"../input/ohmeow-transformers/deberta-v3-small\"\nsep = \" [s] \"\nnew_special_toks = [sep]\n\nhf_arch, hf_config, hf_tokenizer, hf_model = get_hf_objects(model_checkpoint, model_cls=AutoModelForSequenceClassification, config_kwargs={\"num_labels\": 3})\n\nif new_special_toks:\n    # After adding the new tokens, we need to resize the embedding matrix in the model and initialize the weights\n    hf_tokenizer.add_special_tokens({\"additional_special_tokens\": new_special_toks})\n    hf_model.resize_token_embeddings(len(hf_tokenizer))\n\n    with torch.no_grad():\n        emb_size = hf_model.config.to_dict().get(\"embedding_size\", hf_model.config.hidden_size)\n        hf_model.get_input_embeddings().weight[-len(hf_tokenizer), :] = torch.zeros([emb_size])","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:43.726789Z","iopub.execute_input":"2022-07-06T22:07:43.727137Z","iopub.status.idle":"2022-07-06T22:07:50.571772Z","shell.execute_reply.started":"2022-07-06T22:07:43.727098Z","shell.execute_reply":"2022-07-06T22:07:50.570770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**TIP**: Not setting the `num_labels` correctly is the most frequent cause of errors folks encounter encounter when training transformers for classification and regression tasks. Always check this and remember that for regression, it should be set to 1.","metadata":{}},{"cell_type":"markdown","source":"## Step 4: Build your `DataLoaders`\n\nIn the first Kaggle competition I seriously competed in, we could never get our models performing as good as those in the top 20% of the public LB and my team dropped over 400 places when the private LB came out.  The problem was two-fold:\n\n1. We lost trust in our validation set and started fitting to the public LB, which in turn lead to overfitting on the public LB.\n2. We never really implemented a validation metric so that it precisely matched the stated competition metric.  \n\nWhat I learned from this experience, is that after you do some initial EDA, you focus on understanding and implementing the following two things BEFORE all else:\n\n1. A good validation set\n2. The competition metric as stated in the competition\n\n**IF** we had done that, we would have realized our preprocessing and training would have never put us into competitive shape. Instead, we were a bit lazy, assumed a bunch of things about the dataset, and blissfully started building models that never had a chance.  Having learned a good lesson through that brutal experience, we jumped *up* over 100 places on the private LB to finish in the top 7% in our following competition.\n\nSo ... let's not make that mistake with our baseline here :)","metadata":{}},{"cell_type":"markdown","source":"### 4a. Define a **good** validation set\n\nA \"good\" validation is going to require some EDA, and based on my cursory EDA work and others in this competition, it seems sensible to ensure our validation set does not share discourses from the same essays in the training set at minimum.  Therefore, we'll start with that.","metadata":{}},{"cell_type":"code","source":"essays = train_df[\"essay_id\"].unique()\nnp.random.seed(random_seed)\nnp.random.shuffle(essays)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:50.573282Z","iopub.execute_input":"2022-07-06T22:07:50.573654Z","iopub.status.idle":"2022-07-06T22:07:50.581585Z","shell.execute_reply.started":"2022-07-06T22:07:50.573615Z","shell.execute_reply":"2022-07-06T22:07:50.580674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_prop = 0.25\nval_sz = int(len(essays) * val_prop)\nval_essays = essays[:val_sz]","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:50.583024Z","iopub.execute_input":"2022-07-06T22:07:50.583627Z","iopub.status.idle":"2022-07-06T22:07:50.591879Z","shell.execute_reply.started":"2022-07-06T22:07:50.583586Z","shell.execute_reply":"2022-07-06T22:07:50.590779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"is_val = np.isin(train_df[\"essay_id\"], val_essays)\n\nidxs = np.arange(len(train_df))\nval_idxs = idxs[ is_val]\ntrn_idxs = idxs[~is_val]\n\nlen(trn_idxs), len(val_idxs)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:50.593270Z","iopub.execute_input":"2022-07-06T22:07:50.593674Z","iopub.status.idle":"2022-07-06T22:07:52.375084Z","shell.execute_reply.started":"2022-07-06T22:07:50.593637Z","shell.execute_reply":"2022-07-06T22:07:52.373799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**TIP**: How do you define a \"good\" validation set? Here's a couple of resources for you to consider:\n\n1. [How to create good validation and test sets](https://ohmeow.com/posts/2020/11/06/ajtfb-chapter-1.html#How-to-create-good-validation-and-test-sets) \n2. My blog post on lessons learned from chapter 1 of the fastbook, [\"Deep Learning for Coders with fastai & PyTorch\"](https://github.com/fastai/fastbook).\n3. Get some quality EDA in ya (see below)","metadata":{}},{"cell_type":"markdown","source":"Let's check that our target variable, `discourse_effectiveness`, distribution in both our training and validation sets is roughly proportional.","metadata":{}},{"cell_type":"code","source":"train_df.iloc[trn_idxs][\"discourse_effectiveness\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:52.376809Z","iopub.execute_input":"2022-07-06T22:07:52.377265Z","iopub.status.idle":"2022-07-06T22:07:52.394817Z","shell.execute_reply.started":"2022-07-06T22:07:52.377226Z","shell.execute_reply":"2022-07-06T22:07:52.393971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.iloc[val_idxs][\"discourse_effectiveness\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:52.397891Z","iopub.execute_input":"2022-07-06T22:07:52.398147Z","iopub.status.idle":"2022-07-06T22:07:52.410317Z","shell.execute_reply.started":"2022-07-06T22:07:52.398122Z","shell.execute_reply":"2022-07-06T22:07:52.409238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 4b. Define your `DataBlock`","metadata":{}},{"cell_type":"markdown","source":"From my experience, when working on NLP classification tasks, how you structure your text inputs is more of an art form than a science. An art form that shouldn't be overthought early but may well deserve some exploration later down the line.  \n\nAs such, it seems reasonable to include at least the `discourse_type` along with  `discourse_text` as it will likely provide some helpful signal given that the target variable is predicated on it.","metadata":{}},{"cell_type":"code","source":"def build_inputs(example):\n    return f'{example[\"discourse_type\"]}{sep}{example[\"discourse_text\"]}'","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:52.412337Z","iopub.execute_input":"2022-07-06T22:07:52.413246Z","iopub.status.idle":"2022-07-06T22:07:52.418307Z","shell.execute_reply.started":"2022-07-06T22:07:52.413207Z","shell.execute_reply":"2022-07-06T22:07:52.417459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"set_seed(random_seed)\nblocks = (TextBlock(hf_arch, hf_config, hf_tokenizer, hf_model), CategoryBlock)\ndblock = DataBlock(blocks=blocks, get_x=build_inputs, get_y=ColReader(\"discourse_effectiveness\"), splitter=IndexSplitter(val_idxs))","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:52.419506Z","iopub.execute_input":"2022-07-06T22:07:52.419777Z","iopub.status.idle":"2022-07-06T22:07:52.441507Z","shell.execute_reply.started":"2022-07-06T22:07:52.419752Z","shell.execute_reply":"2022-07-06T22:07:52.440556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Note that the only thing we really had to do to turn this from a regression task into a classification task was switch out our target's `TransformBlock` to be a `CategoryBlock` instead of a `RegressionBlock`.  This is one of the reasons I generally like working in fast.ai's mid-level API ... small changes = faster iterations.","metadata":{}},{"cell_type":"markdown","source":"### 4c. Build your `DataLoaders`","metadata":{}},{"cell_type":"code","source":"batch_size = 8\ndls = dblock.dataloaders(train_df, bs=batch_size)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:07:52.443630Z","iopub.execute_input":"2022-07-06T22:07:52.444468Z","iopub.status.idle":"2022-07-06T22:08:32.367945Z","shell.execute_reply.started":"2022-07-06T22:07:52.444420Z","shell.execute_reply":"2022-07-06T22:08:32.366884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's verify the classes we want to predict look right","metadata":{}},{"cell_type":"code","source":"dls.vocab","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:08:32.369400Z","iopub.execute_input":"2022-07-06T22:08:32.369776Z","iopub.status.idle":"2022-07-06T22:08:32.376671Z","shell.execute_reply.started":"2022-07-06T22:08:32.369737Z","shell.execute_reply":"2022-07-06T22:08:32.375659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's verify our our minibatches look as expected as well.","metadata":{}},{"cell_type":"code","source":"dls.show_batch(dataloaders=dls, max_n=2, trunc_at=500)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:08:32.378070Z","iopub.execute_input":"2022-07-06T22:08:32.378642Z","iopub.status.idle":"2022-07-06T22:08:32.504357Z","shell.execute_reply.started":"2022-07-06T22:08:32.378605Z","shell.execute_reply":"2022-07-06T22:08:32.503417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"b = dls.one_batch()\nlen(b), len(b[0][\"input_ids\"]), b[0][\"input_ids\"].shape, len(b[1])","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:08:32.505745Z","iopub.execute_input":"2022-07-06T22:08:32.506371Z","iopub.status.idle":"2022-07-06T22:08:32.564563Z","shell.execute_reply.started":"2022-07-06T22:08:32.506328Z","shell.execute_reply":"2022-07-06T22:08:32.563558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**TIP**: It's almost always a good idea to really look at a mini-batch of data before moving on. Often, this step will be key in troubleshooting issues discovered during training.","metadata":{}},{"cell_type":"code","source":"hf_tokenizer.decode(b[0][\"input_ids\"][0])","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:08:32.570571Z","iopub.execute_input":"2022-07-06T22:08:32.570886Z","iopub.status.idle":"2022-07-06T22:08:32.582380Z","shell.execute_reply.started":"2022-07-06T22:08:32.570860Z","shell.execute_reply":"2022-07-06T22:08:32.581514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**TIP**: When working with transformers, it's always wise to ensure your tokenized text looks right.","metadata":{}},{"cell_type":"code","source":"b[0]","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:08:32.583618Z","iopub.execute_input":"2022-07-06T22:08:32.584622Z","iopub.status.idle":"2022-07-06T22:08:32.606436Z","shell.execute_reply.started":"2022-07-06T22:08:32.584583Z","shell.execute_reply":"2022-07-06T22:08:32.605479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 5: Train\n\nRemember the second most important thing after a good validation set? \n\nThat's right ... making sure we use the same metric for evaluation as defined by the competition.  We'll take care of that during our construction of our fast.ai `Learner.`\n\nThe competition metric is stated as **muticlass logarithmic loss**. Given this, it makes sense to structure this as a multiclass classification problem and use `cross-entropy loss` as both our loss function and evaluation metric.  We can add in other metrics such as `error_rate` as well which may in turn be helpful in further understanding the underlying training data and improving performance.","metadata":{}},{"cell_type":"code","source":"weight_decay = 0.01\nset_seed(random_seed)\n\nmodel = BaseModelWrapper(hf_model)\n\nlearn = Learner(\n    dls,\n    model,\n    opt_func=partial(Adam, wd=weight_decay),\n    loss_func=PreCalculatedCrossEntropyLoss(),\n    metrics=[error_rate],\n    cbs=[BaseModelCallback],\n    splitter=blurr_splitter\n)\n\nlearn.create_opt()\nlearn = learn.to_fp16()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:08:32.607766Z","iopub.execute_input":"2022-07-06T22:08:32.608124Z","iopub.status.idle":"2022-07-06T22:08:32.779037Z","shell.execute_reply.started":"2022-07-06T22:08:32.608086Z","shell.execute_reply":"2022-07-06T22:08:32.777991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Running `Learner.create_opt()`","metadata":{}},{"cell_type":"code","source":"print(len(learn.opt.param_groups))","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:08:32.780667Z","iopub.execute_input":"2022-07-06T22:08:32.781042Z","iopub.status.idle":"2022-07-06T22:08:32.787209Z","shell.execute_reply.started":"2022-07-06T22:08:32.781005Z","shell.execute_reply":"2022-07-06T22:08:32.785934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.lr_find(suggest_funcs=[minimum, steep, valley, slide])","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:08:32.788782Z","iopub.execute_input":"2022-07-06T22:08:32.789705Z","iopub.status.idle":"2022-07-06T22:08:43.830417Z","shell.execute_reply.started":"2022-07-06T22:08:32.789665Z","shell.execute_reply":"2022-07-06T22:08:43.829458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**TIP**: For deberta classification/regression models, I find that setting my minimum LR equal to the middle of the downward slope / 10 or 100 and setting my maximum LR equal to that same point x 5, 10, or 100 ... generally leads to good results. Of course, this doesn't always work.\n\n**TIP**: Using **1cycle** learning allows us to be a bit more aggressive with our learning rates.  Want to learn more? Check out the [\"A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay\" paper](https://arxiv.org/abs/1803.09820).\n","metadata":{}},{"cell_type":"code","source":"set_seed(random_seed)\nlearn.fit_one_cycle(4, lr_max=slice(1e-7, 1e-3))","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:08:43.832203Z","iopub.execute_input":"2022-07-06T22:08:43.832598Z","iopub.status.idle":"2022-07-06T22:32:41.802108Z","shell.execute_reply.started":"2022-07-06T22:08:43.832558Z","shell.execute_reply":"2022-07-06T22:32:41.801060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.recorder.values","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:32:41.803886Z","iopub.execute_input":"2022-07-06T22:32:41.804618Z","iopub.status.idle":"2022-07-06T22:32:41.813374Z","shell.execute_reply.started":"2022-07-06T22:32:41.804574Z","shell.execute_reply":"2022-07-06T22:32:41.812437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"probs, targs = learn.get_preds()\n\nprobs.shape, targs.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:32:41.815051Z","iopub.execute_input":"2022-07-06T22:32:41.815668Z","iopub.status.idle":"2022-07-06T22:33:14.351046Z","shell.execute_reply.started":"2022-07-06T22:32:41.815628Z","shell.execute_reply":"2022-07-06T22:33:14.349605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 6: Inference","metadata":{}},{"cell_type":"markdown","source":"Inference is pretty easy with fastai. `Learner.dls.test_dl()` ensures that the same item/batch transforms defined in our `DataBlock` above are applied to our test set.  `Learner.get_preds()` will return the probabilities for our three classes.","metadata":{}},{"cell_type":"code","source":"test_dl = learn.dls.test_dl(test_df)\ntest_dl = test_dl.to(device)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:33:14.353023Z","iopub.execute_input":"2022-07-06T22:33:14.353614Z","iopub.status.idle":"2022-07-06T22:33:14.371596Z","shell.execute_reply.started":"2022-07-06T22:33:14.353572Z","shell.execute_reply":"2022-07-06T22:33:14.370392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"probs, _ = learn.get_preds(dl=test_dl)\nprobs = probs.numpy()\n\nprint(probs.shape)\nprint(probs)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:33:14.373650Z","iopub.execute_input":"2022-07-06T22:33:14.373982Z","iopub.status.idle":"2022-07-06T22:33:14.587700Z","shell.execute_reply.started":"2022-07-06T22:33:14.373954Z","shell.execute_reply":"2022-07-06T22:33:14.586175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 7: Submission","metadata":{}},{"cell_type":"markdown","source":"We can use `dls.vocab` to align our class probabilities above with the actual classes.  This will allow us to put this in a `DataFrame` we can horizontally concatenate with `test_df` to get the `discourse_id` that we need to include alongside the class probabilities.  From there, we need to simply save our `submission_df` with the required fields to a .csv ... and we're done.","metadata":{}},{"cell_type":"code","source":"dls.vocab","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:33:14.589764Z","iopub.execute_input":"2022-07-06T22:33:14.590181Z","iopub.status.idle":"2022-07-06T22:33:14.600078Z","shell.execute_reply.started":"2022-07-06T22:33:14.590132Z","shell.execute_reply":"2022-07-06T22:33:14.598741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df = pd.concat([test_df, pd.DataFrame(data=probs, columns=dls.vocab)], axis=1)\nsubmission_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:33:14.601484Z","iopub.execute_input":"2022-07-06T22:33:14.602092Z","iopub.status.idle":"2022-07-06T22:33:14.622200Z","shell.execute_reply.started":"2022-07-06T22:33:14.602049Z","shell.execute_reply":"2022-07-06T22:33:14.621310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df[[\"discourse_id\"] + list(dls.vocab)].head()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:33:14.623453Z","iopub.execute_input":"2022-07-06T22:33:14.624030Z","iopub.status.idle":"2022-07-06T22:33:14.641771Z","shell.execute_reply.started":"2022-07-06T22:33:14.624001Z","shell.execute_reply":"2022-07-06T22:33:14.640852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df[[\"discourse_id\"] + list(dls.vocab)].to_csv('submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T22:33:14.643137Z","iopub.execute_input":"2022-07-06T22:33:14.643398Z","iopub.status.idle":"2022-07-06T22:33:14.656666Z","shell.execute_reply.started":"2022-07-06T22:33:14.643373Z","shell.execute_reply":"2022-07-06T22:33:14.655450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Resources and Things to try\n\nAs promised, here are some competition notebooks and discussion threads I'm finding helpful during EDA:\n\n1. [Feedback Prize Effectiveness EDA+DeBERTa baseline](https://www.kaggle.com/code/tanlikesmath/feedback-prize-effectiveness-eda-deberta-baseline)\n2. [Feedback Prize - 📊📉 The Complete Overview](https://www.kaggle.com/code/lextoumbourou/feedback-prize-the-complete-overview)\n3. [📊EDA - FeedBack Prize⭐⭐⭐](https://www.kaggle.com/code/vivek61/eda-feedback-prize)\n4. [Feedback: Inspired EDA](https://www.kaggle.com/code/foolishboi/feedback-inspired-eda)\n5. [This is why you do EDA](https://www.kaggle.com/competitions/feedback-prize-effectiveness/discussion/331544)\n\n\nThings to try:\n- Review similar competition winner's strategies. [Here](https://www.kaggle.com/competitions/feedback-prize-effectiveness/discussion/333063), [here](https://www.kaggle.com/competitions/feedback-prize-effectiveness/discussion/326998) and [here](https://www.kaggle.com/competitions/feedback-prize-effectiveness/discussion/327056) are a few.\n- Spend some EDA time cleaning up the training set (e.g., those essays with only a single discourse look a bit sus) and feature engineering.\n- Experiment with other smaller models that typically work well for clssification tasks (e.g., roberta, deberta, and bart are my go tos)\n- Experiment with some bigger versions of those smaller models that worked well for you.\n- Be creative with your inputs; you can improve your results by adding special or regular tokens and/or structuring your inputs differently\n- Try using K-Fold or Stratified K-Fold cross validation and ensemble your results (see Jeremy's notebook for more info)\n- Once you have a decent set of hyperparameters working for you, you can use an optimization framework like Optuna and/or Weights & Biases to fine-tune your choices.\n- Read the papers related to the architectures you are using. Often you'll find recommended hyperparameter values and other important recommendations to training them well.","metadata":{}},{"cell_type":"markdown","source":"## In conclusion\n\nI hope you've learned a little bit about training transformers with blurr, and may even be encouraged to give it a go on Kaggle or at work. If you enjoyed this notebook, **I would greatly appreciate an upvote**. 🙏\n\nPlease use the comments section below to ask any questions or share insights you may have on using blurr, fastai, and the transformers library to effectively train transformer models. As you can see, with fastai and blurr, it requires less than 3 or so changes to turn a regression task into a classification task ... and that is helpful when you want to iterate quickly.\n\nFor folks new to working with the Hugging Face transformers library with a particular interest in using fastai and blurr, I heartily recommend the study group hosted by Weights&Biases that I've been leading for the past few months.  You can watch the entire playlist [here](https://www.youtube.com/playlist?list=PLD80i8An1OEF8UOb9N9uSoidOGIMKW96t).\n\n\nAnd if you made it this far, thanks for reading all the way to the end :)\n\nYou can find me on tweeting at [@waydegilliam](https://twitter.com/waydegilliam) and blogging ML/Software development at [ohmeow.com](https://ohmeow.com/)","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}