{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### melanoma detection using FastAI - prototype\nCaitlin Ortega Ruble\n\nThis notebook is going to serve as the prototype for training the deep learning model for the [siim-isic-melanoma-classification challenge on Kaggle.](https://www.kaggle.com/competitions/siim-isic-melanoma-classification/) \n\nGoals:\n1. Recreate the cleaned data frame of training image metadata, as explored in an earlier Data Wrangling and EDA notebook [here]( https://www.kaggle.com/code/caitlinruble/melanoma-classification-eda)\n2. Create a subset of negative and positive images from the training set to serve as my \"proof of concept\" set as I set up the modeling process with FastAI\n3. Load, train and validate the miniature data set using FastAI. Basically, go through the whole process I will eventually with the full dataset, except only with a small subset of data to save on computing power and time while I'm tinkering.\n4. Export the best model for deployment prototyping.\n\n","metadata":{"execution":{"iopub.status.busy":"2022-09-19T22:54:28.584787Z","iopub.execute_input":"2022-09-19T22:54:28.585616Z","iopub.status.idle":"2022-09-19T22:54:40.125175Z","shell.execute_reply.started":"2022-09-19T22:54:28.585544Z","shell.execute_reply":"2022-09-19T22:54:40.123667Z"}}},{"cell_type":"markdown","source":"## Step 1: Load and clean train.csv as a Pandas DataFrame","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport os\nimport imageio.v2 as imageio\n\nsns.set_style('darkgrid')\nplt.style.use('seaborn-notebook')","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:45:20.938797Z","iopub.execute_input":"2022-10-11T19:45:20.939268Z","iopub.status.idle":"2022-10-11T19:45:21.918297Z","shell.execute_reply.started":"2022-10-11T19:45:20.939202Z","shell.execute_reply":"2022-10-11T19:45:21.917199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#read in train csv\ntrain = pd.read_csv('../input/siim-isic-melanoma-classification/train.csv')\n\n#check for any missing values:\nfor col in train.columns:\n    print(col + ' missing values: ' + str(train[col].isna().sum()))","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:45:21.924665Z","iopub.execute_input":"2022-10-11T19:45:21.927152Z","iopub.status.idle":"2022-10-11T19:45:22.043513Z","shell.execute_reply.started":"2022-10-11T19:45:21.927112Z","shell.execute_reply":"2022-10-11T19:45:22.042518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code block executes the data cleaning steps which were found to be necessary in the original data wrangling and cleaning notebook. The rationale behind each choice is fully explored and explained there.","metadata":{}},{"cell_type":"code","source":"#fill missing anatom site values with \"unknown or other\"\ntrain.anatom_site_general_challenge.fillna('other or unknown', inplace=True)\n\n#drop all rows from the train dataframe that are associated with 3 patients who are missing some data.\ntrain.drop(train[(train['patient_id'] == 'IP_0550106') | (train['patient_id'] == 'IP_5205991') |(train['patient_id'] == 'IP_9835712')].index, inplace=True)\n\n#check for any more missing values:\nfor col in train.columns:\n    print(col + ' missing values: ' + str(train[col].isna().sum()))","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:45:22.048015Z","iopub.execute_input":"2022-10-11T19:45:22.050273Z","iopub.status.idle":"2022-10-11T19:45:22.100582Z","shell.execute_reply.started":"2022-10-11T19:45:22.050235Z","shell.execute_reply":"2022-10-11T19:45:22.099553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#save cleaned train df as data\n\ndata = train","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:45:22.106205Z","iopub.execute_input":"2022-10-11T19:45:22.108736Z","iopub.status.idle":"2022-10-11T19:45:22.115216Z","shell.execute_reply.started":"2022-10-11T19:45:22.108700Z","shell.execute_reply":"2022-10-11T19:45:22.113986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 2: Create a subset of images that are positive and negative for the target class","metadata":{}},{"cell_type":"code","source":"#check percentage of positive images in entire training data set:\n\nn_pos = len(data[data['target']==1])\nprint('number of images in full train set: {}'.format(len(data)))\nprint('Number of positive images in subset of train set: {}'.format(n_pos))\nprint('Percentage of positive images in subset train set: {:.1%}'.format(n_pos/len(data)))\nprint('Number of negative images in subset train set: {}'.format(len(data) - n_pos))\nprint('Percentage of negative images in subset train set: {:.1%}'.format(1 - n_pos/len(data)))","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:45:22.120328Z","iopub.execute_input":"2022-10-11T19:45:22.122944Z","iopub.status.idle":"2022-10-11T19:45:22.136154Z","shell.execute_reply.started":"2022-10-11T19:45:22.122904Z","shell.execute_reply":"2022-10-11T19:45:22.135013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Proportionally sample out 2.5% of the images (1.8% positive, 98.2% negative):\n\nsub_data = data.groupby('target', group_keys=False).apply(lambda x: x.sample(frac=0.025))\nsub_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:45:22.140696Z","iopub.execute_input":"2022-10-11T19:45:22.143069Z","iopub.status.idle":"2022-10-11T19:45:22.185366Z","shell.execute_reply.started":"2022-10-11T19:45:22.143035Z","shell.execute_reply":"2022-10-11T19:45:22.183885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#check percentage of positive images in the data subset:\n\nn_pos = len(sub_data[sub_data['target']==1])\nprint('Number of images in subset of data: {}'.format(len(sub_data)))\nprint('Subset of data as a percentage of all image data: {:.1%}'.format(len(sub_data)/len(data)))\nprint('Number of positive images: {}'.format(n_pos))\nprint('Percentage of positive images in subset: {:.1%}'.format(n_pos/len(sub_data)))\nprint('Number of negative images: {}'.format(len(sub_data) - n_pos))\nprint('Percentage of negative images in subset: {:.1%}'.format(1 - n_pos/len(sub_data)))","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:45:22.189485Z","iopub.execute_input":"2022-10-11T19:45:22.191812Z","iopub.status.idle":"2022-10-11T19:45:22.205690Z","shell.execute_reply.started":"2022-10-11T19:45:22.191773Z","shell.execute_reply":"2022-10-11T19:45:22.204411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Excellent! Now I have a small subset of the original training data set, proportionally split by the target field to build and test my propotype deep learning model. ","metadata":{}},{"cell_type":"markdown","source":"## Part 3: Build deep-learning classifier prototype with fastAI","metadata":{}},{"cell_type":"code","source":"#hide output\n!pip install -Uqq fastbook --q --q\n!pip install ipywidgets\n\n#jupyter nbextension enable --py widgetsnbextension","metadata":{"_kg_hide-output":false,"_kg_hide-input":false,"scrolled":true,"execution":{"iopub.status.busy":"2022-10-11T19:45:22.210440Z","iopub.execute_input":"2022-10-11T19:45:22.212892Z","iopub.status.idle":"2022-10-11T19:45:53.946333Z","shell.execute_reply.started":"2022-10-11T19:45:22.212855Z","shell.execute_reply":"2022-10-11T19:45:53.944919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import fastbook\nfastbook.setup_book()\nimport fastai\n\nfrom fastbook import *\nfrom fastai.vision.all import *\nfrom fastai.tabular.all import *\nfrom fastai.medical.imaging import *\nimport torch\nfrom pathlib import Path\nfrom PIL import Image\nfrom tqdm import tqdm","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:45:53.948232Z","iopub.execute_input":"2022-10-11T19:45:53.949147Z","iopub.status.idle":"2022-10-11T19:45:56.593969Z","shell.execute_reply.started":"2022-10-11T19:45:53.949101Z","shell.execute_reply":"2022-10-11T19:45:56.592992Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#ensure GPU is running\ndevice = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")\ndevice","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:45:56.597863Z","iopub.execute_input":"2022-10-11T19:45:56.598361Z","iopub.status.idle":"2022-10-11T19:45:56.668474Z","shell.execute_reply.started":"2022-10-11T19:45:56.598330Z","shell.execute_reply":"2022-10-11T19:45:56.667521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#create dataloaders object of class ImageDataLoaders from the sub_data df\njpg_path = '../input/siim-isic-melanoma-classification/jpeg'\n\ndls = ImageDataLoaders.from_df(df = sub_data, #specify df holding image names\n                               path = jpg_path,    #set path for where to find images\n                               folder = 'train',   #specify looking specifically in the 'train' folder for this dls\n                               suff = '.jpg',      #add the .jpg suffix to file names from df\n                               label_col = 6,      #col 6 holds label info 'benign' or 'malignant'\n                               valid_pct = 0.2,    #20% of images will be held for validation\n                               bs = 8,            #set batch size\n                               device=device,      #set device\n                               item_tfms = Resize(128))    #resize images as they're loaded\n\ndls.show_batch()","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:45:56.670003Z","iopub.execute_input":"2022-10-11T19:45:56.670390Z","iopub.status.idle":"2022-10-11T19:46:03.792977Z","shell.execute_reply.started":"2022-10-11T19:45:56.670352Z","shell.execute_reply":"2022-10-11T19:46:03.792013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#create the vision learner\n\nlearn = vision_learner(dls,                  #specify dataloader object\n                       resnet18,             #specify a pre-trained model we want to build off of\n                       metrics=RocAucBinary, #specify metric we want to optimize\n                       model_dir = '/kaggle/working')   #specify output location to store model","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:46:03.794091Z","iopub.execute_input":"2022-10-11T19:46:03.794410Z","iopub.status.idle":"2022-10-11T19:46:08.679642Z","shell.execute_reply.started":"2022-10-11T19:46:03.794380Z","shell.execute_reply":"2022-10-11T19:46:08.678688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#This cell is taking *way* too long to run, and not utilizing the GPU. \n#A forum search shows that others have run into the same issue with this dataset\n#The CPU is being used to resize the images, which it turns out takes a long time\n#The solution: resize the images prior to feeding them to the data learner.\n\n\n#learn.lr_find()","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:46:08.684331Z","iopub.execute_input":"2022-10-11T19:46:08.686584Z","iopub.status.idle":"2022-10-11T19:46:08.692790Z","shell.execute_reply.started":"2022-10-11T19:46:08.686544Z","shell.execute_reply":"2022-10-11T19:46:08.691601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The previous cell is taking *way* too long to run, and not utilizing the GPU. The cell below shows that the GPU is mounted and running, so why isn't it being used when I run the lr_find() method on the learner?\n\nA forum search shows that others have run into the same issue with this dataset.\\\nThe CPU is being used to resize the images in the DataLoader, which it turns out takes a long time.\n\nThe solution: resize the images prior to feeding them to the DataLoader.","metadata":{}},{"cell_type":"code","source":"print(torch.__version__)\nprint(fastai.__version__)\nprint(torch.cuda.is_available())\nprint(torch.cuda.get_device_name(0))\nprint(torch.backends.cudnn.enabled)\nprint(dls.device)","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:46:08.694383Z","iopub.execute_input":"2022-10-11T19:46:08.695179Z","iopub.status.idle":"2022-10-11T19:46:08.712700Z","shell.execute_reply.started":"2022-10-11T19:46:08.695144Z","shell.execute_reply":"2022-10-11T19:46:08.711654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I want to resize all the images in ../jpeg to be a consistent size prior to feeding them into the data loader....how to do this?\n\nidea 1: move copies of the images into my kaggle working directory and resize them there\n\nidea 2: use a resized dataset someone else has already made (e.g. https://www.kaggle.com/datasets/sarques/siimisic-melanoma-1024jpeg) Note: tried these 1024x1024 files and running training the model ate up all my GPU capacity before even one epoch. Not a reasonable option in this environment. Instead, I'll try 224x224.\n\nidea 3: can I somehow resize the images directly before feeding them to dls?","metadata":{}},{"cell_type":"code","source":"#try method 2 first, to see if I can get this working at all\n#for this I actually ended up going with a 224 image size\n\npath_224 = '../input/siic-isic-224x224-images'\n\ndls2 = ImageDataLoaders.from_df(df = sub_data, #specify df holding image names\n                               path = path_224,    #set path for where to find images\n                               folder = 'train',   #specify looking specifically in the 'train' folder for this dls\n                               suff = '.png',      #add the .png suffix to file names from df\n                               label_col = 7,      #col 6 holds label info 0 or 1\n                                seed = 42,\n                                valid_pct = 0.2,\n                               splitter=TrainTestSplitter (\n                                   test_size=0.2,   #20% of images will be held for validation\n                                   random_state=42, #set seed\n                                   stratify=True,   #ensure a proportional split of classes in the train/test split\n                                   shuffle=False),  #no indication that we should shuffle the data\n                               bs = 8,            #set batch size\n                               device=device)      #set device\n                              # item_tfms = Resize(128))    #resize images as they're loaded\n\ndls2.show_batch()","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:46:08.714416Z","iopub.execute_input":"2022-10-11T19:46:08.715108Z","iopub.status.idle":"2022-10-11T19:46:09.772254Z","shell.execute_reply.started":"2022-10-11T19:46:08.715075Z","shell.execute_reply":"2022-10-11T19:46:09.771394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now I can instantiate my first learner. [This resource](https://learnopencv.com/transfer-learning-for-medical-images/) on Transfer Learning for Medical Images from LearnOpenCV.com asserts that VGGNet pretrained models perform well for skin images in the medical context, so I'm going to leverage the vgg16_bn (batch normalized) pretrained model for my first pass.\n\nI'm using the roc_auc_score metric because that is the basis the competition is scored on. It would also be very reasonable to use model recall as the primary metric, because we value avoiding false negatives when it comes to detecting deadly diseases. It it better to flag some lesions as malignant when they're not than to fail to detect lesions that are truly malignant. To be able to fully define the metrics of my model, I'm going to track roc_auc_score, error_rate (%of test samples incorrectly classified), and recall (TP/TP+FN, i.e. percentage of positive cases being accurately classified)","metadata":{}},{"cell_type":"code","source":"#instantiate the rocaucbinary metric\nrocAucBinary = RocAucBinary()\nrecall = Recall()\n\n#create the vision learner\n\nlearn2 = vision_learner(dls2,                  #specify dataloader object\n                       models.vgg16_bn,             #specify a pre-trained model we want to build off of\n                       metrics=[rocAucBinary, error_rate, recall], #specify metrics we want to see\n                       model_dir = '/kaggle/working')   #specify output location to store model","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:46:09.773924Z","iopub.execute_input":"2022-10-11T19:46:09.774562Z","iopub.status.idle":"2022-10-11T19:46:45.617564Z","shell.execute_reply.started":"2022-10-11T19:46:09.774510Z","shell.execute_reply":"2022-10-11T19:46:45.616603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A good first step before training the model is to gauge an appropriate learning rate to begin with. FastAI has a great built-in utility for this: lr_find. This function will default to using the \"valley\" algorithm as the suggested hyperparameter, but others are included as well. The topic of choosing the learning rate for a deep learning model is complex, with many theories existing. Ultimately, my impression as a novice is that it kind of comes down to being an art. Setting the learning rate too high risks missing out on finding the true gradient descent, setting the learning rate lower increases the processing time. There's a sweet spot in the middle.\n\nHere are what the different suggestion functions are for FastAI's .lr_find():\n\n   - [Valley](https://docs.fast.ai/callback.schedule.html#valley): Suggests a learning rate from the longest valley and returns its index. The valley algorithm was developed by ESRI and takes the steepest slope roughly 2/3 through the longest valley in the LR plot, and is also the default for Learner.lr_find\n    \n   - [Slide](https://docs.fast.ai/callback.schedule.html#slide): Suggests a learning rate following an interval slide rule and returns its index. The slide rule is an algorithm developed by Andrew Chang out of Novetta.\n   \n   - [Minimum](https://docs.fast.ai/callback.schedule.html#minimum): Suggests a learning rate one-tenth the minumum before divergance and returns its index\n   \n   - [Steep](https://docs.fast.ai/callback.schedule.html#steep): Suggests a learning rate when the slope is the steepest and returns its index","metadata":{}},{"cell_type":"code","source":"minimum, steep, valley, slide = learn2.lr_find(suggest_funcs=(minimum, steep, valley, slide))\nprint(f\"Minimum/10:\\t{minimum:.2e}\\nSteepest point:\\t{steep:.2e}\\nLongest valley:\\t{valley:.2e}\\nSlide interval:\\t{slide:.2e}\")","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:46:45.618886Z","iopub.execute_input":"2022-10-11T19:46:45.619599Z","iopub.status.idle":"2022-10-11T19:47:02.618075Z","shell.execute_reply.started":"2022-10-11T19:46:45.619560Z","shell.execute_reply":"2022-10-11T19:47:02.616919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"My artistic sensibilities are driving me to **use the learning rate found by the \"minimum\"** in this case. \n\n**Now I am ready to fine-tune the VGG16 pre-trained model.** FastAI's .fine_tune() method automates a lot of the steps of freezing/unfreezing layers, changing weights, and updating learning rates in the last few layers of the pre-trained model, leading to great results in most cases.\n\nAfter some experimentation, **I've opted to use the SaveModelCallback() function** in the fine-tuning callbacks. What this does is evaluate the model performance against a set metric after each epoch. When there is an improvement in performance, the model is saved to the working directory. Using this method, I can ensure that the best epoch in model training is the one that is saved.","metadata":{}},{"cell_type":"code","source":"#set up a 10 epoch fine-tuning run\n\nlearn2.fine_tune(10,                               #set number of epochs \n                 base_lr=minimum,                  #set the initial learning rate\n                 cbs=SaveModelCallback(            #use the SaveModelCallback to save the best model\n                     monitor='roc_auc_score',      #set the roc_auc_score as the montitored metric\n                     fname = 'vgg16_subset',       #choose the name the best model will be saved under\n                     comp = np.greater,            #specify that when the roc_auc_score increases, that's considered better\n                     with_opt=True))               #saves optimizer state, if available, when saving model","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:49:47.298504Z","iopub.execute_input":"2022-10-11T19:49:47.299153Z","iopub.status.idle":"2022-10-11T19:52:17.481846Z","shell.execute_reply.started":"2022-10-11T19:49:47.299106Z","shell.execute_reply":"2022-10-11T19:52:17.480624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For the vgg16_bn pre-trained model, fine-tuning over 10 epochs leads to a max roc_auc_score of 0.75 after epoch 6. The best model has been saved as \"vgg16_subset.pth\" in the kaggle working directory.\nThe error rate is basically stable at only 3%, but the recall score is consistently 0, indicating that the model isn't correctly classifying the positive cases in the validation set. This is good starting place, but I definitely want to improve that performance.\n\nThe very top teams for this competition were able to achieve roc_auc scores of 0.95 in their submission.\n\nWhat can I tinker with?\n\n    1. which pre-trained model I start with\n    2. batch size\n    3. learning rate\n    4. the proportion and quantity of malignant images the model is allowed to train on\n    5. Preprocessing steps with the images; I could try transformations like random croping, rotation, using grayscale, etc.\n    6. The loss function used in fine-tuning\n   \n**First, I'm going to jump over to ResNet34 and see how that does.**\nResNet is a versatile pretrained image model that applies well to a variety of use-cases. It's worth a try to see how it compares to VGG16","metadata":{}},{"cell_type":"code","source":"#leverage the resnet34 pretrained CNN model\nlearn3 = vision_learner(dls2,                  #specify dataloader object\n                       models.resnet34,             #specify a pre-trained model we want to build off of\n                       metrics=[rocAucBinary, error_rate, recall], #specify metric we want to optimize\n                       model_dir = '/kaggle/working')   #specify output location to store model","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:55:30.819367Z","iopub.execute_input":"2022-10-11T19:55:30.819885Z","iopub.status.idle":"2022-10-11T19:55:31.578328Z","shell.execute_reply.started":"2022-10-11T19:55:30.819837Z","shell.execute_reply":"2022-10-11T19:55:31.577260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn3.lr_find(suggest_funcs=(minimum, steep, valley, slide))","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:58:03.866272Z","iopub.execute_input":"2022-10-11T19:58:03.866867Z","iopub.status.idle":"2022-10-11T19:58:12.541394Z","shell.execute_reply.started":"2022-10-11T19:58:03.866823Z","shell.execute_reply":"2022-10-11T19:58:12.538523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#use the valley method to set the base_lr: 0.04786301031708717\nlearn3.fine_tune(10, \n                 base_lr=0.04786301031708717, \n                 cbs=SaveModelCallback(\n                     monitor='roc_auc_score', \n                     fname = 'resnet34_subset', \n                     comp = np.greater, \n                     with_opt=True))","metadata":{"execution":{"iopub.status.busy":"2022-10-11T19:58:26.105665Z","iopub.execute_input":"2022-10-11T19:58:26.106295Z","iopub.status.idle":"2022-10-11T20:00:09.275757Z","shell.execute_reply.started":"2022-10-11T19:58:26.106246Z","shell.execute_reply":"2022-10-11T20:00:09.274509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The highest roc_auc_score for the resnet 34 based-model reached 0.775 after epoch 9. The fine-tuned model has an error rate of just 2.4%, however only scores a recall of 0.2. The epochs trained more quickly, in under 10s for each epoch v. the 13s per epoch when using the vgg16_bn model. While fine-tuning the ResNet34 model, the validation loss was really high and \"all over the place\" for epochs 5-8, but then drastically improved for the final epoch, 9.\n\n**Conclusion:** resnet34 is achieving higher performance with this particular training and validation set (reminder: this is only 2.5% of the overall training data, and only being validated on 0.5% of the overall training data), and does so with lower computational cost. **It's my best model, for now**","metadata":{}},{"cell_type":"markdown","source":"## Part IV: Use the best model to make predictions for the submission.\n\nHere are the submission guidelines for the Kaggle competition:\n\nSubmissions are evaluated on area under the ROC curve between the predicted probability and the observed target.\n\nSubmission File\nFor each image_name in the test set, you must predict the probability (target) that the sample is malignant. The file should contain a header and have the following format:\n\nimage_name,target\n\nISIC_0052060,0.7\n\nISIC_0052349,0.9\n\nISIC_0058510,0.8\n\nISIC_0073313,0.5\n\nISIC_0073502,0.5\n\netc.\n","metadata":{}},{"cell_type":"code","source":"#load the resnet34 learner as \"best_model\"\nbest_model = learn3.load('resnet34_subset', #load model from working directory\n                         device=device,     #ensure the loaded model uses the active cuda:0 device\n                         with_opt=True,     #load optimizer state\n                         strict=True)       #the file must exactly contain weights for every parameter key in model\n\n#read in test csv\ntest = pd.read_csv('../input/siim-isic-melanoma-classification/test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-10-11T20:02:22.738593Z","iopub.execute_input":"2022-10-11T20:02:22.739885Z","iopub.status.idle":"2022-10-11T20:02:23.037838Z","shell.execute_reply.started":"2022-10-11T20:02:22.739833Z","shell.execute_reply":"2022-10-11T20:02:23.036769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Initialize a list of image names containing all the images from test\ntest_ims = test['image_name'].to_list()\nprint(f\"number of test images: {len(test_ims)}\")\n\n#initialize a dict to hold results\npreds = {}\n","metadata":{"execution":{"iopub.status.busy":"2022-10-11T20:02:28.085732Z","iopub.execute_input":"2022-10-11T20:02:28.086201Z","iopub.status.idle":"2022-10-11T20:02:28.099625Z","shell.execute_reply.started":"2022-10-11T20:02:28.086162Z","shell.execute_reply":"2022-10-11T20:02:28.098571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#test_path\ntest_path = '../input/siic-isic-224x224-images/test/'\n\ntest_dl = dls2.test_dl(get_image_files(test_path))\n\npredictions = best_model.get_preds(dl=test_dl, with_preds=True)","metadata":{"execution":{"iopub.status.busy":"2022-10-11T20:50:21.579248Z","iopub.execute_input":"2022-10-11T20:50:21.579790Z","iopub.status.idle":"2022-10-11T20:51:11.031165Z","shell.execute_reply.started":"2022-10-11T20:50:21.579740Z","shell.execute_reply":"2022-10-11T20:51:11.029880Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions","metadata":{"execution":{"iopub.status.busy":"2022-10-11T20:52:54.459891Z","iopub.execute_input":"2022-10-11T20:52:54.460413Z","iopub.status.idle":"2022-10-11T20:52:54.472571Z","shell.execute_reply.started":"2022-10-11T20:52:54.460371Z","shell.execute_reply":"2022-10-11T20:52:54.471114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#this would work if we were just predicting on 1 image\n#categories = ('Benign', 'Malignant')\n\n#def classify_image(img):\n#    pred, idx, probs = best_model.predict(img)\n#    return dict(zip(categories, map(float,probs)))","metadata":{"execution":{"iopub.status.busy":"2022-10-11T20:59:43.860443Z","iopub.execute_input":"2022-10-11T20:59:43.860959Z","iopub.status.idle":"2022-10-11T20:59:43.868052Z","shell.execute_reply.started":"2022-10-11T20:59:43.860919Z","shell.execute_reply":"2022-10-11T20:59:43.866934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"#initialize a dict to hold results\npreds = {}\n\n\n#update keys in dict to hold likelihood\nfor i in range(len(test_ims)):\n    preds[test_ims[i]] = np.round(float(predictions[0][i][1]),1) #gives the likelihood of an image being malignant","metadata":{"execution":{"iopub.status.busy":"2022-10-11T21:06:37.714637Z","iopub.execute_input":"2022-10-11T21:06:37.715163Z","iopub.status.idle":"2022-10-11T21:06:38.970511Z","shell.execute_reply.started":"2022-10-11T21:06:37.715119Z","shell.execute_reply":"2022-10-11T21:06:38.969459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame.from_dict(preds, orient='index')\nsubmission.reset_index(inplace=True)\nsubmission.rename(columns={'index':'image_name',0:'target'},inplace=True)\n\nsubmission.to_csv(\"submission.csv\", index=False)\nsubmission.head()\n","metadata":{"execution":{"iopub.status.busy":"2022-10-11T21:10:39.958971Z","iopub.execute_input":"2022-10-11T21:10:39.959417Z","iopub.status.idle":"2022-10-11T21:10:40.009682Z","shell.execute_reply.started":"2022-10-11T21:10:39.959378Z","shell.execute_reply":"2022-10-11T21:10:40.008739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#check to make sure output looks like I want it to. Looks good.\npd.read_csv('./submission.csv').head()","metadata":{"execution":{"iopub.status.busy":"2022-10-11T21:10:47.686569Z","iopub.execute_input":"2022-10-11T21:10:47.687036Z","iopub.status.idle":"2022-10-11T21:10:47.715331Z","shell.execute_reply.started":"2022-10-11T21:10:47.686996Z","shell.execute_reply":"2022-10-11T21:10:47.714394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Submission scores\nPrivate: .5081\nPublic: .4928","metadata":{"execution":{"iopub.status.busy":"2022-10-11T21:21:14.847278Z","iopub.execute_input":"2022-10-11T21:21:14.847870Z","iopub.status.idle":"2022-10-11T21:21:14.879737Z","shell.execute_reply.started":"2022-10-11T21:21:14.847822Z","shell.execute_reply":"2022-10-11T21:21:14.877223Z"}}},{"cell_type":"markdown","source":"## Part V: export the best model as a pkl file","metadata":{}},{"cell_type":"code","source":"#this exports the model to the kaggle working directory. \n#From there I have to download it to my local device so I can use it for deployment.\nbest_model.export('/kaggle/working/resnet34_subsample_proportional.pkl')","metadata":{"execution":{"iopub.status.busy":"2022-10-11T20:02:55.158686Z","iopub.execute_input":"2022-10-11T20:02:55.159135Z","iopub.status.idle":"2022-10-11T20:02:55.496114Z","shell.execute_reply.started":"2022-10-11T20:02:55.159096Z","shell.execute_reply":"2022-10-11T20:02:55.495033Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Part 6: Next steps\n\nNow that I have a working prototype that is performing reasonably well on a small subset of the data, I have a couple of avenues to work on next.\n    \n    1. Set up a site to deploy the model on. I'm going to follow the instructions given by Jeremy Howard in [Practical Deep Learning for Coders, Lesson 2: Deployment](https://course.fast.ai/Lessons/lesson2.html)\n    \n    2. Build a working model on a larger set of the data. Because I'm not building from scratch, it may make more sense to get a balanced sample of the ~500 malignant images and an equal number of benign images. I'm not sure yet, and I do think this will take some experimentation.\n    \n    3. Create my own image dataset with resized images and upload to kaggle. This will give me greater ownership of the project from start to finish, as I'll be writing the Python code to resize the images on my local machine. ","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}