{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"# SIIM-ISIC Melanoma Classification\n\nThis is my solution to the [SIIM-ISIC Melanoma Classification](https://www.kaggle.com/c/siim-isic-melanoma-classification) competition using ResNet34.","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"%reload_ext autoreload\n%autoreload 2\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from fastai import *\nfrom fastai.vision import *","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import warnings\nwarnings.filterwarnings(\"ignore\", category=UserWarning, module=\"torch.nn.functional\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"path = Path('../input/siim-isic-melanoma-classification')\npath_512 = Path('../input/siim-isic-melanoma-classification-jpeg512')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In order to make it easier to train,we will use [resized images](https://www.kaggle.com/itacdonev/siim-isic-melanoma-classification-jpeg512) (Thank you [stats](https://www.kaggle.com/itacdonev)). Working with the current images in the `jpeg/train` folder, it would take 5 hours to run one complete epoch (See Version 1) because the images are of different sizes and fastai would need to resize each batch on the fly. If the images are resized beforehand, it will take less time to train, meaning we can run more epochs.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"np.random.seed(2)\ndata = ImageDataBunch.from_csv(\n            path_512, folder='train512', csv_labels='train.csv', ds_tfms=get_transforms(), label_col=7, size=128, suffix='.jpg', num_workers=0\n        ).normalize(imagenet_stats)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data.classes, data.c, len(data.train_ds), len(data.valid_ds), data.batch_size","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data.show_batch(rows=3, figsize=(12,9))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn = cnn_learner(data, models.resnet34, metrics=AUROC(), model_dir = '/kaggle/working')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I set up the `stage-8` model I created in the previous version. I will use this model to attempt to create a better model.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"! wget link-to-stage-8.pth","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.load('stage-8')\nlearn.data = data","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We add a callback called `SaveModelCallback` that will save the best model generated by `fit_one_cycle`.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"We will experiment with much lower learning rates.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.fit_one_cycle(\n    10, slice(1e-7), callbacks=[callbacks.SaveModelCallback(learn, every='improvement', monitor='auroc', name='stage-9')]\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.load('stage-9')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learn.export('/kaggle/working/export.pkl')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"learner = load_learner('/kaggle/working')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As has been pointed out [here](https://www.kaggle.com/edkahara/fast-ai-v3-melanoma-classification#912974), the probability of malignancy will always be `outputs[1]`. This means we may have submitted probabilities that were not the target probabilities in previous versions. This terrible oversight has been corrected. ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"img = open_image(path/'jpeg/test/ISIC_0052060.jpg')\n\n\npred_class,pred_idx,outputs = learner.predict(img)\n\n# Get the probability of malignancy\n\nprob_malignant = float(outputs[1])\n\nprint(pred_class)\nprint(prob_malignant)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test = os.listdir(path/'jpeg/test')\ntest.sort(key=lambda f: int(re.sub('\\D', '', f)))\n\nwith open('/kaggle/working/submission.csv', 'w', newline='') as file:\n    writer = csv.writer(file)\n    writer.writerow(['image_name', 'target'])\n    \n    for image_file in test:\n        image = os.path.join(path/'jpeg/test', image_file) \n        image_name = Path(image).stem\n\n        img = open_image(image)\n        pred_class,pred_idx,outputs = learner.predict(img)\n        target = float(outputs[1])\n\n        \n        writer.writerow([image_name, target])","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}