{"cells":[{"metadata":{"_uuid":"9380833a1a2503c5d3518f0ed8d6df8dcf05b7c2"},"cell_type":"markdown","source":"## Overview\nResnet34 is commonly used as an encoder for U-net and SSD, boosting the model performance and training time since you do not need to train the model from scratch. However, in particular cases it makes sense to do fine-tuning of Resnet34 model before using it as a decoder for object localization or image segmentation. In this competition the size of ship masks is much smaller than the size of images that leads to quite unbalanced training with ~1 positive pixel per 1000 negative ones. If images with no ships are used, instead of ~1:1000 you will end up with ~1:10000 unbalance, which is quite tough. Moreover, the training time is ~4 times longer since you need to process more images in each epoch. So, it is reasonable to drop empty images and focus only on ones with ships. Meanwhile, since the current dataset is quite different from ImageNet, the empty images are quite helpful in fine-tuning your encoder on a pseudo task - ship detection. Moreover, when the training of your U-net or SSD model is completed, you can run the model on images without ships, add false positives (~4000 in my case) as negative example to you training set, and train the model for several additional epochs. Finally, a good model focused on a single task, ship detection, can boost the final score when you stack up it with U-net or SSD. If you predict a ship for an empty image you will get automatically zero score for it, and since PLB has ~85% of empty images, prediction of empty images is quite important.\n\nIn this notebook I want to share how to pretrain Resnet34 (or higher end models) on a ship detection task. After training of the head layers of the model on 256x256 rescaled images for one epoch the accuracy has reached 93.7%. The following fine-tuning of entire model for 2 more epochs with learning rate annealing boosted the accuracy to ~97%. If the training is continued for several epochs with a new data set composed of images of 384x384 resolution, the accuracy could be boosted to ~98%. Unfortunately, continuing training the model on full resolution, 768x768, images leaded to reduction of the accuracy that is likely attributed to insufficient model capacity."},{"metadata":{"trusted":true,"_uuid":"2a2f9181ed56a8310f6188ac1254f903574fb115"},"cell_type":"code","source":"from fastai.conv_learner import *\nfrom fastai.dataset import *\n\nimport pandas as pd\nimport numpy as np\nimport os\nfrom sklearn.model_selection import train_test_split","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8fa09c99d9f5b03e8e3a213f6d84902d5e1d59e1"},"cell_type":"markdown","source":"### Data"},{"metadata":{"trusted":true,"_uuid":"85e1116d1d675a8ae7fca878c2ddf4469207665d"},"cell_type":"code","source":"PATH = './'\nTRAIN = '../input/airbus-ship-detection/train_v2/'\nTEST = '../input/airbus-ship-detection/test_v2/'\nSEGMENTATION = '../input/airbus-ship-detection/train_ship_segmentations_v2.csv'\nLOAD_PATH = '../input/fine-tuning-resnet34-on-ship-detection/models/'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f487dd77687f4edb070bd5d2dc9da9a001d62bdb"},"cell_type":"code","source":"nw = 2   #number of workers for data loader\narch = resnet34 #specify target architecture","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"56ed39146115a4767a257fec60a3b367284fa0d6"},"cell_type":"code","source":"train_names = [f for f in os.listdir(TRAIN)]\ntest_names = [f for f in os.listdir(TEST)]\n#5% of data in the validation set is sufficient for model evaluation\ntr_n, val_n = train_test_split(train_names, test_size=0.05, random_state=42)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d1085ab7f398dbcde2c28a4476af73f9d04df85e"},"cell_type":"code","source":"class pdFilesDataset(FilesDataset):\n    def __init__(self, fnames, path, transform):\n        self.segmentation_df = pd.read_csv(SEGMENTATION).set_index('ImageId')\n        super().__init__(fnames, transform, path)\n    \n    def get_x(self, i):\n        img = open_image(os.path.join(self.path, self.fnames[i]))\n        if self.sz == 768: return img \n        else: return cv2.resize(img, (self.sz, self.sz))\n    \n    def get_y(self, i):\n        if(self.path == TEST): return 0\n        masks = self.segmentation_df.loc[self.fnames[i]]['EncodedPixels']\n        if(type(masks) == float): return 0 #NAN - no ship \n        else: return 1\n    \n    def get_c(self): return 2 #number of classes","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"25fa3283c992696575914a5fdb6ebc433a0b5d1f"},"cell_type":"code","source":"def get_data(sz,bs):\n    #data augmentation\n    aug_tfms = [RandomRotate(20, tfm_y=TfmType.NO),\n                RandomDihedral(tfm_y=TfmType.NO),\n                RandomLighting(0.05, 0.05, tfm_y=TfmType.NO)]\n    tfms = tfms_from_model(arch, sz, crop_type=CropType.NO, tfm_y=TfmType.NO, \n                aug_tfms=aug_tfms)\n    ds = ImageData.get_ds(pdFilesDataset, (tr_n[:-(len(tr_n)%bs)],TRAIN), \n                (val_n,TRAIN), tfms, test=(test_names,TEST))\n    md = ImageData(PATH, ds, bs, num_workers=nw, classes=None)\n    return md","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"193381699f5595c916647bfd6c51eaeba699379d"},"cell_type":"markdown","source":"### Model"},{"metadata":{"trusted":true,"_uuid":"ddaf708a2b33a383ccd712d1758718c38eeeb922","scrolled":true},"cell_type":"code","source":"sz = 256 #image size\nbs = 64  #batch size\n\nmd = get_data(sz,bs)\nlearn = ConvLearner.pretrained(arch, md, ps=0.5) #dropout 50%\nlearn.opt_fn = optim.Adam","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d22885312709be8df9b077cc2f3da59c23439b31"},"cell_type":"markdown","source":"I begin with finding the optimal learning rate. The following function runs training with different lr and records the loss. Increase of the loss indicates onset of divergence of training. The optimal lr lies in the vicinity of the minimum of the curve but before the onset of divergence. Based on the following plot, for the current setup the divergence starts at ~0.01, and the recommended learning rate is ~0.002."},{"metadata":{"_uuid":"cec5997ac36f7c56e101a899291f0a8b0edeca1f"},"cell_type":"markdown","source":"I'm going to load a model trained in the original version of the kernel, so I comment lines corresponding to training. Please see the original kernel for more details."},{"metadata":{"trusted":true,"_uuid":"149f05054b0fb3dadc126ea2789653603fd4b7ea","collapsed":true},"cell_type":"code","source":"#learn.lr_find()\n#learn.sched.plot()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"196cf509806bf19556d30514e79e97b41b9b04c1"},"cell_type":"markdown","source":"Training the head part of the model with constant learning rate for one epoch. "},{"metadata":{"trusted":true,"_uuid":"ef68043848b5d05a0bbc2c53b13dd54767b82f5e","collapsed":true},"cell_type":"code","source":"#learn.fit(2e-3, 1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"54961ec96c09672e66ccea9f79c10567dc458df0"},"cell_type":"markdown","source":"Unfreeze the model and train it with differential learning rate. The lr of the head part is still 2e-3, while the middle layers of the model a trained with 5e-4 lr, and the base is trained with even smaller lr, 1e-4, since low level detector do not vary much from one image data set to another."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"33ffa27057b85dd084df994e34874a8a7c131c46"},"cell_type":"code","source":"#learn.unfreeze()\n#lr=np.array([1e-4,5e-4,2e-3])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"132c754c6732d8224cbbc60eae835c73ebf4b63e","collapsed":true},"cell_type":"code","source":"#learn.fit(lr, 1, cycle_len=2, use_clr=(20,8))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6103281a0c1c7db61fae2743ebcdcddaef9023db"},"cell_type":"markdown","source":"The training has been run with learning rate annealing. Periodic lr increase followed by slow decrease drives the system out of steep minima (when lr is high) towards broader ones (which are explored when lr decreases) that enhances the ability of the model to generalize and reduces overfitting. Due to time limit, only once cycle has been run. But ideally several cycles must be run with gradual increase of the image size, etc. 256x256, 384x384, 768x768, to reach better performance of the model."},{"metadata":{"trusted":true,"_uuid":"ccfd2cf5e8edeceeec185f988ecd4ebd24d14b80","collapsed":true},"cell_type":"code","source":"#learn.sched.plot_lr()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"56a469df5b92590a3a5fd4d3015c2f838c307813"},"cell_type":"code","source":"#learn.save('Resnet34_lable_256_1')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"583311b2639aee0c32a81328ab47d4df21f675db"},"cell_type":"code","source":"path = learn.models_path\nlearn.models_path = LOAD_PATH\nlearn.load('Resnet34_lable_256_1')\nlearn.models_path = path","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a5af78f6512ab4f514818ef47b3481ef67a65e46"},"cell_type":"markdown","source":"### Prediction"},{"metadata":{"trusted":true,"_uuid":"4aece9e7181c6eafeba7fa831fa73efeecbd0b8d"},"cell_type":"code","source":"log_preds,y = learn.predict_with_targs(is_test=True)\nprobs = np.exp(log_preds)[:,1]\npred = (probs > 0.5).astype(int)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b50fc52ed7f49e83e79d9ff20cbb2cbf75d3c1bb"},"cell_type":"code","source":"df = pd.DataFrame({'id':test_names, 'p_ship':probs})\ndf.to_csv('ship_detection.csv', header=True, index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b1eb1634fee993062bc1386da8ce80b54895c60e"},"cell_type":"markdown","source":"### Training on high resolution images"},{"metadata":{"_uuid":"872e04eaecc268599dab225b43753503ad40910d"},"cell_type":"markdown","source":"Since each epoch on higher resolution images, like 384x384 or 768x768, takes quite long time, training the model from scratch on these images is quite inefficient. Fortunately, modern convolutional nets support input images of arbitrary resolution. To decrease the training time, one can start training the model on low resolution images first and continue training on higher resolution images for only a few epochs. In addition, a model pretrained on low resolution images first generalizes better since a pixel information is less available and high order features are tended to be used."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"df8ae8c3b76092761c6eb813a98aca5e22e18c8e"},"cell_type":"code","source":"#sz = 384 #image size\n#bs = 32  #batch size\n\n#md = get_data(sz,bs)\n#learn = ConvLearner.pretrained(arch, md, ps=0.5) #dropout 50%\n#learn.opt_fn = optim.Adam\n#learn.unfreeze()\n#lr=np.array([1e-4,5e-4,2e-3])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"ee4662b81bbc6d288027850e3ccef5426b72ec66"},"cell_type":"code","source":"#learn.load('Resnet34_lable_256_1')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"59f3f312fccf272970bafe5d4295094ef6e1af96"},"cell_type":"markdown","source":"Due to the kernal run time limit, the following code does not have enough time to be executed. I hope it is possible to rerun this kernel with disabled cells from \"Model\" part and included data from the previous commit to complete training of higher resolution images. Each epoch on images 383x384 takes about 1.5 hours. Probably, some paths, like \"learn.models_path\", may need to be changed."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"6fffd57bcafceeed806f5d627af53d10a7367b36"},"cell_type":"code","source":"#learn.fit(lr/2, 1, cycle_len=2, use_clr=(20,8)) #lr is smaller since bs is only 32\n#learn.save('Resnet34_lable_384_1')","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}