{"cells":[{"metadata":{},"cell_type":"markdown","source":"# GENet : GPU Efficient Network\n\n![](https://raw.githubusercontent.com/idstcv/GPU-Efficient-Networks/master/misc/genet_acc_speed_curve.jpg)\n\nThe proposed design space is optimized for fast GPU inference. In this space, it uses a semi-automatic NAS to\nhelp us design GPU-Efficient Networks. GENets use full convolutions in low-level stages and depth-wise\nconvolution and/or bottleneck structure in high-level stages. \n\nThis design is inspired by the observation that convolutional kernels in the high-level stages are more likely to have low intrinsic rank and different types of convolutions have different kinds of efficiency on GPU."},{"metadata":{},"cell_type":"markdown","source":"# Albumentations\nAlbumentations is a Python library for image augmentation. Image augmentation is used in deep learning and computer vision tasks to increase the quality of trained models. The purpose of image augmentation is to create new training samples from the existing data.\n\n* Albumentations supports all common computer vision tasks such as classification, semantic segmentation, instance segmentation, object detection, and pose estimation.\n* The library provides a simple unified API to work with all data types: images (RBG-images, grayscale images, multispectral images), segmentation masks, bounding boxes, and keypoints.\n* The library contains more than 70 different augmentations to generate new training samples from the existing data.\n* Albumentations is fast.\n\n* Installation:- pip install -U albumentations"},{"metadata":{},"cell_type":"markdown","source":"# Load GENets"},{"metadata":{"trusted":true},"cell_type":"code","source":"!git clone https://github.com/idstcv/GPU-Efficient-Networks.git","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"cd ./GPU-Efficient-Networks","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import GENet","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"cd ../","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Import Libraries"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# Python library to interact with the file system.\nimport os\n\n# Software library written for data manipulation and analysis.\nimport pandas as pd\n\n# fastai library for computer vision tasks\nfrom fastai.vision.all import *\n\n# Python library for image augmentation\nimport albumentations as A\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Load training data"},{"metadata":{"trusted":true},"cell_type":"code","source":"path = Path('../input/cassava-leaf-disease-classification')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv(path/'train.csv')\ntrain_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['image_id'] = train_df['image_id'].map(lambda x : path /'train_images'/x )\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Create Dataloaders"},{"metadata":{"trusted":true},"cell_type":"code","source":"# obtain the input images.\ndef get_x(r):\n    return r['image_id']\n\n# obtain the targets.\ndef get_y(r):\n    return r['label']","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The Albumentation code has been borrowed from Fastai [docs](https://docs.fast.ai/tutorial.albumentations.html). It's very common to use different transforms on the training dataset versus the validation dataset. Lets see how!"},{"metadata":{"trusted":true},"cell_type":"code","source":"'''AlbumentationsTransform will perform different transforms over both\n   the training and validation datasets ''' \nclass AlbumentationsTransform(RandTransform):\n    \n    '''split_idx is None, which allows for us to say when we're setting our split_idx.\n       We set an order to 2 which means any resize operations are done first before our new transform. '''\n    split_idx, order = None, 2\n    \n    def __init__(self, train_aug, valid_aug): store_attr()\n    \n    # Inherit from RandTransform, allows for us to set that split_idx in our before_call.\n    def before_call(self, b, split_idx):\n        self.idx = split_idx\n    \n    # If split_idx is 0, run the trainining augmentation, otherwise run the validation augmentation. \n    def encodes(self, img: PILImage):\n        if self.idx == 0:\n            aug_img = self.train_aug(image=np.array(img))['image']\n        else:\n            aug_img = self.valid_aug(image=np.array(img))['image']\n        return PILImage.create(aug_img)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_train_aug(size): \n    \n    return A.Compose([\n            # allows to combine RandomCrop and RandomScale\n            A.RandomResizedCrop(size,size),\n            \n            # Transpose the input by swapping rows and columns.\n            A.Transpose(p=0.5),\n        \n            # Flip the input horizontally around the y-axis.\n            A.HorizontalFlip(p=0.5),\n        \n            # Flip the input horizontally around the x-axis.\n            A.VerticalFlip(p=0.5),\n        \n            # Randomly apply affine transforms: translate, scale and rotate the input.\n            A.ShiftScaleRotate(p=0.5),\n        \n            # Randomly change hue, saturation and value of the input image.\n            A.HueSaturationValue(hue_shift_limit=0.2, sat_shift_limit=0.2, val_shift_limit=0.2, p=0.5),\n        \n            # Randomly change brightness and contrast of the input image.\n            A.RandomBrightnessContrast(brightness_limit=(-0.1,0.1), contrast_limit=(-0.1, 0.1), p=0.5),\n        \n            # CoarseDropout of the rectangular regions in the image.\n            A.CoarseDropout(p=0.5),\n        \n            # CoarseDropout of the square regions in the image.\n            A.Cutout(p=0.5) ])\n\ndef get_valid_aug(size): \n    \n    return A.Compose([\n    # Crop the central part of the input.   \n    A.CenterCrop(size, size, p=1.),\n    \n    # Resize the input to the given height and width.    \n    A.Resize(size,size)], p=1.)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"'''The first step item_tfms resizes all the images to the same size (this happens on the CPU) \n   and then batch_tfms happens on the GPU for the entire batch of images. '''\n# Transforms we need to do for each image in the dataset\nitem_tfms = [Resize(256), AlbumentationsTransform(get_train_aug(256), get_valid_aug(256))]\n\n# Transforms that can take place on a batch of images\nbatch_tfms = [Normalize.from_stats(*imagenet_stats)]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_data(bs=32, data_df=train_df):\n    dblock = DataBlock(blocks=(ImageBlock, CategoryBlock),\n                       splitter=RandomSplitter(seed=42), # split data into training and validation subsets.\n                       get_x=get_x, # obtain the input images.\n                       get_y=get_y, # obtain the targets.\n                       item_tfms = item_tfms,\n                       batch_tfms = batch_tfms)\n    return dblock.dataloaders(data_df,bs=bs)\n\ndls = get_data()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# We can call show_batch() to see what a sample of a batch looks like.\ndls.show_batch()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Model Definition\n\n* There are three pre-trained models, GENet-large/normal/small.\n* GENet_large (31M param)\n* GENet_normal (21M param)\n* GENet_small (8.17M)\n\n    * GENet-large/normal/small use different input image resolutions.\n    * GENet-large, size = 256\n    * GENet-normal, size = 192\n    * GENet-small, size = 192\n    "},{"metadata":{"trusted":true},"cell_type":"code","source":"model = GENet.genet_large(pretrained=True, root='../input/genetparam/')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Group together some dls, a model, and metrics to handle training\nlearn = Learner(dls, model, metrics = accuracy) ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Choosing a good learning rate\nlearn.lr_find()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# We can use the fine_tune function to train a model with this given learning rate\nlearn.fine_tune(4, base_lr=0.0012022644514217973)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 85% accuracy in 5 epochs! Not Bad!!"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Plot training and validation losses.\nlearn.recorder.plot_loss()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Interpretation methods for classification models.\ninterp = ClassificationInterpretation.from_learner(learn)\n\n# Show images in top_losses along with their prediction, actual, loss, and probability of actual class.\ninterp.plot_top_losses(5, nrows=5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Make Submission file"},{"metadata":{"trusted":true},"cell_type":"code","source":"sample = pd.read_csv(path/'sample_submission.csv')\nsample","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"_sample = sample.copy()\n_sample['image_id'] = _sample['image_id'].map(lambda x:path/'test_images'/x)\ntest_dl = dls.test_dl(_sample)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"_sample.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_dl.show_batch()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Test Time Augmentation (TTA)\nSimilar to what Data Augmentation is doing to the training set, the purpose of Test Time Augmentation is to perform random modifications to the test images. Thus, instead of showing the regular, “clean” images, only once to the trained model, we will show it the augmented images several times. We will then average the predictions of each corresponding image and take that as our final guess.\n\nThe reason why it works is that, by averaging our predictions, on randomly modified images, we are also averaging the errors. The error can be big in a single vector, leading to a wrong answer, but when averaged, only the correct answer stand out."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"a, _ = learn.tta(dl=test_dl, n=8)\npred = a.argmax(dim=1).numpy()\nsample['label'] = pred","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample.to_csv('submission.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## *Upvote the kernel if you found it insightful!*"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}