{"cells":[{"metadata":{},"cell_type":"markdown","source":"<img src=\"https://storage.googleapis.com/kaggle-media/competitions/MaxPlanck/Teaser_AnimationwLabels.gif\" width=\"800px\">\n\n"},{"metadata":{},"cell_type":"markdown","source":"**About the project**"},{"metadata":{},"cell_type":"markdown","source":"\nIn this challenge, we have to build a model to classify cloud organization patterns from satellite images.\nThere are many ways in which clouds can organize, but the boundaries between different forms of organization are murky. This makes it challenging to build traditional rule-based algorithms to separate cloud features. The human eye, however, is really good at detecting features—such as clouds that resemble flowers.\nThe input data are images with 4 different types of clouds (fish, flower, gravel and sugar) and a csv file which describes the clouds position in the image.\nThe predicted encodings should be against images that are scaled by 0.25 per side. In other words, while the images in Train and Test are 1400 x 2100 pixels, the predictions should be scaled down to a 350 x 525 pixel image. The reduction is required to achieve reasonable submission evaluation times"},{"metadata":{},"cell_type":"markdown","source":"**Libraries: **  \n\nTorch(with Catalyst and segmentation_models_pytorch), OpenCV, Albumentation and other standard Python libraries(numpy, pandas, matplotlib)  \n\n**Data visualization: ** \n\ncheck how many fish,flower,gravel and sugar valid data are in the training set  \ncheck how many clouds are per picture in the training set  \nshow a specific or random image with the segmentation data plotted over the original image  \n\n**Image augmentation**  \n\nFinal version: Resize(320, 640), HorizontalFlip,VerticalFlip,ShiftScaleRotate  \nOther attempts: Blur, MedianBlur, GridDistortion  \n  \n**Loss function**  \n\nAlthough the project is evaluated by Dice loss I have tested multiple loss function:  \nDice loss  \nIoU loss  \nFocal loss  \nA custom metric which contains a linear combination of Dice loss, IoU loss and Focal loss each with a specific weight  \n\n**Network arhitecture**  \n\nI have build the segmentation arhitecture based on different types of networks:  \nResnet50  \nResnet101\nEfficientnet-b2\nEfficientnet-b7  \nDensenet121  \n\n**Other aspects**  \n  \n  I have tried various learning rate options starting from 1e-3 for Encoder and 1e-2 for Decoder to smaller ones(5e-4 and 5e-3) and even equal learning rates for Encoder and Decoder  \n  Also, I have used and really helped ReduceLROnPlateau with factor=0.3 and patience=5  \n  \n**Problems and Issues**  \n\nI faced a problem with the Catalyst callbacks (DiceCallback(),InferCallback()). The RAM memory keep increasing along with the training iterations until it reaches the kernel limits. It seems that it is a memory leak somewhere or the I am not using them right\n\n**Acknowledgements and Inspirations**  \n\nA lot of thanks to Andrew Lukyanenko. His great kernel was a source of inspiration for general aproach and a lot of usefull functions(https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools#Exploring-augmentations-with-albumentations) \n  \n**Next steps**  \n\nDo more data visualization and analytics(cloud types corellations per picture), Aleksandra Deis has a great kernel on data exploration and visualization (https://www.kaggle.com/aleksandradeis/understanding-clouds-eda)  \nTry other data augmentation (Affine, Fliplr, ElasticTransformation, etc.)  \nTry new network arhitectures  \nImplement the same design using PyTorch from scratch(Dhananjay Raut has a great kernel on this: https://www.kaggle.com/dhananjay3/image-segmentation-from-scratch-in-pytorch)\n\n**Best score so far**  \n\n0.6331"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"!pip install git+https://github.com/qubvel/segmentation_models.pytorch","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nfrom PIL import Image\nimport numpy as np\nimport random\nfrom torch.utils.data import Dataset\nimport os\nimport cv2\nimport albumentations as albu\nfrom sklearn.model_selection import train_test_split\nimport segmentation_models_pytorch as smp\nfrom torch.utils.data import DataLoader\nimport torch\nfrom catalyst.dl.runner import SupervisedRunner\nfrom catalyst.dl.callbacks import DiceCallback, InferCallback, CheckpointCallback,JaccardCallback,IouCallback\nfrom torch.optim.lr_scheduler import ReduceLROnPlateau\nimport tqdm\nfrom catalyst.contrib.criterion import DiceLoss, IoULoss, FocalLossBinary\nimport torch.nn as nn\nimport pandas as pd\nfrom catalyst.dl import utils\nimport gc","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#plot descriptive stats\ndef DescriptiveStats(trainData):\n    #check how many fish,flower,gravel and sugar valid data are in the training set\n    occuranceDict={\"Fish\":0,\"Flower\":0,\"Gravel\":0,\"Sugar\":0}\n    for i in range(len(trainData['EncodedPixels'])):\n        if (pd.isnull(trainData['EncodedPixels'][i]))==False:\n            cloudType=trainData[\"Image_Label\"][i].split('_')[1]\n            occuranceDict[cloudType]+=1   \n    print(\"how many fish,flower,gravel and sugar valid data are in the training set\")\n    labels = occuranceDict.keys()\n    sizes = occuranceDict.values()\n    explode = (0.1, 0.1, 0.1, 0.1) \n    \n    fig1, ax1 = plt.subplots()\n    ax1.pie(sizes, explode=explode, labels=labels, autopct='%1.1f%%',\n            shadow=True, startangle=90)\n    ax1.axis('equal')  \n    plt.show()\n    print(\"how many clouds are per picture\")\n    #check how many clouds are per picture in the training set\n    occurancePerPic=trainData.loc[trainData['EncodedPixels'].isnull() == False, 'Image_Label'].apply(lambda x: x.split('_')[0]).value_counts().value_counts()\n    labels2 = occurancePerPic.keys() \n    fig2, ax2 = plt.subplots()\n    ax2.pie(occurancePerPic, explode=explode, labels=labels2, autopct='%1.1f%%',\n            shadow=True, startangle=90)\n    ax2.axis('equal')  \n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#show a specific image\ndef ShowImg(trainData,imgName):\n    image = Image.open(path+\"/train_images/\"+imgName)\n    print(\"Original Img\")\n    plt.imshow(image)\n    plt.show()\n    ss=trainData.loc[trainData['im_id']==imgName, 'EncodedPixels']\n    for row in ss:\n        print(trainData.loc[trainData['EncodedPixels']==row, \"label\"])\n        try: # label might not be there!\n            mask = rle_decode(row)\n        except Exception as exception:\n            mask = np.zeros((1400, 2100))\n            continue\n            \n        plt.imshow(image)\n        plt.imshow(mask, alpha=0.3, cmap='gray')\n        plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# create a custom loss (total loss=alfa*IoULoss + beta*DiceLoss + gamma*FocalLossBinary)\nclass CustomLoss(nn.Module):\n    def __init__(self, alpha=.7, beta=1.5, gamma=.4):\n        super().__init__()\n        self.alpha=alpha\n        self.beta=beta\n        self.gamma=gamma\n        self.lossIOU=IoULoss()\n        self.lossDice=DiceLoss()\n        self.lossFocal=FocalLossBinary()\n    def forward(self, input, target):\n        loss=self.alpha*self.lossIOU(input.cpu(), target.cpu()) + self.beta*self.lossDice(input.cpu(), target.cpu()) + self.beta*self.lossFocal(input.cpu(), target.cpu())\n        return loss.mean()/(alpha+beta+gamma)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Create mask based on df, image name and shape\ndef make_mask(df, image_name= 'img.jpg',\n              shape= (1400, 2100)):\n    encoded_masks = df.loc[df['im_id'] == image_name, 'EncodedPixels']\n    masks = np.zeros((shape[0], shape[1], 4), dtype=np.float32)\n\n    for idx, label in enumerate(encoded_masks.values):\n        if label is not np.nan:\n            mask = rle_decode(label)\n            masks[:, :, idx] = mask\n\n    return masks","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#Decode rle encoded mask    \ndef rle_decode(mask_rle='', shape=(1400, 2100)):\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int)\n                       for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0] * shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape, order='F')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def mask2rle(img):\n    pixels = img.T.flatten()\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    return ' '.join(str(x) for x in runs)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#reshape after using albumentation\ndef ConvertToTensorFormat(x, **kwargs):\n    return x.transpose(2, 0, 1).astype('float32')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def post_process(probability, threshold, min_size):\n    mask = cv2.threshold(probability, threshold, 1, cv2.THRESH_BINARY)[1]\n    num_component, component = cv2.connectedComponents(mask.astype(np.uint8))\n    predictions = np.zeros((350, 525), np.float32)\n    num = 0\n    for c in range(1, num_component):\n        p = (component == c)\n        if p.sum() > min_size:\n            predictions[p] = 1\n            num += 1\n    return predictions, num","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def dice(img1, img2):\n    img1 = np.asarray(img1).astype(np.bool)\n    img2 = np.asarray(img2).astype(np.bool)\n\n    intersection = np.logical_and(img1, img2)\n\n    return 2. * intersection.sum() / (img1.sum() + img2.sum())\n\ndef sigmoid(x): return 1/(1+np.exp(-x))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"class CloudDataset(Dataset):\n    def __init__(self,data,dataSetType,transforms,img_ids,preprocessing):\n        self.data=data\n        self.preprocessing=preprocessing\n        self.transforms=transforms\n        self.img_ids=img_ids\n        if (dataSetType==\"train\"):\n            self.imgFolder=path+\"/train_images\"\n        if (dataSetType==\"test\"):\n            self.imgFolder=path+\"/test_images\"\n\n    def __getitem__(self, idx):\n        image_name = self.img_ids[idx]\n\n        mask = make_mask(self.data, image_name)\n        image_path = os.path.join(self.imgFolder, image_name)\n        \n        img = cv2.imread(image_path)\n        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n        \n        augmented = self.transforms(image=img, mask=mask)\n        img = augmented['image']\n        mask = augmented['mask']\n        \n        if self.preprocessing:\n            preprocessed = self.preprocessing(image=img, mask=mask)\n            img = preprocessed['image']\n            mask = preprocessed['mask']\n            \n            \n        return img, mask\n\n    def __len__(self):\n        return len(self.img_ids)      ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#preprocess for specific network used\ndef get_preprocessing(preprocessing_fn=None):\n    if preprocessing_fn is not None:\n        _transform = [\n            albu.Lambda(image=preprocessing_fn),\n            albu.Lambda(image=ConvertToTensorFormat, mask=ConvertToTensorFormat),\n        ]\n    else:\n        _transform = [\n            albu.Normalize(),\n            albu.Lambda(image=ConvertToTensorFormat, mask=ConvertToTensorFormat),\n        ]\n    return albu.Compose(_transform)\n#augmentation for training data\ndef get_training_augmentation(p=0.5):\n    train_transform = [\n        albu.Resize(320, 640),\n        albu.HorizontalFlip(p=0.25),\n        albu.VerticalFlip(p=0.25),\n        albu.ShiftScaleRotate(scale_limit=0.5, rotate_limit=0, shift_limit=0.1, p=0.5, border_mode=0)\n    ]\n    return albu.Compose(train_transform)\n\n#for validation dataset it is just resize\ndef get_validation_augmentation():\n    train_transform = [\n        albu.Resize(320, 640)\n    ]\n    return albu.Compose(train_transform)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"path = '../input/understanding_cloud_organization'\n\n#read files\ntrainData = pd.read_csv(path+'/train.csv')\nsubSample = pd.read_csv(path+'/sample_submission.csv')\n\n#We can see that are not major imbalance in the dataset\nDescriptiveStats(trainData)\n\n#rearrange dataframe\ntrainData['label'] = trainData['Image_Label'].apply(lambda x: x.split('_')[1])\ntrainData['im_id'] = trainData['Image_Label'].apply(lambda x: x.split('_')[0])\n\nsubSample['label'] = subSample['Image_Label'].apply(lambda x: x.split('_')[1])\nsubSample['im_id'] = subSample['Image_Label'].apply(lambda x: x.split('_')[0])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#plot a random image\nimageToPlot=random.choice(trainData['im_id'])\nShowImg(trainData,imageToPlot)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"uniqueImgId=trainData.im_id.unique()\n#split in train-test\ntrain_ids, valid_ids = train_test_split(\n        uniqueImgId,\n        random_state=42,\n        test_size=0.1)\n#using efficientnet-b3 with imagenet weights\nENCODER = 'efficientnet-b2'\nENCODER_WEIGHTS = 'imagenet'\nDEVICE = 'cuda'\npreprocessing_fn = smp.encoders.get_preprocessing_fn(ENCODER, ENCODER_WEIGHTS)\nnum_workers = 0\nbs = 1\n\nACTIVATION = None\nmodel = smp.Unet(\n    encoder_name=ENCODER, \n    encoder_weights=ENCODER_WEIGHTS, \n    classes=4, \n    activation=ACTIVATION,\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_dataset = CloudDataset(data=trainData, dataSetType='train', img_ids=train_ids, transforms = get_training_augmentation(), preprocessing=get_preprocessing(preprocessing_fn))\nvalid_dataset = CloudDataset(data=trainData, dataSetType='train', img_ids=valid_ids, transforms = get_validation_augmentation(), preprocessing=get_preprocessing(preprocessing_fn))\n\ntrain_loader = DataLoader(train_dataset, batch_size=bs, shuffle=True, num_workers=num_workers)\nvalid_loader = DataLoader(valid_dataset, batch_size=bs, shuffle=False, num_workers=num_workers)\nloaders = {\n    \"train\": train_loader,\n    \"valid\": valid_loader\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"num_epochs = 10\nlogdir = \"./logs/CloudsSegmentation\"\n\noptimizer = torch.optim.Adam([\n    {'params': model.decoder.parameters(), 'lr': 5e-3}, \n    {'params': model.encoder.parameters(), 'lr': 7e-4},  \n])\n            \n#criterion=CustomLoss()\ncriterion=smp.utils.losses.BCEDiceLoss()\nscheduler = ReduceLROnPlateau(optimizer, factor=0.3, patience=5)\n\nrunner = SupervisedRunner()\ntorch.cuda.empty_cache()\ngc.collect()    ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"runner.train(\n    model=model,\n    criterion=criterion,\n    optimizer=optimizer,\n    scheduler=scheduler,\n    loaders=loaders,\n    #callbacks=[DiceCallback(),InferCallback()],\n    logdir=logdir,\n    num_epochs=num_epochs,\n    verbose=True\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#load best checkpoint\n\nencoded_pixels = []\nloaders = {\"infer\": valid_loader}\nrunner.infer(\n    model=model,\n    loaders=loaders,\n    callbacks=[\n        CheckpointCallback(\n            resume=f\"{logdir}/checkpoints/best.pth\"),\n        InferCallback()\n    ],\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"resizedMasks=[]\nprobabilities = np.zeros((len(valid_dataset)*4, 350, 525))\n# for each valid set and prediction on valid set:\n# the predictions should be scaled down to a 350 x 525 pixel image\n# make a resizedMasks of the valid elements\n#make a probabilities mask with 4*len(valid_dataset) (labels)\nfor i in range(len(valid_dataset)):\n    batch=valid_dataset[i]\n    output=runner.callbacks[0].predictions[\"logits\"][i]\n    image, mask = batch\n    for m in mask:\n        m = cv2.resize(m, dsize=(525, 350), interpolation=cv2.INTER_LINEAR)\n        resizedMasks.append(m)\n        \n    for j in range(len(output)):\n        probability = cv2.resize(output[j], dsize=(525, 350), interpolation=cv2.INTER_LINEAR)\n        probabilities[i * 4 + j, :, :] = probability","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#find the optimum threshold for each class\n\nclass_params = {}\nfor class_id in range(4):\n    print(\"Calculating optimum threshold for class:\",class_id)\n    attempts = []\n    for t in range(0, 100, 30):\n        t /= 100\n        for ms in [10000]:\n            masks = []\n            for i in range(class_id, len(probabilities), 4):\n                probability = probabilities[i]\n                predict, num_predict = post_process(sigmoid(probability), t, ms)\n                masks.append(predict)\n\n            d = []\n            for i, j in zip(masks, resizedMasks[class_id::4]):\n                if (i.sum() == 0) & (j.sum() == 0):\n                    d.append(1)\n                else:\n                    d.append(dice(i, j))\n\n            attempts.append((t, ms, np.mean(d)))\n\n    attempts_df = pd.DataFrame(attempts, columns=['threshold', 'size', 'dice'])\n\n\n    attempts_df = attempts_df.sort_values('dice', ascending=False)\n    print(attempts_df.head())\n    best_threshold = attempts_df['threshold'].values[0]\n    best_size = attempts_df['size'].values[0]\n    \n    class_params[class_id] = (best_threshold, best_size)\n        \n        \nprint(class_params)     ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"del probabilities\ndel resizedMasks\ndel attempts_df\ntorch.cuda.empty_cache()\ngc.collect()   ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"torch.cuda.empty_cache()\ngc.collect()    \n\ntest_ids = subSample['Image_Label'].apply(lambda x: x.split('_')[0]).drop_duplicates().values\n\ntest_dataset = CloudDataset(data=subSample, dataSetType='test', img_ids=test_ids, transforms = get_validation_augmentation(), preprocessing=get_preprocessing(preprocessing_fn))\ntest_loader = DataLoader(test_dataset, batch_size=1, shuffle=False, num_workers=0)\n\nloaders = {\"test\": test_loader}  \n\n\nencoded_pixels = []\nimage_id = 0\nfor i, test_batch in enumerate(tqdm.tqdm(loaders['test'])):\n    runner_out = runner.predict_batch({\"features\": test_batch[0].cuda()})['logits']\n    for i, batch in enumerate(runner_out):\n        for probability in batch:\n            \n            probability = probability.cpu().detach().numpy()\n            if probability.shape != (350, 525):\n                probability = cv2.resize(probability, dsize=(525, 350), interpolation=cv2.INTER_LINEAR)\n            predict, num_predict = post_process(sigmoid(probability), class_params[image_id % 4][0], class_params[image_id % 4][1])\n            if num_predict == 0:\n                encoded_pixels.append('')\n            else:\n                r = mask2rle(predict)\n                encoded_pixels.append(r)\n            image_id += 1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"subSample['EncodedPixels'] = encoded_pixels\nsubSample.to_csv('sub12.csv', columns=['Image_Label', 'EncodedPixels'], index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}