{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Basic EDA of Satellite clouds\n\n![Competition label](https://storage.googleapis.com/kaggle-media/competitions/MaxPlanck/Teaser_AnimationwLabels.gif)\n\nHere I am performing a simple _Exploratory data analysis_ on [Understanding Clouds from Satellite Images](https://www.kaggle.com/c/understanding_cloud_organization) dataset provided with the same contest on [Kaggle](https://www.kaggle.com/).\n> The images were downloaded from [NASA Worldview](https://worldview.earthdata.nasa.gov/). Three regions, spanning 21 degrees longitude and 14 degrees latitude, were chosen. The true-color images were taken from two polar-orbiting satellites, TERRA and AQUA, each of which pass a specific region once a day. Due to the small footprint of the imager (MODIS) on board these satellites, an image might be stitched together from two orbits. The remaining area, which has not been covered by two succeeding orbits, is marked black."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"import os\nimport pandas as pd\nimport cv2\nimport numpy as np\nfrom PIL import Image\nfrom matplotlib import pyplot as plt\nos.listdir('../input')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Exploring files in dataset\n\nSo we have the data split into the training and the testing dataset. Lets now explore the dataset more with numbers"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"print('We have {} files in dataset'.format(len(os.listdir('../input/understanding_cloud_organization/train_images/'))))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Reading the training dataset\ndf = pd.read_csv('../input/understanding_cloud_organization/train.csv')\ndf.tail()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Split the labels from the Images ID\nnew = df.Image_Label.str.split('_', expand=True).rename(columns={0:'id',1:'labels'})\ndf['id']=new['id']\ndf['labels']=new['labels']\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can clearly see that the dataset contains the images has been mapped with all the labels and each image has been discretely provided us all the segmetation maps of the images\nThe masks of the images have been encoded and then fed into the training file. For this we do not have seperate mask images. So images without some specific class has _NaN_ in its place. "},{"metadata":{},"cell_type":"markdown","source":"Now lets verify our preposition here"},{"metadata":{"trusted":true},"cell_type":"code","source":"# All individual labels\nlabels_counts = df.labels.value_counts()\nlabels_counts","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We see 4 seperate classes:\n1. Sugar\n2. Fish\n3. Gravel\n4. Flower\n\nAnd as each image has been given all the labels, we have 5546 labels for each classes"},{"metadata":{"trusted":true},"cell_type":"code","source":"print('We have {} NaN classes'.format(df.EncodedPixels.isna().sum()))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Plotting the nan class occurance\nvalue_count = df[df.EncodedPixels.isna()]['labels'].value_counts()\nvalue_count.plot.bar()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Plotting the classes occurances\nnon_nan_labels = labels_counts - value_count\nnon_nan_labels.plot.bar()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## What's under those masks\n\nLets now decode the masks provided in the training dataset and view what's beneath those masks.\n\nFor this I would like to thank [xhlulu](https://www.kaggle.com/xhlulu) for making his awesome kernels publically available. \n"},{"metadata":{"_kg_hide-input":false,"trusted":true},"cell_type":"code","source":"def np_resize(img, input_shape):\n    \"\"\"\n    Reshape a numpy array, which is input_shape=(height, width), \n    as opposed to input_shape=(width, height) for cv2\n    \"\"\"\n    height, width = input_shape\n    return cv2.resize(img, (width, height))\n    \ndef mask2rle(img):\n    '''\n    img: numpy array, 1 - mask, 0 - background\n    Returns run length as string formated\n    '''\n    pixels= img.T.flatten()\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    return ' '.join(str(x) for x in runs)\n\ndef rle2mask(rle, input_shape):\n    width, height = input_shape[:2]\n    \n    mask= np.zeros( width*height ).astype(np.uint8)\n    \n    array = np.asarray([int(x) for x in rle.split()])\n    starts = array[0::2]\n    lengths = array[1::2]\n\n    current_position = 0\n    for index, start in enumerate(starts):\n        mask[int(start):int(start+lengths[index])] = 1\n        current_position += lengths[index]\n        \n    return mask.reshape(height, width).T\n\ndef build_masks(rles, input_shape, reshape=None):\n    depth = len(rles)\n    if reshape is None:\n        masks = np.zeros((*input_shape, depth))\n    else:\n        masks = np.zeros((*reshape, depth))\n    \n    for i, rle in enumerate(rles):\n        if type(rle) is str:\n            if reshape is None:\n                masks[:, :, i] = rle2mask(rle, input_shape)\n            else:\n                mask = rle2mask(rle, input_shape)\n                reshaped_mask = np_resize(mask, reshape)\n                masks[:, :, i] = reshaped_mask\n    \n    return masks\n\ndef build_rles(masks, reshape=None):\n    width, height, depth = masks.shape\n    \n    rles = []\n    \n    for i in range(depth):\n        mask = masks[:, :, i]\n        \n        if reshape:\n            mask = mask.astype(np.float32)\n            mask = np_resize(mask, reshape).astype(np.int64)\n        \n        rle = mask2rle(mask)\n        rles.append(rle)\n        \n    return rles","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"os.listdir('../input/understanding_cloud_organization/train_images/')[:10]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_filename = '8db703a.jpg'\nsample_image_df = df[df['id'] == sample_filename]\nsample_path = f\"../input/understanding_cloud_organization/train_images/{sample_image_df['id'].iloc[0]}\"\nsample_img = cv2.imread(sample_path)\nsample_rles = sample_image_df['EncodedPixels'].values\nsample_masks = build_masks(sample_rles, input_shape=(1400, 2100))\n\nfig, axs = plt.subplots(5, figsize=(12, 12))\naxs[0].imshow(sample_img, cmap='gray')\naxs[0].axis('off')\n\nfor i in range(4):\n    axs[i+1].imshow(sample_masks[:, :, i])\n#     axs[i+1].axis('off')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"maskid=2\nymin = sample_masks[:,:,maskid].argmax(axis=1).argmax()\nxmin = sample_masks[:,:,maskid].argmax(axis=0).argmax()\nymax = sample_masks[ymin:,xmin:,maskid].argmin(axis=1).argmin()+ymin\nxmax = sample_masks[ymin:,xmin:,maskid].argmin(axis=0).argmin()+xmin\n\nprint(xmin, ymin, xmax, ymax)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_masks[ymin:,:,maskid].argmin(axis=1).shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# (sample_img[ymin:ymax, xmin:xmax], cmap='gray')\n# cv2.rectangle(sample_img, (xmin, ymin), (xmax, ymax), (0,255,0), 5)\nplt.imshow(sample_img[ymin:ymax,xmin:xmax], cmap='gray')\nplt.axis('off')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"This is all that I thought would be required for understanding the dataset. But if you would like to help improve the kernel, you are most welcome to contribute to this kernel. If there's anything  else you would want me to add, do comment."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}