{"cells":[{"metadata":{},"cell_type":"markdown","source":"## Introduction & Loading The Data\n\nClimate change is one of the most important and pressing issues of our time. In this competition, we have the opportunity to use data to understand important climatic variables. The idea is if we can understand the cloud, we can better understand our environment and how it is changing. \n\nThe idea of the competition is fairly simple. There are different types or patterns of clouds, and our task is to identify them and classify these types.\n\n**If you like my work please upvote this Kernel. This encourages or motivates people like me, who contributes to Kaggle on their own time with the intention to share knowledge, to continue the effort. Furthermore, if I made a mistake or can do something more, please leave a comment in the comments section to help me out. Many thanks in advance!**"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np \nimport pandas as pd\nimport os\nimport cv2\n# visualization\nimport matplotlib.pyplot as plt\nfrom matplotlib import patches as patches\n# plotly offline imports\nfrom plotly.offline import download_plotlyjs, init_notebook_mode, iplot\nfrom plotly import subplots\nimport plotly.express as px\nimport plotly.figure_factory as ff\nfrom plotly.graph_objs import *\nfrom plotly.graph_objs.layout import Margin, YAxis, XAxis\ninit_notebook_mode()\n# frequent pattern mining\nfrom mlxtend.frequent_patterns import fpgrowth","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data_path = '/kaggle/input/understanding_cloud_organization'\ntrain_csv_path = os.path.join('/kaggle/input/understanding_cloud_organization','train.csv')\ntrain_image_path = os.path.join('/kaggle/input/understanding_cloud_organization','train_images')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"pd.read_csv(train_csv_path).head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# load full data and label no mask as -1\ntrain_df = pd.read_csv(train_csv_path).fillna(-1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# image id and class id are two seperate entities and it makes it easier to split them up in two columns\ntrain_df['ImageId'] = train_df['Image_Label'].apply(lambda x: x.split('_')[0])\ntrain_df['Label'] = train_df['Image_Label'].apply(lambda x: x.split('_')[1])\n# lets create a dict with class id and encoded pixels and group all the defaults per image\ntrain_df['Label_EncodedPixels'] = train_df.apply(lambda row: (row['Label'], row['EncodedPixels']), axis = 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now we have our modified train.csv, we answer a few basic questions below:"},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Total number of images: %s' % len(train_df['ImageId'].unique()))\nprint('Images with at least one label: %s' % len(train_df[train_df['EncodedPixels'] != -1]['ImageId'].unique()))\nprint('Total instance or examples of defects: %s' % len(train_df[train_df['EncodedPixels'] != -1]))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Different Types of Clouds\n\nThe first important thing I want to understand are the followings:\n1. What are the different types of clouds we have in our dataset.\n2. How does the mask for these formation looks like.\n2. Value cout of how many types of clouds we have per image.\n3. Distribution / frequency of each types of clouds in our dataset."},{"metadata":{"trusted":true},"cell_type":"code","source":"# different types of clouds we have in our dataset\ntrain_df['Label'].unique()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# visualize steel image with four classes of faults in seperate columns\ndef viz_two_instance_of_a_mask(encoded_masks_1, encoded_masks_2):\n    '''\n    visualize an image with two types of defects by plotting them on two columns\n    with the defect overlayed on top of the original image.\n\n    Parameters: \n    img_path (str): path of images\n    img_id (str): image id or filename of the path\n    encoded_masks (list): a list of strings of encoded masks \n    \n    Returns: \n    matplotlib image plot in columns for two classes iwth defect\n    '''\n    \n    fig, ax = plt.subplots(nrows=1, ncols=2, sharey=True, figsize=(20,10))\n    axid = 0\n    \n    for idx, encoded_mask in enumerate(encoded_masks):\n        class_id = idx + 1\n        if encoded_mask == -1:\n            pass\n        else:\n            mask_decoded = rle_to_mask(encoded_mask, 256, 1600)\n            ax[axid].get_xaxis().set_ticks([])\n            ax[axid].get_yaxis().set_ticks([])\n            ax[axid].text(0.25, 0.25, 'Image Id: %s - Class Id: %s' % (img_id, class_id), fontsize=12)\n            ax[axid].imshow(img)\n            ax[axid].imshow(mask_decoded, alpha=0.15, cmap=\"Blues\")\n            axid += 1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# lets group each of the types and their mask in a list so we can do more aggregated counts\ngrouped_EncodedPixels = train_df.groupby('ImageId')['Label_EncodedPixels'].apply(list)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# count the number of labels per image has\nlabels_per_image_count = grouped_EncodedPixels.apply(lambda x: len([x[0] for x in x if x[1]!=-1])).value_counts()\n# count frequency of each type of cloud\nlabel_type_per_image = train_df[train_df['EncodedPixels']!=-1]['Label'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# now we have the data ready lets plot them to answer our questions\ntrace0 = Bar(x=labels_per_image_count.index, y=labels_per_image_count.values, name = 'Number of Cloud Types Per Image')\ntrace1 = Bar(x=label_type_per_image.index, y=label_type_per_image.values, name = 'Frequency of Different Clouds')\nfig = subplots.make_subplots(rows=1, cols=2)\nfig.append_trace(trace0, 1, 1)\nfig.append_trace(trace1, 1, 2)\nfig['layout'].update(height=400, width=900, title='Label Count and Frequency Per Image')\niplot(fig)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Looks like most of the time we have 2-3 types of cloud formation in one image. 4 types of cloud formation in one image is very rare. Only one type of cloud formation in the image is also somewhat common. Furthermore, the data looks very evenly distributed for all four types of cloud formation. This is going to be an awesome compeition!"},{"metadata":{},"cell_type":"markdown","source":"## Drawing Clouds\n\nNow we know a little bit about the distribution of our data, we need to take a look at it and get an understanding of what it's all about. First, the masks are encoded so we will need the following function to decode the mask."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"def rle_to_mask(rle_string, height, width):\n    '''\n    convert RLE(run length encoding) string to numpy array\n\n    Parameters: \n    rle_string (str): string of rle encoded mask\n    height (int): height of the mask\n    width (int): width of the mask \n\n    Returns: \n    numpy.array: numpy array of the mask\n    '''\n    \n    rows, cols = height, width\n    \n    if rle_string == -1:\n        return np.zeros((height, width))\n    else:\n        rle_numbers = [int(num_string) for num_string in rle_string.split(' ')]\n        rle_pairs = np.array(rle_numbers).reshape(-1,2)\n        img = np.zeros(rows*cols, dtype=np.uint8)\n        for index, length in rle_pairs:\n            index -= 1\n            img[index:index+length] = 255\n        img = img.reshape(cols,rows)\n        img = img.T\n        return img","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Lets just plot a single image and a mask to get an idea of what it looks like."},{"metadata":{"trusted":true},"cell_type":"code","source":"img = cv2.imread(os.path.join(train_image_path, train_df['ImageId'][0]))\nmask_decoded = rle_to_mask(train_df['Label_EncodedPixels'][0][1], img.shape[0], img.shape[1])\nfig, ax = plt.subplots(nrows=1, ncols=2, sharey=True, figsize=(20,10))\nax[0].imshow(img)\nax[1].imshow(mask_decoded)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As it seems, there are images of clouds, and a mask not outlining the exact clouds but roughly the area with the same kind of patterns. And from our last section, we know an image can have more than one type of cloud patterns. So, I propose we visualize an image in two columns. First, shows the different types of cloud formation with a bounding box. On the second column, we visualize the cloud picture with the mask segments as an overlay. Lets go!"},{"metadata":{"trusted":true},"cell_type":"code","source":"def bounding_box(img):\n    # return max and min of a mask to draw bounding box\n    rows = np.any(img, axis=1)\n    cols = np.any(img, axis=0)\n    rmin, rmax = np.where(rows)[0][[0, -1]]\n    cmin, cmax = np.where(cols)[0][[0, -1]]\n\n    return rmin, rmax, cmin, cmax","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_cloud(img_path, img_id, label_mask):\n    img = cv2.imread(os.path.join(img_path, img_id))\n    \n    fig, ax = plt.subplots(nrows=1, ncols=2, sharey=True, figsize=(20,10))\n    ax[0].imshow(img)\n    ax[1].imshow(img)\n    cmaps = {'Fish': 'Blues', 'Flower': 'Reds', 'Gravel': 'Greys', 'Sugar':'Purples'}\n    colors = {'Fish': 'Blue', 'Flower': 'Red', 'Gravel': 'Gray', 'Sugar':'Purple'}\n    for label, mask in label_mask:\n        mask_decoded = rle_to_mask(mask, img.shape[0], img.shape[1])\n        if mask != -1:\n            rmin, rmax, cmin, cmax = bounding_box(mask_decoded)\n            bbox = patches.Rectangle((cmin,rmin),cmax-cmin,rmax-rmin,linewidth=1,edgecolor=colors[label],facecolor='none')\n            ax[0].add_patch(bbox)\n            ax[0].text(cmin, rmin, label, bbox=dict(fill=True, color=colors[label]))\n            ax[1].imshow(mask_decoded, alpha=0.3, cmap=cmaps[label])\n            ax[0].text(cmin, rmin, label, bbox=dict(fill=True, color=colors[label]))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Lets print out 10 random samples!"},{"metadata":{"trusted":true},"cell_type":"code","source":"for image_id, label_mask in grouped_EncodedPixels.sample(10).iteritems():\n    plot_cloud(train_image_path, image_id, label_mask)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Zooming In To The Cloud Formations\n\nIn the last section, we visualized many types of cloud formation at once. This visualization is good as it gives us a good high level picture of our training images. Also, excuse me for the terrible pun. However, in this section we will zoom into the different type of cloud formations by using the mask to mute out the section that doesn't belong to the type of cloud formation we are interested in. Our main question for this section is the following:\n\n* How do different types of cloud formations look like to the naked eye."},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_mask_cloud(img_path, img_id, label, mask):\n    img = cv2.imread(os.path.join(img_path, img_id), 0)\n    mask_decoded = rle_to_mask(mask, img.shape[0], img.shape[1])\n    mask_decoded = (mask_decoded > 0.0).astype(int)\n    img = np.multiply(img, mask_decoded)\n    return img","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def draw_label_only(label):\n    samples_df = train_df[(train_df['EncodedPixels']!=-1) & (train_df['Label']==label)].sample(2)\n    count = 0\n    fig, ax = plt.subplots(nrows=1, ncols=2, sharey=True, figsize=(20,10))\n    for idx, sample in samples_df.iterrows():\n        img = get_mask_cloud(train_image_path, sample['ImageId'], sample['Label'],sample['EncodedPixels'])\n        ax[count].imshow(img, cmap=\"gray\")\n        count += 1","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Fish Cloud Formation"},{"metadata":{"trusted":true},"cell_type":"code","source":"draw_label_only('Fish')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Flower Cloud Formation"},{"metadata":{"trusted":true},"cell_type":"code","source":"draw_label_only('Flower')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Gravel Cloud Formation"},{"metadata":{"trusted":true},"cell_type":"code","source":"draw_label_only('Gravel')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Sugar Cloud Formation"},{"metadata":{"trusted":true},"cell_type":"code","source":"draw_label_only('Sugar')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Well, I sort of get the sense why the cloud formations are called what they are called. What I am afraid of is often times, it seems like there is a little bit of cloud formation from one type overlaps with cloud formation of another type. I am interested to see how our algorithm performs in these scenarios."},{"metadata":{},"cell_type":"markdown","source":"## Cloud That Forms Together Stays Together\nAs we saw in the previous visualizations, it's very common for multiple cloud formations to be present in a single image. In this section we want to answer:\n1. Which cloud formations occur frequently.\n2. Which cloud formations hardly ever appears together.\n\nTo do this, we will use a really simple data mining algorithm called Frequent Pattern Mining. FP Growth is a particular algorithmic implementation of Frequent Pattern, which aims to identify items that appear frequently together in a list. "},{"metadata":{"trusted":true},"cell_type":"code","source":"# create a series with fault classes\nlabel_per_image = grouped_EncodedPixels.apply(lambda encoded_list: [x[0] for x in encoded_list if x[1] != -1])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# create a list of dict with count of each fault class\nlabel_per_image_list = []\nfor r in label_per_image.iteritems():\n    label_count = {'Fish':0,'Flower':0,'Gravel':0,'Sugar':0}\n    # go over each class and \n    for image_label in r[1]:\n        label_count[image_label] = 1\n    label_per_image_list.append(label_count)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# do FP calculation with all image\nlabel_per_image_df = pd.DataFrame(label_per_image_list)\nlabel_fp_df = fpgrowth(label_per_image_df, use_colnames=True, min_support=0.001)\nlabel_fp_df = label_fp_df.sort_values(by=['support'])\nlabel_combi_fp_df = label_fp_df[label_fp_df['itemsets'].apply(lambda x: len(x) > 1)]\nlabel_combi_fp_df['itemsets'] = label_combi_fp_df['itemsets'].apply(lambda x: ', '.join(x))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.bar(label_combi_fp_df, x=\"support\", y=\"itemsets\", orientation='h', \\\n            title='Frequent Patterns of The Cloud Formation')\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Sugar cloud formation frequently appears together in the images with Gravel or Fish cloud formation. Sugar also appears with Flower cloud formation but is less frequent. Gravels and Fish cloud formation also appears with other cloud formation. Sugar, Gravel, and Fish also appears all together in some instances.\n\nFlower tends to occur less frequently with other clouds, and the combination of Gravel and Flower occurs but at much less frequency compared to others. In fact, Sugar, Gravel, and Fish appear all together more frequently than Grave and Flower. However, it's not like Flower cloud formation never occurs with other cloud formation, just occurs less frequently compared to others.\n\nIn summary, they are all combination of cloud formations appearing together is a possibility, and the combinations between Sugar, Fish, and Gravel are more likely than with Flower cloud formation."},{"metadata":{},"cell_type":"markdown","source":"## Surface Area Ratio Per Cloud Formation\n\nIn this section we are interested in the following questions:\n1. How does the distribution of the surface area for different cloud formation masks look like?\n2. Are there any types of cloud formations that appear very frequently but have a very small surface area in our training dataset?"},{"metadata":{"trusted":true},"cell_type":"code","source":"# we will use the following function to decode our mask to binary and count the sum of the pixels for our mask.\ndef get_binary_mask_sum(encoded_mask):\n    mask_decoded = rle_to_mask(encoded_mask, width=2100, height=1400)\n    binary_mask = (mask_decoded > 0.0).astype(int)\n    return binary_mask.sum()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# calculate sum of the pixels for the mask per cloud formation\ntrain_df['mask_pixel_sum'] = train_df.apply(lambda x: get_binary_mask_sum(x['EncodedPixels']), axis=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# plot a histogram and boxplot combined of the mask pixel sum per cloud formation\nfig = px.histogram(train_df[train_df['mask_pixel_sum']!=0][['Label','mask_pixel_sum']], \n                   x=\"mask_pixel_sum\", y=\"Label\", color=\"Label\", marginal=\"box\")\n\nfig['layout'].update(title='Histogram and Boxplot of Sum of Mask Pixels Per Cloud Formation')\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Analysis to come soon!"},{"metadata":{},"cell_type":"markdown","source":"*Well, thats awesome! Now you can see what cloud formations that occur frequently together. I will be back soon with more updates, and good luck for the competition!!*"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}