{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Summary of notebook","metadata":{}},{"cell_type":"markdown","source":"**Purpose of EDA**<br>\n1. Explore label distribution<br>\n2. Show images for each label<br>\n3. Summarize what I understood<br>\n\n**Summary of EDA**\n1. Number of training images are 18632 and image size are not uniformed\n2. Main categories are bellow\n    1. Healthy\n    2. Scab\n    3. Frog_eye_leaf_spot\n    4. Rust\n    5. Powdery_mildew\n    6. *Mixed* (label \"complex\" or ones with multiple label)\n3. *Mixed* labels consists of 20% of total\n4. “complex” is leaves with too many diseases to classify (Mentioned in official discription)\n5. Labels \"frog_eye_leaf_spot\", \"rust\", \"powdery_mildew\" seems easy to identify\n6. “scab” seems difficult to identify spots<br>\n   -> Consider applying filter to emphasize scab  ","metadata":{}},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"import os\nimport glob\nimport math\nfrom PIL import Image\n\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read data","metadata":{}},{"cell_type":"code","source":"!ls ../input/plant-pathology-2021-fgvc8","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"BASE_DIR = \"../input/plant-pathology-2021-fgvc8\"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Training data image path and label\ntrain_image_path = glob.glob(os.path.join(BASE_DIR, \"train_images/*.jpg\"))\nlabel_df = pd.read_csv(os.path.join(BASE_DIR, \"train.csv\"))\n# Number od trainin images\nprint(\"Number of training images: {}\".format(len(train_image_path)))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show first 10 rows\nprint(train_image_path[:10])\ndisplay(label_df.head(10))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Label distribution","metadata":{}},{"cell_type":"code","source":"labels = label_df.labels.unique()\nprint(labels)\nprint(len(labels))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count number of data for each labels\nlabel_count = label_df.labels.value_counts()\nlabel_ratio = label_df.labels.value_counts(normalize=True, sort=True)\nprint(label_count)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show distribution for label by bar graph and pie chart \nfig, ax = plt.subplots(nrows=1, ncols=2, figsize=(20, 8))\nlabel_count.plot.bar(ax=ax[0], title=\"Number of data per label\", rot=90, fontsize=15)\nlabel_ratio.plot.pie(ax=ax[1], title=\"Ratio of label\", autopct='%1.1f%%', fontsize=15)\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Note from label distribution**\n1. Both healthy and scrab images consists of 25% of total data\n2. There are label with mix disease. (e.g. scab frog_eye_leaf_spot complex)","metadata":{}},{"cell_type":"markdown","source":"# Image properties","metadata":{}},{"cell_type":"markdown","source":"## Number of length of label.csv and images","metadata":{}},{"cell_type":"code","source":"# Get set of labels and image path\nlabels = set(label_df[\"image\"].to_list())\nimages = set(os.path.basename(full_path) for full_path in train_image_path)\n\nprint(labels ^ images)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Number of images in \"train.csv\" and image files are the same","metadata":{}},{"cell_type":"markdown","source":"## Image size","metadata":{}},{"cell_type":"code","source":"# Get image size\nimage_property = {\"image\": [], \"width\": [], \"height\": []}\nfor image_path in train_image_path:\n    im = Image.open(image_path)\n    width, height = im.size\n    file_name = os.path.basename(image_path)\n    image_property[\"image\"].append(file_name)\n    image_property[\"width\"].append(width)\n    image_property[\"height\"].append(height)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merge width, height info with label data\nimage_size = pd.DataFrame(image_property)\nimage_size = pd.merge(image_size, label_df, on=\"image\", how=\"outer\")\nimage_size","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Describe image property\nimage_size.info()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot (width, height) pair\nimage_size.plot(x=\"width\", y=\"height\", linestyle=\"none\", marker = \"x\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Count value pair for (width, height)\nimage_size.groupby([\"width\", \"height\"]).count()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- All the image do not have the same size","metadata":{}},{"cell_type":"markdown","source":"## Show images","metadata":{}},{"cell_type":"code","source":"labels = label_df.labels.value_counts()\ndisplay(labels)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display_num = 9\nfor label in labels.index:\n    df_tmp = label_df[label_df.labels == label]\n    images = df_tmp.sample(display_num).image.values\n    images = [os.path.join(BASE_DIR, \"train_images/\", file_name) for file_name in images]\n    \n    plt.figure(figsize=(20, 20))\n    for idx, image in enumerate(images):\n        plt.subplot(math.ceil(display_num / 3), 3, idx+1)\n        im = plt.imread(image)\n        plt.imshow(im)\n    print()\n    print(\"==================== Label:{} ===================\".format(label))\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observation from image**\n- Main categories are bellow\n    1. Healthy\n    2. Scab\n    3. Frog_eye_leaf_spot\n    4. Rust\n    5. Powdery_mildew\n    6. *Mixed* (label \"complex\" or ones with multiple label)\n- *Mixed* labels consists of 20% of total\n- “complex” is leaves with too many diseases to classify (Mentioned in official discription)\n- Labels \"frog_eye_leaf_spot\", \"rust\", \"powdery_mildew\" seems easy to identify\n- “scab” seems difficult to identify spots<br>\n   -> Consider applying filter to emphasize scab  ","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}