{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\n#for dirname, _, filenames in os.walk('/kaggle/input'):\n#    for filename in filenames:\n#        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Purpose of this note book\nThe purpose of this notebook is to explore data provide in a very simple way, so that kaggle beginner (like me) can understand the overview of competition. This notebook includes contents bellow.\n\n**Contents**<br>\n- Reading input data\n- Visualize & Analyze data\n- Feautures of cassava desease\n- Cassava classification quize to train your own nueral network!!"},{"metadata":{},"cell_type":"markdown","source":"# Imports"},{"metadata":{},"cell_type":"markdown","source":"Importing neccssary pakages for explore"},{"metadata":{"trusted":true},"cell_type":"code","source":"import glob\nimport json\nimport PIL\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport cv2","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Reading data"},{"metadata":{},"cell_type":"markdown","source":"We are going to read data bellow \n1. Json file to map desease name and label data\n2. Label csv file\n3. File path of train images"},{"metadata":{"trusted":true},"cell_type":"code","source":"# ID code and desease name mapping\nwith open(\"../input/cassava-leaf-disease-classification/label_num_to_disease_map.json\") as f:\n    desease_map = json.load(f)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Read label CSV \ntrain_label = pd.read_csv(\"../input/cassava-leaf-disease-classification/train.csv\")\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"# Get jpeg file path\ntrain_image_files = glob.glob(\"../input/cassava-leaf-disease-classification/train_images/*.jpg\")\nprint(len(train_image_files))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Visualize & Analyze data"},{"metadata":{},"cell_type":"markdown","source":"Here I am going to visualize data loaded at above cells. I explore label data first, and then image data."},{"metadata":{},"cell_type":"markdown","source":"## Explore label data"},{"metadata":{},"cell_type":"markdown","source":"Let's take a look at label data first to understand the disribution of data."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Show head and tails in label data.\ndisplay(train_label.head())\ndisplay(train_label.tail())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Looks like there is 21396 data. <br>\nEach row conists of \"file name\" and \"label id\". Let's find out what each label represents."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Here are the label (that is defined by competition) and its actual desease name.\ndisplay(desease_map)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"lable=4 indicates \"Healthy\" and others are desease. I am going to show the features for each desease at \"Feautures on cassava desease\" part in this notebook.<br>\n\nNext, let's get statisical overivew of label data."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Display summary info for label data\ndisplay(train_label.info())\ndisplay(train_label.describe())\nprint(train_label.isna().sum()) # Counting number of misiing values","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"From output above we can say...\n1. No misiing value in label data\n2. Max value for label is 4, which means undefined label is not included.\n3. Maen value for label is 2.65 so that number of data is little iclined to higher value.<br>\n   (We are going to see the distribution next for precise discussion)"},{"metadata":{},"cell_type":"markdown","source":"Next I am going to see the distribution of labels. This inspection is very important because, biased data whitout reason lead to bad model."},{"metadata":{"trusted":true},"cell_type":"code","source":"# See distribition of label data\ncount_per_label = train_label[\"label\"].value_counts(sort=False)\nratio_per_label = train_label[\"label\"].value_counts(sort=False, normalize=True)\ndesease_id = pd.DataFrame(desease_map.keys())\ndesease_name = pd.DataFrame(desease_map.values())\n\n# Show distribution for label by tabel\nprint(\"Number of data per label \")\ncomb = pd.concat([desease_id, desease_name, count_per_label, ratio_per_label], axis=1)\ncomb.columns = [\"Label id\", \"Desease name\", 'Number counts', 'Ratio']\ndisplay(comb)\n\n# Show distribution for label by bar graph and pie chart \nfig, ax = plt.subplots(nrows=1, ncols=2, figsize=(20, 8))\n# Plot Bar graph\ncomb.plot.bar(x=\"Label id\", y=\"Number counts\",ax=ax[0], title=\"Number of desease per id\", rot=0)\n# Plot Pie chart\ncomb.set_index(\"Desease name\").plot.pie(y=\"Ratio\", ax=ax[1], title=\"Ratio of desease\",autopct='%1.1f%%')\nplt.plot()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Bellow is the information that I got from above figures and numbers\n\n**Summary for label data**<br>\n- Key is image file name, Value is code to indentify disease.\n- 21397 datas with no missing value.\n- id=3('Cassava Mosaic Disease (CMD)') is the most common desease (60%), and number of the others are almost same (5-10%)\n- ID code and desease name mapping is below <br>\n  '0': 'Cassava Bacterial Blight (CBB)',<br>\n  '1': 'Cassava Brown Streak Disease (CBSD)',<br>\n  '2': 'Cassava Green Mottle (CGM)',<br>\n  '3': 'Cassava Mosaic Disease (CMD)',<br>\n  '4': 'Healthy'"},{"metadata":{},"cell_type":"markdown","source":"## Images"},{"metadata":{},"cell_type":"markdown","source":"### Image size"},{"metadata":{},"cell_type":"markdown","source":"First I am going to see the size (width and height) of an image."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Image size distribution\nimage_size = []\nfor path in train_image_files:\n    size = PIL.Image.open(path).size\n    image_size.append(size)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Create DataFrame to for image size\nimg_size_df = pd.DataFrame(image_size, columns=[\"Width\", \"Hight\"])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Plot image size distribution on 2D graph\nimg_size_df.plot(kind=\"scatter\", x=\"Width\", y=\"Hight\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Unique values in img_size_df\nprint(img_size_df.value_counts())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Looks like all the train images has the same size (width=800, height=600)! <br>\nThis means that we don't have to extract too big/small images before training."},{"metadata":{},"cell_type":"markdown","source":"## Show images"},{"metadata":{},"cell_type":"markdown","source":"Finally I am going to show some image per each desease label.<br>\nFirst I define functios."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Extract image file path with certain label\ndef pick_img_by_label(X, y, label, num=10):\n    label_i = X[y==label]\n    return label_i[np.random.choice(len(label_i), num, replace=False)]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Show image with certain label for image number specified. \ndef show_image_per_label(X, y, label, num=9, figsize=(20,20)):\n    X_label_i = pick_img_by_label(X, y, label, num)\n    rows, cols = np.ceil(num/3), 3\n    \n    plt.figure(figsize=figsize)\n    for j, file in enumerate(X_label_i):\n        img = cv2.imread(file)\n        plt.subplot(rows, cols, j+1)\n        plt.axis(\"off\")\n        plt.imshow(img)\n    plt.tight_layout()\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_label","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Add file path column\ntrain_label[\"file_path\"] = \"../input/cassava-leaf-disease-classification/train_images/\" + train_label[\"image_id\"]\n# Create X, y data\nX_files = train_label[\"file_path\"].values\ny = train_label[\"label\"].values\n\n# Show 10 images per label\nfor label in range(5):\n    print(\"Label: {}, Desease name: {}\".format(label, desease_map[str(label)]))\n    show_image_per_label(X_files, y, label, num=9, figsize=(10,10))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Feautures on cassava desease"},{"metadata":{},"cell_type":"markdown","source":"**Breif servey on each desease**<br>\n\nBellow are the breif surevey summary for each desease.\n\nLabel 0: Cassava Bacterial Blight (CBB)\n - Symptoms include blight, wilting, dieback, and vascular necrosis.<br>\n  (source:https://en.wikipedia.org/wiki/Bacterial_blight_of_cassava)\n\nLabel 1: Cassava Brown Streak Disease (CBSD)\n - CBSD is characterized by severe chlorosis and necrosis on infected leaves, giving them a yellowish, mottled appearance.\n - Symptons in leaf are sumillar to that of CMD, however CMD shows more apparent symptons.\n - CBSD is as serious as CMD in Africa, however CBSD is not found in Asia.<br>\n (sorce: <br>\n  - https://en.wikipedia.org/wiki/Cassava_brown_streak_virus_disease\n  - https://www.jica.go.jp/project/all_asia/005/materials/ku57pq000025s2lv-att/cassava_about.pdf)\n\nLable 2: Cassava Green Mottle (CGM)\n - Sympotons includes yellow spots, green mosiac patter and twisted margin.\n - \n(sorce: http://www.pestnet.org/fact_sheets/v6/cassava_green_mottle_068.htm)\n\nLable 3: Cassava Mosaic Disease (CMD)\n - CMD makes yellow or yellow-green spot on leaf.\n - Cassava Mosaic Virus is  virus that decreases crop yeild.<br>\n (source:https://en.wikipedia.org/wiki/Cassava_mosaic_virus)\n\n**Disscussion**\n- Data with label=3 (Cassava Mosaic Disease) has the most number might be because label=3 is a seirious desease for cassava in terms of yeild ammount.\n- Looks difficult to classify label 1, 2 and 3 as there sympton are simillar."},{"metadata":{},"cell_type":"markdown","source":"# Cassava desease quize"},{"metadata":{},"cell_type":"markdown","source":"This part is a quize for clasifying cassava desease by own nueral network! I hope this quize gives you an insight on features of each desease or images"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Define function to run quize\ndef cassava_quize(X=X_files, y=y, quize_num=5, figsize=(10,10)):\n    rand_idx = np.random.choice(len(X), quize_num, replace=False)\n    ans = [str(y[img_idx]) for img_idx in rand_idx]\n\n    nrows, ncols = int(np.ceil(quize_num / 3)), 3\n    plt.figure(figsize=figsize)\n    for i, img_idx in enumerate(rand_idx):\n        img = cv2.imread(X[img_idx])\n        plt.subplot(nrows, ncols, i+1)\n        plt.imshow(img)\n        plt.axis('off')\n        plt.title(\"Image {}\".format(i+1))\n    plt.tight_layout()\n    plt.show()\n\n    return ans\n\ndef your_answer(ans):\n    my_ans = []\n    print(\"Answer selections\")\n    for key, val in desease_map.items():\n        print(\"Label:{}, Name:{}\".format(key, val))\n    for j in range(len(ans)):\n        ans_j = input(\"Type your answer for Image {} \".format(j+1))\n        my_ans.append(ans_j)\n    print(\"Your answer: {}\".format(my_ans))\n    print(\"Correct answer: {}\".format(ans))\n    print(\"Results: {}/{}\".format(str(sum(np.array(my_ans)==np.array(ans))), str(len(ans))))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**How to play quize**\n1. Run first cell to show problem\n2. Run second cell to type your answer\n3. Look at the result and get insight for cassava desease or images!"},{"metadata":{"trusted":true},"cell_type":"code","source":"# 1st Step: Run here to show problem\nans = cassava_quize()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# 2nd-3rd Step: Run here to type your answer and check result!!\nyour_answer(ans)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}