{"cells":[{"metadata":{},"cell_type":"markdown","source":"## CASSAVA LEAF DISEASE DETECTION\n\n\nTask --> Your task is to classify each cassava image into four disease categories or a fifth category indicating a healthy leaf. \n\nThis competition is somewhat similar to SIIM - Melanoma classification hosted on kaggle previously this year.\nSo , if you didn't participated in that one , kindly have a look once and understand some of the concepts discussed in discussion forum.\n"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt \nimport os\nimport cv2","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Creating Variables for some paths to be used."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"input_dir = \"../input/cassava-leaf-disease-classification\"\n\ntrain_images_path = os.path.join(input_dir,\"train_images\")\ntest_images_path = os.path.join(input_dir,'test_images')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Loading Training Data"},{"metadata":{"trusted":true},"cell_type":"code","source":"train = pd.read_csv(os.path.join(input_dir,'train.csv'))\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Plotting a countplot that shows, some the distribution of all classes in the data provided in the data."},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.countplot(train.label)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"Total Number of Images in Training Data : \",train.shape[0])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Checking Class Mappings provided to each class in the train.csv file."},{"metadata":{"trusted":true},"cell_type":"code","source":"import json\n\nwith open(os.path.join(input_dir,\"label_num_to_disease_map.json\"),'r') as f:\n    class_mapping = json.load(f)\n\nclass_mapping","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Below i will be providing some plots showing some samples of images belonging to each class. \nThis will just provide an overview of how those images really are and may be hlpeful. "},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_samples(class_):\n    \n    print(f'Some Sample Images belonging to Class {class_mapping[f\"{class_}\"]}')\n    \n    sample_images = train[train.label == class_].sample(8)\n\n    fig,ax = plt.subplots(nrows=2,ncols=4,figsize=(16,8))\n\n    for e,img in enumerate(sample_images.image_id):\n        image_path = os.path.join(input_dir,f'train_images/{img}')\n        image = cv2.imread(image_path)\n        ax[e//4][e%4].imshow(image)\n    \n    plt.show()\n\nplot_samples(0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_samples(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_samples(2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_samples(3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_samples(4)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The way we load or read any image matters a lot on our model final results. So here i have a taken a sample image from train_images folder and shown 3 different reading types , namely BGR , RGB and HSV. You can clearly see the difference in these 3 different plots and examine the difference of how can our model can perform on such images."},{"metadata":{"trusted":true},"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras.preprocessing.image import *\nimage = cv2.imread(os.path.join(input_dir,\"train_images/1000015157.jpg\"))\n\nfig,ax = plt.subplots(nrows=1,ncols=3,figsize=(21,7))\n\nax[0].imshow(image)\nax[0].title.set_text(\"Original Image in BGR Format\")\n\nsample_image = cv2.cvtColor(image,cv2.COLOR_BGR2RGB)\nax[1].imshow(sample_image)\nax[1].title.set_text(\"Original Image in RGB Format\")\n\nsample_image = cv2.cvtColor(sample_image,cv2.COLOR_BGR2HSV)\nax[2].imshow(sample_image)\nax[2].title.set_text(\"Original Image in HSV Format\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig,ax = plt.subplots(nrows=1,ncols=3,figsize=(21,7))\nsns.distplot(image[:,:,0],ax=ax[0],color=\"b\")\nsns.distplot(image[:,:,1],ax=ax[1],color=\"g\")\nsns.distplot(image[:,:,2],ax=ax[2],color=\"r\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The above plots show distribution of pixel values of a sample image. Channel color is used as repective to the color of the plot."},{"metadata":{},"cell_type":"markdown","source":"This was a sample Starter notebook giving some EDA about the data. I will be also adding a sample baseline model with this notebook only.\n\nKindly provide our suggestions in the comment section. It would be very helpful in my learning.\nThank You !!!"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}