{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Introduction**","metadata":{"id":"e1pKGE2btd2k"}},{"cell_type":"markdown","source":"By Yvtsan Levy ","metadata":{"id":"4wNm97NGtlEI"}},{"cell_type":"markdown","source":"Misdiagnosis of the many diseases impacting agricultural crops can lead to misuse of chemicals leading to the emergence of resistant pathogen strains, increased input costs, and more outbreaks with significant economic loss and environmental impacts. Current disease diagnosis based on human scouting is time-consuming and expensive, and although computer-vision based models have the promise to increase efficiency, the great variance in symptoms due to age of infected tissues, genetic variations, and light conditions within trees decreases the accuracy of detection.\n\nThe purpose of the project is to diagnose apple tree diseases solely based on leaf images. The data is a set of images that is sapareted to those categories: \"healthy\", \"scab\", \"rust\", and \"multiple diseases\". Solving this problem is important because diagnosing plant diseases early can save tonnes of agricultural produce every year. This will benefit not only the general population by reducing hunger, but also the farmers by ensuring they get the harvest they deserve.\n\nThe project is based on the data set \"[Plant Pathology 2020 - FGVC7](https://www.kaggle.com/c/plant-pathology-2020-fgvc7/overview/description)\" Contains 3642 images of apple leaves divided into training data and test data.","metadata":{"id":"g3yLTNeus5RQ"}},{"cell_type":"markdown","source":"# **Domain knowledge**","metadata":{"id":"iLG7NpcTVQKY"}},{"cell_type":"markdown","source":"Global human population growth amounts to around 83 million annually, and as the population  grows, so does the global demand for food is rising. In order to meet the demand, agricultural productivity must increase. One of the factors for decrease in crop yield are plants diseases, and identifing them is one of the bottle necks in the process of treating them. Current disease diagnosis based on human scouting is time-consuming and expensive. Diagnosing plants diseases using computer-vision based models can improve the efficeincy of those processes and increace crop yield.\n\nApple tree are one of the most common fruit trees, and worldwide production of apples in 2018 was 86 million tonnes. 2 of the most common diseases of apple trees are apple scab and rust. Apple scab is a common disease of plants in the rose family (Rosaceae) that is caused by the ascomycete fungus Venturia inaequalis. Although apple scab rarely kills its host, infection typically leads to fruit deformation and premature leaf and fruit drop. The reduction of fruit quality and yield may result in crop losses of up to 70%. Rusts are plant diseases caused by pathogenic fungi of the order Pucciniales. Rusts are considered among the most harmful pathogens to agriculture, horticulture and forestry. Rust fungi are major concerns and limiting factors for successful cultivation of agricultural and forest crops.\n\nadd image of rust and scabs***\n","metadata":{"id":"a0pVo_WdT8cr"}},{"cell_type":"markdown","source":"# **Project's Target**","metadata":{"id":"OePBPP9eyv5O"}},{"cell_type":"markdown","source":"\nOur goal is to produce a model that classify correcly the health of a leaf and can say in great confidence what disease does it have. We will use images of apple leaves that are divided into 4 classes: healty leaves,leaves with scab,leaves with rust, and leaves with scab and rust. Our data-base contains 3642 images of apple leaves, 1821 classified images and 1821 images with no classification. We will try to train a model that can classify correcly images of apple leaves to those classes.","metadata":{"id":"lzv78dpB69Ea"}},{"cell_type":"markdown","source":"# **Install and Import Necessary Libraries**","metadata":{"id":"k1RVIoYM69Ec"}},{"cell_type":"markdown","source":"Importing the libraries that we will use","metadata":{"id":"hfsu46GUTCSZ"}},{"cell_type":"code","source":"","metadata":{"id":"1EQ_3Kb769Eh","outputId":"de597814-1a81-4cdb-9ffd-3fa3bfd55f6c","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install -q efficientnet\nimport os\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport cv2\nimport matplotlib.pyplot as plt\nimport matplotlib.pyplot\n%matplotlib inline\nimport sklearn\nimport sklearn.cluster\nimport plotly.graph_objects as go\nimport keras\nfrom sklearn.preprocessing import MultiLabelBinarizer\nimport tensorflow as tf\nimport tensorflow.keras.layers as L\nimport tensorflow.keras.backend as K\nimport imblearn.over_sampling\nimport warnings\nfrom tqdm.notebook import tqdm\ntqdm.pandas()\n\n#from tensorflow.keras.preprocessing import image\n\n#from tensorflow.keras.models import Model","metadata":{"id":"GRyJncMO69Ek","outputId":"88399709-5868-4e0f-eba3-68187804cac0","execution":{"iopub.status.busy":"2021-05-25T19:58:33.268588Z","iopub.execute_input":"2021-05-25T19:58:33.269095Z","iopub.status.idle":"2021-05-25T20:01:11.92995Z","shell.execute_reply.started":"2021-05-25T19:58:33.268974Z","shell.execute_reply":"2021-05-25T20:01:11.928741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#**Data Loading and Constant Defining** ","metadata":{"id":"ONWGIztv69Er"}},{"cell_type":"markdown","source":"Firstly we will load our data and define constants","metadata":{"id":"BUvm9c5y4m-o"}},{"cell_type":"code","source":"IMAGE_PATH = '../input/plant-pathology-2021-fgvc8/train_images/'\nTRAIN_PATH = '../input/plant-pathology-2021-fgvc8/train.csv'\nseed=1994\n\ntrain_data = pd.read_csv(TRAIN_PATH)","metadata":{"id":"cqQbG6Pe69Er","execution":{"iopub.status.busy":"2021-05-25T20:01:11.931845Z","iopub.execute_input":"2021-05-25T20:01:11.932281Z","iopub.status.idle":"2021-05-25T20:01:11.978558Z","shell.execute_reply.started":"2021-05-25T20:01:11.932233Z","shell.execute_reply":"2021-05-25T20:01:11.977433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#lets count the instances of each class we have :\n\nfig,ax=plt.subplots(figsize=(16,8))\nsns.countplot(train_data['labels'])\n#rotate labels\nplt.setp(ax.get_xticklabels(),rotation=45)\n\nplt.title('Label counts')","metadata":{"execution":{"iopub.status.busy":"2021-05-25T20:01:11.981082Z","iopub.execute_input":"2021-05-25T20:01:11.981554Z","iopub.status.idle":"2021-05-25T20:01:12.270858Z","shell.execute_reply.started":"2021-05-25T20:01:11.981507Z","shell.execute_reply":"2021-05-25T20:01:12.269108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#converting the labels as multiple labels:\ntrain_data['labels']=train_data['labels'].str.split(' ')\n\nmlb = MultiLabelBinarizer()\n\n# one hot encode labels\nlab=mlb.fit_transform(train_data['labels'])\nlab[:10]","metadata":{"execution":{"iopub.status.busy":"2021-05-25T20:01:12.272476Z","iopub.execute_input":"2021-05-25T20:01:12.272758Z","iopub.status.idle":"2021-05-25T20:01:12.313081Z","shell.execute_reply.started":"2021-05-25T20:01:12.272732Z","shell.execute_reply":"2021-05-25T20:01:12.312114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data[:10]","metadata":{"execution":{"iopub.status.busy":"2021-05-25T20:01:12.314352Z","iopub.execute_input":"2021-05-25T20:01:12.314607Z","iopub.status.idle":"2021-05-25T20:01:12.333616Z","shell.execute_reply.started":"2021-05-25T20:01:12.314583Z","shell.execute_reply":"2021-05-25T20:01:12.332564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#classes for OHE encoded var.\nclasses=mlb.classes_\nclasses","metadata":{"execution":{"iopub.status.busy":"2021-05-25T20:01:12.334933Z","iopub.execute_input":"2021-05-25T20:01:12.335297Z","iopub.status.idle":"2021-05-25T20:01:12.34077Z","shell.execute_reply.started":"2021-05-25T20:01:12.335263Z","shell.execute_reply":"2021-05-25T20:01:12.339599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data['image']","metadata":{"execution":{"iopub.status.busy":"2021-05-25T20:01:12.342098Z","iopub.execute_input":"2021-05-25T20:01:12.342397Z","iopub.status.idle":"2021-05-25T20:01:12.359249Z","shell.execute_reply.started":"2021-05-25T20:01:12.342368Z","shell.execute_reply":"2021-05-25T20:01:12.358107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DF = pd.DataFrame(\n    data=lab,\n    columns=classes\n)\n\nprint(DF.head(),DF.shape)","metadata":{"execution":{"iopub.status.busy":"2021-05-25T20:01:12.361592Z","iopub.execute_input":"2021-05-25T20:01:12.361941Z","iopub.status.idle":"2021-05-25T20:01:12.373534Z","shell.execute_reply.started":"2021-05-25T20:01:12.361911Z","shell.execute_reply":"2021-05-25T20:01:12.372019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DF['image'] =  train_data['image']\nprint(DF.head(),DF.shape)","metadata":{"execution":{"iopub.status.busy":"2021-05-25T20:01:12.375781Z","iopub.execute_input":"2021-05-25T20:01:12.37609Z","iopub.status.idle":"2021-05-25T20:01:12.388262Z","shell.execute_reply.started":"2021-05-25T20:01:12.376059Z","shell.execute_reply":"2021-05-25T20:01:12.386918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = DF","metadata":{"execution":{"iopub.status.busy":"2021-05-25T20:01:12.390017Z","iopub.execute_input":"2021-05-25T20:01:12.390519Z","iopub.status.idle":"2021-05-25T20:01:12.399408Z","shell.execute_reply.started":"2021-05-25T20:01:12.390469Z","shell.execute_reply":"2021-05-25T20:01:12.397935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#**Functions**","metadata":{"id":"3TSGQc-m69Ey"}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will use utility functions several times in our project, so we will define them here. Most of them are used to process images and reshape them or the data-sets that contins them","metadata":{"id":"kjZrYQeK5dsP"}},{"cell_type":"code","source":"#prints a image\n# gets an image to prints\n# returns nothing\ndef show_image(img):\n  fig = plt.imshow(img)\n\n#loads an image and returns it \n# gets the image id of the image we wants to load from the raw images and the shape of the output image\n# returns an image with the shape of image_size\ndef load_image(image_id,image_size=(100,100)):\n    file_path = image_id\n    image = cv2.imread(IMAGE_PATH + file_path,1)\n    image = cv2.resize(image, image_size)\n    image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n    return image\n\n#loads an image from returns it \n# gets the image id of the image we wants to load, the path to the dictunary and the shape of the output image\n# returns an image with the shape of image_size\ndef load_dif_image(image_id,path,image_size=(100,100)):\n    file_path = image_id \n    image = cv2.imread(path + file_path,1)\n    image = cv2.resize(image, image_size)\n    image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n    return image\n\n#creat from image objects list a reshaped 4-dimentional array(image_number,row,column,color_channel)\n# get a list of image objects\n# returns a 4-dimentional np array of the same images\ndef images_4d_array(files_list):\n  images_list = [img[np.newaxis, :, :, :3] for img in files_list]\n  images_array = np.vstack(images_list)\n  return (images_array)\n\n# retrive a list of images that fullfil the cond and show up to 9 pics of them. cond is a health situation name name. \n# gets health situation name:'scab','rust','multiple_diseases','healthy'\n# returns a list of images that fullfil the condition, and prints up to 9 images that fullfil them\ndef show_cond(cond):\n  cond_list=train_images[train_data[cond]==1]\n  cols, rows = 3,3  \n  fig, ax = plt.subplots(nrows=rows, ncols=cols, figsize=(15, rows*10/3))\n  for i in range(cols*rows):\n    if(i>=cond_list.size):\n      break\n    ax[int(i/cols), int(i%rows)].imshow(cond_list.iloc[i])\n  plt.show()\n  return (cond_list)\n\n# retrive a list of images that fullfil the cond. cond is a health situation name name. \n# gets health situation name:'scab','rust','multiple_diseases','healthy'\n# returns a list of images that fullfil the condition\ndef get_cond(cond):\n  cond_list=train_images[train_data[cond]==1] \n  return (cond_list)\n\n# calculate the avarge per color channel\n# gets a 4-dimensional  array and an int  \n# returns the avarage value of the cells by the 4th dimension channel position \ndef get_average_channel(x, channel):\n    return x[:,:,:,channel].mean(axis = (1, 2)).reshape(-1, 1)\n\n# claculate the avarge per color channel for x\n# gets image 4-dimentional array \n# returns array of \ndef get_channels(x):\n    return np.hstack([get_average_channel(x, i) for i in range(3)])\n\n# reshapes the 4-dimentional array into 2-dimentional array\n# gets a 4-dimentional array\n# returns a 2-dimentional array\ndef get_all_pixels(x):\n    return x.reshape(-1, np.prod(x.shape[1:]))\n\n# merges images to one image\n# gets an 4-dimentional array and the numbers of images to print as number per rows and number of images per row\n# returns an image that is a marge of the other ones\ndef merge_images(image_batch, size = [20, 20]):\n    h,w = image_batch.shape[1], image_batch.shape[2]\n    c = image_batch.shape[3]\n    img = np.zeros((int(h*size[0]), w*size[1], c))\n    for idx, im in enumerate(image_batch):\n        i = idx % size[1]\n        j = idx // size[1]\n        img[j*h:j*h+h, i*w:i*w+w,:] = im/255\n    return img\n\n# get the data and creates a numerated classes array\n# gets the images data\n# returns an array of numerated classes\ndef get_target_array(data):\n  targets=[]\n  for index, row in data.drop('image',axis=1).iterrows():\n    if (row[0]==1):\n      targets.append(0)\n    elif (row[1]==1):\n      targets.append(1)\n    elif (row[2]==1):\n      targets.append(2)\n    elif (row[3]==1):\n      targets.append(3)\n    elif (row[4]==1):\n      targets.append(4)\n    elif (row[5]==1):\n      targets.append(5)\n  return (np.array(targets))\n\n\n# get the data and creates a numerated classes array\n# gets the images data\n# returns an array of numerated classes\ndef get_target_array_no_im_id(data):\n  targets=[]\n  for index, row in data.iterrows():\n    if (row[0]==1):\n      targets.append(0)\n    elif (row[1]==1):\n      targets.append(1)\n    elif (row[2]==1):\n      targets.append(2)\n    elif (row[3]==1):\n      targets.append(3)\n    elif (row[4]==1):\n      targets.append(4)\n    elif (row[5]==1):\n      targets.append(5)\n  return (np.array(targets))","metadata":{"id":"SA5QGPcN69Ey","execution":{"iopub.status.busy":"2021-05-25T20:01:12.401289Z","iopub.execute_input":"2021-05-25T20:01:12.401729Z","iopub.status.idle":"2021-05-25T20:01:12.422584Z","shell.execute_reply.started":"2021-05-25T20:01:12.401679Z","shell.execute_reply":"2021-05-25T20:01:12.421701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Load Images**","metadata":{"id":"E8Ia42eWbegZ"}},{"cell_type":"markdown","source":"Thus far, we talk about how our images fit in our classes, but we haven't has a look at them, so we shall do it now.","metadata":{"id":"Qu-Qrih3begZ"}},{"cell_type":"code","source":"#loads all of the images\ntrain_images = train_data[\"image\"].progress_apply(load_image, args=((100,100),))\n\n#load a random sample\n#train_images = train_data[\"image_id\"].sample(n = n_sample, random_state = seed).progress_apply(load_image)\n\n#Build the dataset of current images\ncurrent_train_data=train_data.loc[train_images.index]\ntargets=get_target_array(current_train_data)","metadata":{"id":"MpLhGQ5DbegZ","outputId":"c39e833a-2caf-419f-f7b7-890f5bb1776a","execution":{"iopub.status.busy":"2021-05-25T20:01:12.423835Z","iopub.execute_input":"2021-05-25T20:01:12.424096Z","iopub.status.idle":"2021-05-25T20:50:21.932303Z","shell.execute_reply.started":"2021-05-25T20:01:12.424063Z","shell.execute_reply":"2021-05-25T20:50:21.930689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"current_train_data=train_data.loc[train_images.index]\ntargets=lab","metadata":{"execution":{"iopub.status.busy":"2021-05-25T20:50:21.934457Z","iopub.execute_input":"2021-05-25T20:50:21.934787Z","iopub.status.idle":"2021-05-25T20:50:21.941639Z","shell.execute_reply.started":"2021-05-25T20:50:21.934753Z","shell.execute_reply":"2021-05-25T20:50:21.940702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Data Exploration and Analysis**","metadata":{"id":"qOWZneep-xNo"}},{"cell_type":"markdown","source":"As a first step with our data, should explore our data and check what classes and data we have, the number of classes and the amount of images can tell us what dificulties can occur further in the project","metadata":{"id":"2ZYKFE2z_Uom"}},{"cell_type":"markdown","source":"Lets print our data of the images","metadata":{"id":"UEXgPpjrQKam"}},{"cell_type":"code","source":"print(train_data.head(),train_data.shape)","metadata":{"id":"gGSz36b169Ev","outputId":"c74d5405-06e7-4fc0-cb61-f3ef27d536d9","execution":{"iopub.status.busy":"2021-05-25T20:50:21.943069Z","iopub.execute_input":"2021-05-25T20:50:21.945457Z","iopub.status.idle":"2021-05-25T20:50:21.96563Z","shell.execute_reply.started":"2021-05-25T20:50:21.945401Z","shell.execute_reply":"2021-05-25T20:50:21.964296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen, we have a list of images and their class out of 4, healthy, multiple_diseases, rust and scab. Every entry is an image and have the value 1 for bieng in the class and 0 for not.\n","metadata":{"id":"yeEQo95UABEO"}},{"cell_type":"markdown","source":"We have 1821 images, lets find the distribution of the data in the classes.","metadata":{"id":"JQ44iChhQPic"}},{"cell_type":"code","source":"","metadata":{"id":"XJECIotUBbEB","outputId":"25595839-ae9b-40da-d416-cd350396cbcb","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen, 3 of the 4 classes have similar representation in our sata set. The multiple_diseases class is highly underrepresented, we have less then 100 images of this class, out of 1821 images. Our data is imbalanced, and we have to remember it as we draw conclusions","metadata":{"id":"4n5Rb2RmCwKm"}},{"cell_type":"markdown","source":"## **First Look at Our Images**","metadata":{"id":"1ot1NWOh69E0"}},{"cell_type":"markdown","source":"Thus far, we talk about how our images fit in our classes, but we haven't has a look at them, so we shall do it now.","metadata":{"id":"WsJTJSoY69E1"}},{"cell_type":"markdown","source":"First, lest see some leaves with scab","metadata":{"id":"78-FkI2aUZ2i"}},{"cell_type":"code","source":"scab_list=show_cond('scab') ","metadata":{"id":"5DkLVk3E69E4","execution":{"iopub.status.busy":"2021-05-25T20:50:21.967012Z","iopub.execute_input":"2021-05-25T20:50:21.967293Z","iopub.status.idle":"2021-05-25T20:50:22.985849Z","shell.execute_reply.started":"2021-05-25T20:50:21.967265Z","shell.execute_reply":"2021-05-25T20:50:22.985102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen from the images, scabs manifest itself as dark brown spots on the leaves. one can easily think that those are just dirty leaves, but a professional will know that they have scab.","metadata":{"id":"mDhU8TLIUoRe"}},{"cell_type":"markdown","source":"Lest see some leaves with rust","metadata":{"id":"15P_uWnuUmtx"}},{"cell_type":"code","source":"rust_list=show_cond('rust')","metadata":{"id":"pM96HfCb69E6","outputId":"bd1311be-e5b8-411f-a531-948fcc952200","execution":{"iopub.status.busy":"2021-05-25T20:50:22.986934Z","iopub.execute_input":"2021-05-25T20:50:22.987332Z","iopub.status.idle":"2021-05-25T20:50:23.891462Z","shell.execute_reply.started":"2021-05-25T20:50:22.987298Z","shell.execute_reply":"2021-05-25T20:50:23.890676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen from the images, rust manifest itself as yellow-brownish spots on the leaves. The spots are much more identifiable and noticeable.","metadata":{"id":"D4KAm_4-V1kD"}},{"cell_type":"markdown","source":"Lest see some leaves with multiple diseases","metadata":{"id":"w6a3q7XxWpyf"}},{"cell_type":"code","source":"multi_list=show_cond('complex')","metadata":{"id":"Ezf-ewY_69E-","outputId":"4358b52c-efc8-4061-b43b-67c04123021e","execution":{"iopub.status.busy":"2021-05-25T20:50:23.892519Z","iopub.execute_input":"2021-05-25T20:50:23.892896Z","iopub.status.idle":"2021-05-25T20:50:24.820133Z","shell.execute_reply.started":"2021-05-25T20:50:23.892866Z","shell.execute_reply":"2021-05-25T20:50:24.819214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen from the images, those leaves have features of both diseases.","metadata":{"id":"hrD9uTv3WzD1"}},{"cell_type":"markdown","source":"Lets see some healthy leaves","metadata":{"id":"cWaEhGGlZHMn"}},{"cell_type":"code","source":"healthy_list=show_cond('healthy')","metadata":{"id":"-LCV2wLQ69FC","outputId":"71c1aa81-b604-4c27-83ca-de60f85a98d7","execution":{"iopub.status.busy":"2021-05-25T20:50:24.82149Z","iopub.execute_input":"2021-05-25T20:50:24.821781Z","iopub.status.idle":"2021-05-25T20:50:26.03015Z","shell.execute_reply.started":"2021-05-25T20:50:24.821749Z","shell.execute_reply":"2021-05-25T20:50:26.029124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"AS we expected the leaves has no features of the diseases","metadata":{"id":"RaV9moOqZUF8"}},{"cell_type":"markdown","source":"## **Color Analysis**","metadata":{"id":"FIDjsf6h69FE"}},{"cell_type":"markdown","source":"Colors in pictures are the basic features, thus analyising them is our first step. We already saw that the dieases has characteristics that are connected to color, so it is naturally a good direction.","metadata":{"id":"Ao_xZOdOZoFQ"}},{"cell_type":"markdown","source":"Changing the images array to 4 dimentional (image number, row, column, color channel) array so we could analyze them ","metadata":{"id":"_nnKo-z969FF"}},{"cell_type":"code","source":"healthy_array=images_4d_array(healthy_list)\nscab_array=images_4d_array(scab_list)\nmulti_array=images_4d_array(multi_list)\nrust_array=images_4d_array(rust_list)\n\nprint(healthy_array.shape,scab_array.shape,multi_array.shape,rust_array.shape)","metadata":{"id":"QyfSirIG69FF","outputId":"58a151df-0403-4672-81e4-bd63825f16e6","execution":{"iopub.status.busy":"2021-05-25T20:50:26.031506Z","iopub.execute_input":"2021-05-25T20:50:26.031992Z","iopub.status.idle":"2021-05-25T20:50:26.692708Z","shell.execute_reply.started":"2021-05-25T20:50:26.031954Z","shell.execute_reply":"2021-05-25T20:50:26.691705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" Lets plot the color histograms for each class","metadata":{"id":"etWdC--S69FI"}},{"cell_type":"code","source":"#print histograms of colors\ndef plot_hist_normed(images, channel, col):\n    vals = images[:,:,:,channel].flatten()\n    matplotlib.pyplot.ylabel(col)\n    matplotlib.pyplot.hist(vals)\n    matplotlib.pyplot.yticks([])\n\n\n#red  \nmatplotlib.pyplot.figure(figsize =(30, 15))\nmatplotlib.pyplot.subplot(3, 4, 1)\nplot_hist_normed(healthy_array, 0, 'red')\nmatplotlib.pyplot.title('healthy leaves')\nmatplotlib.pyplot.subplot(3, 4, 2)\nplot_hist_normed(scab_array, 0, 'red')\nmatplotlib.pyplot.title('scab leaves')\nmatplotlib.pyplot.subplot(3, 4, 3)\nplot_hist_normed(rust_array, 0, 'red')\nmatplotlib.pyplot.title('rust leaves')\nmatplotlib.pyplot.subplot(3, 4, 4)\nplot_hist_normed(multi_array, 0, 'red')\nmatplotlib.pyplot.title('multiple diseases leaves')\n\n#green\nmatplotlib.pyplot.subplot(3, 4, 5)\nplot_hist_normed(healthy_array, 1, 'green')\nmatplotlib.pyplot.title('healthy leaves')\nmatplotlib.pyplot.subplot(3, 4, 6)\nplot_hist_normed(scab_array, 1, 'green')\nmatplotlib.pyplot.title('scab leaves')\nmatplotlib.pyplot.subplot(3, 4, 7)\nplot_hist_normed(rust_array, 1, 'green')\nmatplotlib.pyplot.title('rust leaves')\nmatplotlib.pyplot.subplot(3, 4, 8)\nplot_hist_normed(multi_array, 1, 'green')\nmatplotlib.pyplot.title('multiple diseases leaves')\n\n#blue\nmatplotlib.pyplot.subplot(3, 4, 9)\nplot_hist_normed(healthy_array, 2, 'blue')\nmatplotlib.pyplot.title('healthy leaves')\nmatplotlib.pyplot.subplot(3, 4, 10)\nplot_hist_normed(scab_array, 2, 'blue')\nmatplotlib.pyplot.title('scab leaves')\nmatplotlib.pyplot.subplot(3, 4, 11)\nplot_hist_normed(rust_array, 2, 'blue')\nmatplotlib.pyplot.title('rust leaves')\nmatplotlib.pyplot.subplot(3, 4, 12)\nplot_hist_normed(multi_array, 2, 'blue')\nmatplotlib.pyplot.title('multiple diseases leaves')\n\n\nmatplotlib.pyplot.show()","metadata":{"id":"r8xjRhRi69FI","outputId":"110fd0eb-5880-4dee-c5de-60f5f32c54d3","execution":{"iopub.status.busy":"2021-05-25T20:50:26.693996Z","iopub.execute_input":"2021-05-25T20:50:26.694295Z","iopub.status.idle":"2021-05-25T20:50:34.48786Z","shell.execute_reply.started":"2021-05-25T20:50:26.694263Z","shell.execute_reply":"2021-05-25T20:50:34.486822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is not very informative, it is hard to notice the differences as the images are mostly similar. maybe we sould check the median and the mean of each channel","metadata":{"id":"7mnSa0S8gtqV"}},{"cell_type":"markdown","source":"Lest plot the mean and median values for the classes","metadata":{"id":"SzKW1Hpe69FK"}},{"cell_type":"code","source":"def summary(images, channel, col):\n    vals = images[:,:,:,channel].flatten()\n    chan_mean = np.mean(vals)\n    chan_median = np.median(vals)\n    print('{} mean: {}, median: {}'.format(col, str(chan_mean), str(chan_median)))\n\nprint('healthy leaves:')\nsummary(healthy_array, 0, 'red')\nsummary(healthy_array, 1, 'green')\nsummary(healthy_array, 2, 'blue')\nprint('scab leaves:')\nsummary(scab_array, 0, 'red')\nsummary(scab_array, 1, 'green')\nsummary(scab_array, 2, 'blue')\nprint('rust leaves:')\nsummary(rust_array, 0, 'red')\nsummary(rust_array, 1, 'green')\nsummary(rust_array, 2, 'blue')\nprint('multiple diseases leaves:')\nsummary(multi_array, 0, 'red')\nsummary(multi_array, 1, 'green')\nsummary(multi_array, 2, 'blue')\n","metadata":{"id":"rJmLsYDM69FL","outputId":"4e8b7f57-7fdf-4d6f-ee22-5ce92669c2c8","execution":{"iopub.status.busy":"2021-05-25T20:50:34.489126Z","iopub.execute_input":"2021-05-25T20:50:34.489404Z","iopub.status.idle":"2021-05-25T20:50:38.020328Z","shell.execute_reply.started":"2021-05-25T20:50:34.489376Z","shell.execute_reply":"2021-05-25T20:50:38.019147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen the blue channels mean and median of the diseased leaves are lower, and the red channels mean and median are higher then the healthy ones. We should explore the difference further by comparing the healty and diseased leaves pictures","metadata":{"id":"47PTATpA69FN"}},{"cell_type":"code","source":"diseased_leaves_array=np.concatenate((rust_array,multi_array,scab_array))\ndiseased_leaves_array.shape","metadata":{"id":"j2ljA_bK69FN","outputId":"a9afe2ee-a00d-4848-da5c-b4d4494d226e","execution":{"iopub.status.busy":"2021-05-25T20:50:38.0218Z","iopub.execute_input":"2021-05-25T20:50:38.022206Z","iopub.status.idle":"2021-05-25T20:50:38.299787Z","shell.execute_reply.started":"2021-05-25T20:50:38.022145Z","shell.execute_reply":"2021-05-25T20:50:38.298742Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#red  \nmatplotlib.pyplot.figure(figsize =(30, 15))\nmatplotlib.pyplot.subplot(3, 2, 1)\nplot_hist_normed(healthy_array, 0, 'red')\nmatplotlib.pyplot.title('healthy leaves')\nmatplotlib.pyplot.subplot(3, 2, 2)\nplot_hist_normed(diseased_leaves_array, 0, 'red')\nmatplotlib.pyplot.title('diseased leaves')\n\n#green\nmatplotlib.pyplot.subplot(3, 2, 3)\nplot_hist_normed(healthy_array, 1, 'green')\nmatplotlib.pyplot.title('healthy leaves')\nmatplotlib.pyplot.subplot(3, 2, 4)\nplot_hist_normed(diseased_leaves_array, 1, 'green')\nmatplotlib.pyplot.title('diseased leaves')\n\n#blue\nmatplotlib.pyplot.subplot(3, 2, 5)\nplot_hist_normed(healthy_array, 2, 'blue')\nmatplotlib.pyplot.title('healthy leaves')\nmatplotlib.pyplot.subplot(3, 2, 6)\nplot_hist_normed(diseased_leaves_array, 2, 'blue')\nmatplotlib.pyplot.title('diseased leaves')\n\nmatplotlib.pyplot.show()\n\n\nprint('healthy leaves:')\nsummary(healthy_array, 0, 'red')\nsummary(healthy_array, 1, 'green')\nsummary(healthy_array, 2, 'blue')\nprint('diseased leaves:')\nsummary(diseased_leaves_array, 0, 'red')\nsummary(diseased_leaves_array, 1, 'green')\nsummary(diseased_leaves_array, 2, 'blue')","metadata":{"id":"nIxW2C6j69FQ","outputId":"32477521-308a-44de-b7d9-057d327881a7","execution":{"iopub.status.busy":"2021-05-25T20:50:38.305579Z","iopub.execute_input":"2021-05-25T20:50:38.305892Z","iopub.status.idle":"2021-05-25T20:50:49.608416Z","shell.execute_reply.started":"2021-05-25T20:50:38.305863Z","shell.execute_reply":"2021-05-25T20:50:49.607436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we expected the blue values are lower and the red values are higher. We can see that the diseased leaves have brown or yellow spots on them. That might be the reason for the discrepancy as in those colors the blue channel has a lower value and the red color has a higher value, and if that is the reason, those differences might help us to identify the healthy leaves. This hypothesis is bold and in order to try and see if it has any truth in it more exploration is needed\n","metadata":{"id":"XG-KeL1l69FS"}},{"cell_type":"markdown","source":"## **PCA**","metadata":{"id":"3ICkr3mv69FS"}},{"cell_type":"markdown","source":"We want to further explore the main difference between the classes, and to try and check our hypothesis. Principal component analysis (PCA) is an effective way to do so. PCA is a method to reduce the dimension of our features, and find the most dominant ones ","metadata":{"id":"M-D0JbI169FT"}},{"cell_type":"markdown","source":"First, we will arrange the data so it will be easier to analyse it","metadata":{"id":"ko7id8gDiMdt"}},{"cell_type":"code","source":"images_array=images_4d_array(train_images)\nimages_array_all = get_all_pixels(images_array)","metadata":{"id":"5i_TTzvG69FT","execution":{"iopub.status.busy":"2021-05-25T20:50:49.610704Z","iopub.execute_input":"2021-05-25T20:50:49.610991Z","iopub.status.idle":"2021-05-25T20:50:50.209934Z","shell.execute_reply.started":"2021-05-25T20:50:49.610958Z","shell.execute_reply":"2021-05-25T20:50:50.208824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" Lets Compute the first 10 PCs and the projection on them","metadata":{"id":"inl4de2S69FV"}},{"cell_type":"code","source":"images_array_all_centered=sklearn.preprocessing.StandardScaler().fit_transform(images_array_all)\n\npca = sklearn.decomposition.PCA(n_components = 10)\npca.fit(images_array_all_centered)\nx_train_reduced = pca.transform(images_array_all_centered)\nprincipalDf=pd.DataFrame(data=x_train_reduced,columns=['PC1','PC2','PC3','PC4','PC5','PC6','PC7','PC8','PC9','PC10'] )\nprincipalDf.head()","metadata":{"id":"toYeuic969FW","outputId":"e2507a26-be89-4dcf-dded-cb844e49bc17","execution":{"iopub.status.busy":"2021-05-25T20:50:50.211606Z","iopub.execute_input":"2021-05-25T20:50:50.212169Z","iopub.status.idle":"2021-05-25T20:51:54.873483Z","shell.execute_reply.started":"2021-05-25T20:50:50.212123Z","shell.execute_reply":"2021-05-25T20:51:54.872512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we have the first 10 PCs we can check how much each of them discribes the variance of the data, by showing the percentage of variance explained by each of the selected components","metadata":{"id":"D9vE93RGldym"}},{"cell_type":"code","source":"pca.explained_variance_ratio_\n\n\ndf1 = pd.DataFrame({'var':pca.explained_variance_ratio_,\n             'PC':['PC1','PC2','PC3','PC4','PC5','PC6','PC7','PC8','PC9','PC10']})\nsns.barplot(x='PC',y=\"var\", \n           data=df1, color=\"c\");","metadata":{"id":"UkUjVp6aujSp","outputId":"c14e4267-63de-44ae-ec08-8789091a84b1","execution":{"iopub.status.busy":"2021-05-25T20:51:54.87499Z","iopub.execute_input":"2021-05-25T20:51:54.875543Z","iopub.status.idle":"2021-05-25T20:51:55.071611Z","shell.execute_reply.started":"2021-05-25T20:51:54.8755Z","shell.execute_reply":"2021-05-25T20:51:55.069342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen, the first 2 PCs explaies about 18 precent of the varience, mabye we can coralate it to our classes","metadata":{"id":"WBh8D9morsCg"}},{"cell_type":"markdown","source":"Lets plot the data on the first 2 PCs, this can visualize easily clusters of data in on those axis if they exist","metadata":{"id":"53-50oSU69FY"}},{"cell_type":"code","source":"matplotlib.pyplot.xlabel('PC 1')\nmatplotlib.pyplot.ylabel('PC 2')\nmatplotlib.pyplot.title('2D PCA')\nmatplotlib.pyplot.scatter(x_train_reduced[:, 0], x_train_reduced[:, 1],c = targets, cmap=matplotlib.pyplot.cm.get_cmap('YlGnBu', 10))\nmatplotlib.pyplot.colorbar();","metadata":{"id":"sTl-AEDW69FY","outputId":"4bcdd4f5-7fec-4ee0-e74b-b696bef6ffeb","execution":{"iopub.status.busy":"2021-05-25T20:51:55.073409Z","iopub.execute_input":"2021-05-25T20:51:55.073766Z","iopub.status.idle":"2021-05-25T20:51:55.578451Z","shell.execute_reply.started":"2021-05-25T20:51:55.073734Z","shell.execute_reply":"2021-05-25T20:51:55.571384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The dispersal of data does not correspond with 4 classes or some kind of clusters, which suggests that the PCs does not discribe the difference in the classes. \n","metadata":{"id":"8gMC0cLs69Fa"}},{"cell_type":"markdown","source":"Since we're dealing with images the best way to get what a given PC's \"subject\", what it discribes, is  simply to view those images which have a high or low score for this PC. We will show the first 4 PCs, because each of the rest represents less then 4 precent of the variance, which is very low","metadata":{"id":"B4a9DJjGlCvv"}},{"cell_type":"markdown","source":"### **First PC**","metadata":{"id":"kqRam5_F69Fb"}},{"cell_type":"markdown","source":"First we should find the top and bottom 16 images in the PC","metadata":{"id":"PxXh2YuP69Fb"}},{"cell_type":"code","source":"highest_score_ids = np.argpartition(x_train_reduced[:, 0], -16)[-16:]\nlowest_score_ids = np.argpartition(x_train_reduced[:, 0], 16)[:16]","metadata":{"id":"qe9-NZia69Fc","execution":{"iopub.status.busy":"2021-05-25T20:51:55.579388Z","iopub.status.idle":"2021-05-25T20:51:55.579766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we will print the classification of them, maybe the this could tell us something","metadata":{"id":"t_JUTErd69Ff"}},{"cell_type":"code","source":"print(current_train_data.iloc[highest_score_ids])\nprint(current_train_data.iloc[lowest_score_ids])","metadata":{"id":"BW5SFuWC69Ff","outputId":"0152be1e-478e-41e7-8938-89f2b17f22e1","execution":{"iopub.status.busy":"2021-05-25T20:51:55.583221Z","iopub.status.idle":"2021-05-25T20:51:55.583815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There isn't a clear difference betweet the groups","metadata":{"id":"CafL1A5V69Fr"}},{"cell_type":"markdown","source":"Lets print the images and check if the difference can be identify","metadata":{"id":"XnZF-H4I69Fs"}},{"cell_type":"code","source":"highest_images_4d=images_array[highest_score_ids]\nlowest_images_4d=images_array[lowest_score_ids]\n\nhighest_images_merged = merge_images(highest_images_4d, size = [4, 4])\nlowest_images_merged = merge_images(lowest_images_4d, size = [4, 4])\n\nmatplotlib.pyplot.figure(figsize=(10,5))\nmatplotlib.pyplot.subplot(1, 2, 1)\nmatplotlib.pyplot.title('Highest PC1 Score')\nmatplotlib.pyplot.axis('off')\nmatplotlib.pyplot.imshow(highest_images_merged)\n\nmatplotlib.pyplot.subplot(1, 2, 2)\nmatplotlib.pyplot.title('Lowest PC1 Score')\nmatplotlib.pyplot.axis('off')\nmatplotlib.pyplot.imshow(lowest_images_merged)","metadata":{"id":"szCHOfWl69Fu","outputId":"f58e706e-a1cb-4356-d049-bdfcade0c4fa","execution":{"iopub.status.busy":"2021-05-25T20:51:55.585541Z","iopub.status.idle":"2021-05-25T20:51:55.586136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the images we can infer that the first PC is about the background color. The highest images have brighter brackgraoud, and the lowest have darker background ","metadata":{"id":"zRTwYM5I69F0"}},{"cell_type":"markdown","source":"### **Second PC**\n\n\n","metadata":{"id":"PYtz4piB69F1"}},{"cell_type":"markdown","source":"We will repeat the process","metadata":{"id":"ILOMSX0069F2"}},{"cell_type":"code","source":"highest_score_ids = np.argpartition(x_train_reduced[:, 1], -16)[-16:]\nlowest_score_ids = np.argpartition(x_train_reduced[:, 1], 16)[:16]\nprint(current_train_data.iloc[highest_score_ids])\nprint(current_train_data.iloc[lowest_score_ids])\nhighest_images_4d=images_array[highest_score_ids]\nlowest_images_4d=images_array[lowest_score_ids]\n\nhighest_images_merged = merge_images(highest_images_4d, size = [4, 4])\nlowest_images_merged = merge_images(lowest_images_4d, size = [4, 4])\n\nmatplotlib.pyplot.figure(figsize=(10,5))\nmatplotlib.pyplot.subplot(1, 2, 1)\nmatplotlib.pyplot.title('Highest PC2 Score')\nmatplotlib.pyplot.axis('off')\nmatplotlib.pyplot.imshow(highest_images_merged)\n\nmatplotlib.pyplot.subplot(1, 2, 2)\nmatplotlib.pyplot.title('Lowest PC2 Score')\nmatplotlib.pyplot.axis('off')\nmatplotlib.pyplot.imshow(lowest_images_merged)","metadata":{"id":"8rdcVFLn69F4","outputId":"6c37418c-55c1-44d4-8ba8-712d20afc068","execution":{"iopub.status.busy":"2021-05-25T20:51:55.587432Z","iopub.status.idle":"2021-05-25T20:51:55.58798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There isn't a clear difference between the groups' classes. From the imags we can infer that the difference is again in the background's shape and lighting changes","metadata":{"id":"0FNksHRj69F-"}},{"cell_type":"markdown","source":"### **Third PC**\n\n\n","metadata":{"id":"qZGJsYkt69F_"}},{"cell_type":"markdown","source":"We will repeat the process","metadata":{"id":"PkoEZDHK69GA"}},{"cell_type":"code","source":"highest_score_ids = np.argpartition(x_train_reduced[:, 2], -16)[-16:]\nlowest_score_ids = np.argpartition(x_train_reduced[:, 2], 16)[:16]\nprint(current_train_data.iloc[highest_score_ids])\nprint(current_train_data.iloc[lowest_score_ids])\nhighest_images_4d=images_array[highest_score_ids]\nlowest_images_4d=images_array[lowest_score_ids]\n\nhighest_images_merged = merge_images(highest_images_4d, size = [4, 4])\nlowest_images_merged = merge_images(lowest_images_4d, size = [4, 4])\n\nmatplotlib.pyplot.figure(figsize=(10,5))\nmatplotlib.pyplot.subplot(1, 2, 1)\nmatplotlib.pyplot.title('Highest PC3 Score')\nmatplotlib.pyplot.axis('off')\nmatplotlib.pyplot.imshow(highest_images_merged)\n\nmatplotlib.pyplot.subplot(1, 2, 2)\nmatplotlib.pyplot.title('Lowest PC3 Score')\nmatplotlib.pyplot.axis('off')\nmatplotlib.pyplot.imshow(lowest_images_merged)","metadata":{"id":"nnTy0Yca69GB","outputId":"a9026e39-118e-4735-f526-1cb603d079ab","execution":{"iopub.status.busy":"2021-05-25T20:51:55.589342Z","iopub.status.idle":"2021-05-25T20:51:55.589902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There isn't a clear difference between the groups' classes. It is hard to infer from the images what features are promenent in this PC.","metadata":{"id":"k8dUqEG569GH"}},{"cell_type":"markdown","source":"### **Forth PC**\n\n\n","metadata":{"id":"pxlDPnas69GK"}},{"cell_type":"markdown","source":"We will repeat the process","metadata":{"id":"-aOcHyhK69GM"}},{"cell_type":"code","source":"highest_score_ids = np.argpartition(x_train_reduced[:, 3], -16)[-16:]\nlowest_score_ids = np.argpartition(x_train_reduced[:, 3], 16)[:16]\nprint(current_train_data.iloc[highest_score_ids])\nprint(current_train_data.iloc[lowest_score_ids])\nhighest_images_4d=images_array[highest_score_ids]\nlowest_images_4d=images_array[lowest_score_ids]\n\nhighest_images_merged = merge_images(highest_images_4d, size = [4, 4])\nlowest_images_merged = merge_images(lowest_images_4d, size = [4, 4])\n\nmatplotlib.pyplot.figure(figsize=(10,5))\nmatplotlib.pyplot.subplot(1, 2, 1)\nmatplotlib.pyplot.title('Highest PC2 Score')\nmatplotlib.pyplot.axis('off')\nmatplotlib.pyplot.imshow(highest_images_merged)\n\nmatplotlib.pyplot.subplot(1, 2, 2)\nmatplotlib.pyplot.title('Lowest PC2 Score')\nmatplotlib.pyplot.axis('off')\nmatplotlib.pyplot.imshow(lowest_images_merged)","metadata":{"id":"l-iArP0o69GN","outputId":"6147119b-bb68-4f68-b1c2-d2fba01d8ccc","execution":{"iopub.status.busy":"2021-05-25T20:51:55.59112Z","iopub.status.idle":"2021-05-25T20:51:55.591701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There isn't a clear difference between the groups' classes. From the imags we can infer that the difference is in the color of the images, the highest group is has darker and less yellowish leaves and the lowest group has bighterand yellowish leaves which is the result of the sorce of lighting and not the different classes","metadata":{"id":"KGKZa5KL69GU"}},{"cell_type":"markdown","source":"## **Image managing**","metadata":{"id":"pTtmDJ9UkzZX"}},{"cell_type":"markdown","source":"### **Cropping the Leaf Out**","metadata":{"id":"GIB43yaJsNuA"}},{"cell_type":"markdown","source":"In the last section we witnessed that the the background is a big contributer to the varience in the feature space. In order to deal with it we will try to crop the leaf out of the image and make a new imge out of it, this should reduce the background contribution","metadata":{"id":"F7kcxHkljnXg"}},{"cell_type":"code","source":"targets=get_target_array(train_data)","metadata":{"id":"IWenbIRrXxBO","execution":{"iopub.status.busy":"2021-05-25T20:51:55.592978Z","iopub.status.idle":"2021-05-25T20:51:55.593566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **creating the new image data set**","metadata":{"id":"_v-9wjRQog_B"}},{"cell_type":"markdown","source":"We have too much images for manualy determine how to crop them, so we sholud try and automatize it. We will first try to identify each leaf in their respected image. Our assumption is that the leaf is the main object in the image, and thus it's conture is the biggest. We will try to find the biggest object from it contures and try to crop a rectange around it.","metadata":{"id":"2A3gSZH4mIGc"}},{"cell_type":"markdown","source":"Finding the contours","metadata":{"id":"keyP-ygLoGqv"}},{"cell_type":"code","source":"# convert to grayscale\ngray_train_images = [cv2.cvtColor(x, cv2.COLOR_BGR2GRAY) for x in train_images] \n# threshold to get just the signature (INVERTED)\nthresh_gray = [cv2.threshold(x, 100, maxval=255, type=cv2.THRESH_BINARY_INV)[1] for x in gray_train_images]\ncontours = [cv2.findContours(x,cv2.RETR_LIST,cv2.CHAIN_APPROX_SIMPLE)[0] for x in thresh_gray]","metadata":{"id":"-Tz10K73sLMW","execution":{"iopub.status.busy":"2021-05-25T20:51:55.594678Z","iopub.status.idle":"2021-05-25T20:51:55.595267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A function that crops a rectangle out of a image","metadata":{"id":"2CVLVI_goKpE"}},{"cell_type":"code","source":"def crop_minAreaRect(img, rect):\n    # Source: https://stackoverflow.com/questions/37177811/\n\n    # rotate img\n    angle = rect[2]\n    rows,cols = img.shape[0], img.shape[1]\n    matrix = cv2.getRotationMatrix2D((cols/2,rows/2),angle,1)\n    img_rot = cv2.warpAffine(img,matrix,(cols,rows))\n\n    # rotate bounding box\n    rect0 = (rect[0], rect[1], 0.0)\n    box = cv2.boxPoints(rect)\n    pts = np.int0(cv2.transform(np.array([box]), matrix))[0]\n    pts[pts < 0] = 0\n\n    # crop and return\n    return img_rot[pts[1][1]:pts[0][1], pts[1][0]:pts[2][0]]\n\n","metadata":{"id":"B18SAM68uLef","execution":{"iopub.status.busy":"2021-05-25T20:51:55.596368Z","iopub.status.idle":"2021-05-25T20:51:55.596917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finding the biggest object out of an image, cropping it out and saveing the image","metadata":{"id":"_qlFO5vqoXX_"}},{"cell_type":"code","source":"crop_train_images_list=[]\nfor ind,img_cont in enumerate(contours):\n  # Find object with the biggest bounding box\n  mx_rect = (0,0,0,0)      # biggest skewed bounding box\n  mx_area = 0\n  for cont in img_cont:\n      arect = cv2.minAreaRect(cont)\n      area = arect[1][0]*arect[1][1]\n      if area > mx_area:\n          mx_rect, mx_area = arect, area\n\n  tmp_img=crop_minAreaRect(train_images[ind], mx_rect)\n  cv2.imwrite(\"/content/drive/My Drive/Apple-class/Data/croped_images/from_300_2/\"+train_data[\"image_id\"][ind]+\".jpg\",cv2.cvtColor(tmp_img, cv2.COLOR_RGB2BGR))\n","metadata":{"id":"PnkME9mGy420","outputId":"5f783c1b-b3dc-4aab-8140-47406aab9995","execution":{"iopub.status.busy":"2021-05-25T20:51:55.598472Z","iopub.status.idle":"2021-05-25T20:51:55.599036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **Cropped Images** ","metadata":{"id":"NF7oAxon_CYo"}},{"cell_type":"markdown","source":"Now we have a data set of cropped images by the biggest object, we should check if that help and if the cropping was successful; Lets load the images.","metadata":{"id":"FxsNznOhAhMr"}},{"cell_type":"code","source":"CROPED_PATH=\"/content/drive/My Drive/Apple-class/Data/croped_images/from_300_2/\"\ncrop_train_images=train_data[\"image_id\"].progress_apply(load_dif_image, args=(CROPED_PATH,(100,100)))\ncrop_images_array=images_4d_array(crop_train_images)\ncrop_images_array_all = get_all_pixels(crop_images_array)","metadata":{"id":"aFKOIV9o_AhF","outputId":"a6660444-6e05-4340-cf1c-63858855e935","execution":{"iopub.status.busy":"2021-05-25T20:51:55.600224Z","iopub.status.idle":"2021-05-25T20:51:55.600777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"current_crop_train_data=train_data.loc[crop_train_images.index]\ntargets=get_target_array(current_crop_train_data)","metadata":{"id":"onbjecNgKrie","execution":{"iopub.status.busy":"2021-05-25T20:51:55.602232Z","iopub.status.idle":"2021-05-25T20:51:55.60279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets take a look at the first images, and try to determant the succese of the cropping","metadata":{"id":"3_6BMLLFIQC_"}},{"cell_type":"code","source":"first_images_merged = merge_images(crop_images_array[0:16], size = [4, 4])\n\nmatplotlib.pyplot.figure(figsize=(10,5))\nmatplotlib.pyplot.subplot(1, 2, 1)\nmatplotlib.pyplot.title('Highest PC1 Score')\nmatplotlib.pyplot.axis('off')\nmatplotlib.pyplot.imshow(first_images_merged)","metadata":{"id":"4mx5SBa7n7ii","outputId":"60b836bc-24be-4e6c-b6f7-a0314acd8fa7","execution":{"iopub.status.busy":"2021-05-25T20:51:55.60393Z","iopub.status.idle":"2021-05-25T20:51:55.604516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen, the result is bad,lets check closely one of the images that our algoritm could not crop the leaf properly.\n","metadata":{"id":"yfUZdewN_AhJ"}},{"cell_type":"code","source":"show_image(train_images[3])","metadata":{"id":"iPi4UcpXM1ht","outputId":"f1c19fa7-80d4-451b-caa8-3e29de2fc39e","execution":{"iopub.status.busy":"2021-05-25T20:51:55.605735Z","iopub.status.idle":"2021-05-25T20:51:55.606308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_image(crop_train_images[3])","metadata":{"id":"2DtcHJiRQWvi","outputId":"b5b66887-5d4f-4b31-d955-0ec177775df6","execution":{"iopub.status.busy":"2021-05-25T20:51:55.607422Z","iopub.status.idle":"2021-05-25T20:51:55.607965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"our algorithm couldn't identify the leaves, as it confused the veins of the leaves as their edges. Maybe the right approch would be to nullify the background and that would give us better results","metadata":{"id":"Bp_Mw1ahM3-A"}},{"cell_type":"markdown","source":"### **Background Removal**","metadata":{"id":"42lc0likk9Fj"}},{"cell_type":"markdown","source":"Now lets try and remove the backgroung by replacing it with black pixels. Again we have too much images for manually determine how to set the parametr that remove the backgriond, so we sholud try and automatize it. We tried to manually adjust the parameter that determans hoe much og the image will become black. Obviously there is no one value that is optimal for all images, so we used the one that by random sampaling the leaf itels didn't became black.","metadata":{"id":"PSZXsxRPS1TD"}},{"cell_type":"markdown","source":"#### **Creating Images with Minimal Background**","metadata":{"id":"7AGN5ys3pMUR"}},{"cell_type":"markdown","source":"Now we will try to remove as nuch as we can from the background of all the images and save them","metadata":{"id":"YH8nXTT7akWq"}},{"cell_type":"code","source":"def remove_background1(name,img1):\n  #https://stackoverflow.com/questions/59824040/\n    #== Parameters =======================================================================\n    BLUR = 5\n    CANNY_THRESH_1 = 10\n    CANNY_THRESH_2 = 100\n    MASK_DILATE_ITER = 20\n    MASK_ERODE_ITER = 20\n    MASK_COLOR = (0.0,0.0,0.0) # In BGR format\n    \n    #== Processing =======================================================================\n    \n    #-- Read image -----------------------------------------------------------------------\n    #img = cv2.imread(img)\n    img=img1.copy()\n    gray = cv2.cvtColor(img,cv2.COLOR_BGR2GRAY)\n    \n    #-- Edge detection -------------------------------------------------------------------\n    edges = cv2.Canny(gray, CANNY_THRESH_1, CANNY_THRESH_2)\n    edges = cv2.dilate(edges, None)\n    edges = cv2.erode(edges, None)\n    \n    #-- Find contours in edges, sort by area ---------------------------------------------\n    contour_info = []\n    contours, _ = cv2.findContours(edges, cv2.RETR_LIST, cv2.CHAIN_APPROX_NONE)\n    \n    for c in contours:\n        contour_info.append((\n            c,\n            cv2.isContourConvex(c),\n            cv2.contourArea(c),\n        ))\n    contour_info = sorted(contour_info, key=lambda c: c[2], reverse=True)\n    \n    \n    #-- Create empty mask, draw filled polygon on it corresponding to largest contour ----\n    # Mask is black, polygon is white\n    mask = np.zeros(edges.shape)\n    for c in contour_info:\n        cv2.fillConvexPoly(mask, c[0], (255))\n    # cv2.fillConvexPoly(mask, max_contour[0], (255))\n    \n    #-- Smooth mask, then blur it --------------------------------------------------------\n    mask = cv2.dilate(mask, None, iterations=MASK_DILATE_ITER)\n    mask = cv2.erode(mask, None, iterations=MASK_ERODE_ITER)\n    mask = cv2.GaussianBlur(mask, (BLUR, BLUR), 0)\n    mask_stack = np.dstack([mask]*3)    # Create 3-channel alpha mask\n    \n    #-- Blend masked img into MASK_COLOR background --------------------------------------\n    mask_stack  = mask_stack.astype('float32') / 255.0          # Use float matrices, \n    img         = img.astype('float32') / 255.0                 #  for easy blending\n    \n    masked = (mask_stack * img) + ((1 - mask_stack) * MASK_COLOR) # Blend\n    masked = (masked * 255).astype('uint8')                     # Convert back to 8-bit \n    cv2.imwrite(\"/content/drive/My Drive/Apple-class/Data/deleted_background/from_300/\"+name+\".jpg\",cv2.cvtColor(masked, cv2.COLOR_RGB2BGR))\n    #show_image(masked)","metadata":{"id":"J2mmKRdelBAf","execution":{"iopub.status.busy":"2021-05-25T20:51:55.609301Z","iopub.status.idle":"2021-05-25T20:51:55.609851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for ind,tmp_img in enumerate(train_images):\n  remove_background1(train_data[\"image_id\"][ind],tmp_img)","metadata":{"id":"Cp76nFnIle8D","execution":{"iopub.status.busy":"2021-05-25T20:51:55.611059Z","iopub.status.idle":"2021-05-25T20:51:55.611625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{"id":"1eqQodwKcHoI"}},{"cell_type":"markdown","source":"#### **No Background Images** ","metadata":{"id":"hV-b0QlIcS03"}},{"cell_type":"markdown","source":"Now we have a data set of most background removed images, we should check if that help and if the removal was successful; Lets load the images.","metadata":{"id":"NidkWzwjcS03"}},{"cell_type":"code","source":"NO_BACK_PATH=\"/content/drive/My Drive/Apple-class/Data/deleted_background/from_300/\"\nno_back_train_images=train_data[\"image_id\"].progress_apply(load_dif_image, args=(NO_BACK_PATH,(100,100)))\nno_back_images_array=images_4d_array(no_back_train_images)\nno_back_images_array_all = get_all_pixels(no_back_images_array)","metadata":{"id":"PTKSu1zecS03","outputId":"c8c6b549-1549-41d9-97a3-51b08871ba37","execution":{"iopub.status.busy":"2021-05-25T20:51:55.612628Z","iopub.status.idle":"2021-05-25T20:51:55.613181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"current_no_back_train_data=train_data.loc[no_back_train_images.index]\ntargets=get_target_array(current_no_back_train_data)","metadata":{"id":"UZK-RHUrcS04","execution":{"iopub.status.busy":"2021-05-25T20:51:55.614261Z","iopub.status.idle":"2021-05-25T20:51:55.61481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets take a look at the first images, and try to determant the succese of the cropping","metadata":{"id":"UlNXh7eFcS05"}},{"cell_type":"code","source":"first_images_merged = merge_images(no_back_images_array[0:16], size = [4, 4])\n\nmatplotlib.pyplot.figure(figsize=(10,5))\nmatplotlib.pyplot.subplot(1, 2, 1)\nmatplotlib.pyplot.title('Highest PC1 Score')\nmatplotlib.pyplot.axis('off')\nmatplotlib.pyplot.imshow(first_images_merged)","metadata":{"id":"IyLF3YvwcS05","outputId":"1b4d85bc-423b-41fc-ae6b-f0741cdfd654","execution":{"iopub.status.busy":"2021-05-25T20:51:55.615951Z","iopub.status.idle":"2021-05-25T20:51:55.616527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen, the result us ok, in some of the images the background is pretty much removed, but in others the background is basicly untouched ,lets check closely one of the images that our algoritm has successfuly removed the background.\n","metadata":{"id":"C2GtElX0cS05"}},{"cell_type":"code","source":"show_image(train_images[3])","metadata":{"id":"i71iJCjLcS05","outputId":"b84ae7f7-88b0-4bb9-fe94-0825d0ce8dfe","execution":{"iopub.status.busy":"2021-05-25T20:51:55.618063Z","iopub.status.idle":"2021-05-25T20:51:55.618707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_image(no_back_train_images[3])","metadata":{"id":"az_PB-LzcS05","outputId":"22f252fb-1819-486f-9040-cc6a9a01967f","execution":{"iopub.status.busy":"2021-05-25T20:51:55.620225Z","iopub.status.idle":"2021-05-25T20:51:55.620863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen the algorithm work pretty well on that image, maybe that could help us","metadata":{"id":"Dx8VWV-DcS05"}},{"cell_type":"markdown","source":" Lets Compute the first 10 PCs and the projection on them","metadata":{"id":"tyv9FxKJgmLk"}},{"cell_type":"code","source":"no_back_images_array_all_centered=sklearn.preprocessing.StandardScaler().fit_transform(no_back_images_array_all)\n\npca = sklearn.decomposition.PCA(n_components = 10)\npca.fit(no_back_images_array_all_centered)\nx_train_reduced = pca.transform(no_back_images_array_all_centered)\nprincipalDf=pd.DataFrame(data=x_train_reduced,columns=['PC1','PC2','PC3','PC4','PC5','PC6','PC7','PC8','PC9','PC10'] )\nprincipalDf.head()","metadata":{"id":"G-FwkUIxgmLk","outputId":"507dacd5-e599-4a3e-91b0-71bb54140433","execution":{"iopub.status.busy":"2021-05-25T20:51:55.622471Z","iopub.status.idle":"2021-05-25T20:51:55.623096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we have the first 10 PCs we can check how much each of them discribes the variance of the data, by showing the percentage of variance explained by each of the selected components","metadata":{"id":"TG0K9EIXgmLl"}},{"cell_type":"code","source":"pca.explained_variance_ratio_\n\n\ndf1 = pd.DataFrame({'var':pca.explained_variance_ratio_,\n             'PC':['PC1','PC2','PC3','PC4','PC5','PC6','PC7','PC8','PC9','PC10']})\nsns.barplot(x='PC',y=\"var\", \n           data=df1, color=\"c\");","metadata":{"id":"DqP3GTU1gmLl","outputId":"9d2ba3eb-3650-4ff4-fe3d-c736f796001f","execution":{"iopub.status.busy":"2021-05-25T20:51:55.624514Z","iopub.status.idle":"2021-05-25T20:51:55.625284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen, the first 2 PCs explaies about 19 precent of the varience, mabye we can coralate it to our classes. We can that the variance explained has remaind pretty much the same","metadata":{"id":"UJDD2dy5gmLl"}},{"cell_type":"markdown","source":"Lets plot the data on the first 2 PCs, this can visualize easily clusters of data in on those axis if they exist","metadata":{"id":"XNWYaqn3gmLl"}},{"cell_type":"code","source":"matplotlib.pyplot.xlabel('PC 1')\nmatplotlib.pyplot.ylabel('PC 2')\nmatplotlib.pyplot.title('2D PCA')\nmatplotlib.pyplot.scatter(x_train_reduced[:, 0], x_train_reduced[:, 1],c = targets, cmap=matplotlib.pyplot.cm.get_cmap('YlGnBu', 10))\nmatplotlib.pyplot.colorbar();","metadata":{"id":"RqNlRjPCgmLm","outputId":"8ebab414-01f5-4019-f74e-135e34965292","execution":{"iopub.status.busy":"2021-05-25T20:51:55.626451Z","iopub.status.idle":"2021-05-25T20:51:55.627011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The dispersal of data does not correspond with 4 classes or some kind of clusters, which suggests that the PCs does not discribe the difference in the classes. \n","metadata":{"id":"-VgaMdR0gmLm"}},{"cell_type":"markdown","source":"Since we're dealing with images the best way to get what a given PC's \"subject\", what it discribes, is  simply to view those images which have a high or low score for this PC. We will show the first 2 PCs, because each of the rest represents less the 4 precent of the variance, which is very low","metadata":{"id":"77KHahwHgmLm"}},{"cell_type":"markdown","source":"##### **First PC**","metadata":{"id":"8om0zfT2kgGu"}},{"cell_type":"markdown","source":"First we should find the top and bottom 16 images in the PC\n\n\n","metadata":{"id":"wiSABug8kgGu"}},{"cell_type":"code","source":"highest_score_ids = np.argpartition(x_train_reduced[:, 0], -16)[-16:]\nlowest_score_ids = np.argpartition(x_train_reduced[:, 0], 16)[:16]","metadata":{"id":"UCxLwvBLkgGu","execution":{"iopub.status.busy":"2021-05-25T20:51:55.628112Z","iopub.status.idle":"2021-05-25T20:51:55.628699Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we will print the classification of them, maybe the this could tell us something","metadata":{"id":"ErPyMOSmkgGv"}},{"cell_type":"code","source":"print(current_no_back_train_data.iloc[highest_score_ids])\nprint(current_no_back_train_data.iloc[lowest_score_ids])","metadata":{"id":"V6u0kq4GkgGv","outputId":"f5b9e2a3-8b3a-414c-8ef9-142f76e88661","execution":{"iopub.status.busy":"2021-05-25T20:51:55.629912Z","iopub.status.idle":"2021-05-25T20:51:55.630486Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There isn't a clear difference betweet the groups","metadata":{"id":"E8XnRmJvkgGw"}},{"cell_type":"markdown","source":"Lets print the images and check if the difference can be identify","metadata":{"id":"7aRPtUGzkgGw"}},{"cell_type":"code","source":"highest_images_4d=no_back_images_array[highest_score_ids]\nlowest_images_4d=no_back_images_array[lowest_score_ids]\n\nhighest_images_merged = merge_images(highest_images_4d, size = [4, 4])\nlowest_images_merged = merge_images(lowest_images_4d, size = [4, 4])\n\nmatplotlib.pyplot.figure(figsize=(10,5))\nmatplotlib.pyplot.subplot(1, 2, 1)\nmatplotlib.pyplot.title('Highest PC1 Score')\nmatplotlib.pyplot.axis('off')\nmatplotlib.pyplot.imshow(highest_images_merged)\n\nmatplotlib.pyplot.subplot(1, 2, 2)\nmatplotlib.pyplot.title('Lowest PC1 Score')\nmatplotlib.pyplot.axis('off')\nmatplotlib.pyplot.imshow(lowest_images_merged)","metadata":{"id":"-Vzse2UQkgGw","outputId":"1836484c-bee9-480d-bcf3-28be3ae907d1","execution":{"iopub.status.busy":"2021-05-25T20:51:55.631737Z","iopub.status.idle":"2021-05-25T20:51:55.632308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the images we can infer that the first PC is now about how much of the background is black, and the fact that we could not remove the background properly for every image has manifested.","metadata":{"id":"C3i2x7UykgGw"}},{"cell_type":"markdown","source":"##### **Secound PC**","metadata":{"id":"35wBIc8U2t5e"}},{"cell_type":"markdown","source":"First we should find the top and bottom 16 images in the PC\n\n\n","metadata":{"id":"Aki-4vTJ2t5k"}},{"cell_type":"code","source":"highest_score_ids = np.argpartition(x_train_reduced[:, 1], -16)[-16:]\nlowest_score_ids = np.argpartition(x_train_reduced[:, 1], 16)[:16]","metadata":{"id":"YKe24Y8c2t5l","execution":{"iopub.status.busy":"2021-05-25T20:51:55.633516Z","iopub.status.idle":"2021-05-25T20:51:55.634091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we will print the classification of them, maybe the this could tell us something","metadata":{"id":"OBfSFbLz2t5q"}},{"cell_type":"code","source":"print(current_no_back_train_data.iloc[highest_score_ids])\nprint(current_no_back_train_data.iloc[lowest_score_ids])","metadata":{"id":"bW0EKpgy2t5r","outputId":"025a91a0-ac38-44de-ac6f-01b9a3ca53a6","execution":{"iopub.status.busy":"2021-05-25T20:51:55.635343Z","iopub.status.idle":"2021-05-25T20:51:55.636006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There isn't a clear difference betweet the groups","metadata":{"id":"xsWGfdpM2t5v"}},{"cell_type":"markdown","source":"Lets print the images and check if the difference can be identify","metadata":{"id":"tBwpNLPD2t5v"}},{"cell_type":"code","source":"highest_images_4d=no_back_images_array[highest_score_ids]\nlowest_images_4d=no_back_images_array[lowest_score_ids]\n\nhighest_images_merged = merge_images(highest_images_4d, size = [4, 4])\nlowest_images_merged = merge_images(lowest_images_4d, size = [4, 4])\n\nmatplotlib.pyplot.figure(figsize=(10,5))\nmatplotlib.pyplot.subplot(1, 2, 1)\nmatplotlib.pyplot.title('Highest PC2 Score')\nmatplotlib.pyplot.axis('off')\nmatplotlib.pyplot.imshow(highest_images_merged)\n\nmatplotlib.pyplot.subplot(1, 2, 2)\nmatplotlib.pyplot.title('Lowest PC2 Score')\nmatplotlib.pyplot.axis('off')\nmatplotlib.pyplot.imshow(lowest_images_merged)","metadata":{"id":"qVG__Mv82t5w","outputId":"b410e560-ea44-423d-86df-747c99aa371f","execution":{"iopub.status.busy":"2021-05-25T20:51:55.63711Z","iopub.status.idle":"2021-05-25T20:51:55.637694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the images we can infer that the secound PC is now about the orientation of the black background, and the fact that we could noe remove the background properlly for every image has manifested in it.","metadata":{"id":"CSz0IgMN2t5z"}},{"cell_type":"markdown","source":"## **Section Summary**","metadata":{"id":"b2C_TzgDlgog"}},{"cell_type":"markdown","source":"At this section we tried to get to know our data and to manipulate it so it will be easier to work with. we learned that we have imbalance in our data in our first section. An hypothisys we had, after analusing the color profile of the images is that the spots on the diseased leaves was prominent in the images, but as we explored the PCs we were proven wrong. We saw that the colors of the images was more infuence by the background and by the lighing on the moment that the image was taken. Then we tried to manipulate the images so that those 2 factor could be nullified with no success. The images were proven to hard to manipulate automaticly due to them having small color diference between the color of the leaves and the color of the background, and duo to some images having bad focus over the edges of the leavs. Furthermore by applying the manipulation ,as best as we could do, didn't yielded any benefit or adventege over the basic data-set. There are far more sophisticated methods to automate the manipulation we tried, but those methods requires far greater knowledge, proficiency, resources and requires a project-size work in order to achieve them.","metadata":{"id":"jO9su4FRmEVU"}},{"cell_type":"markdown","source":"# **Modeling**","metadata":{"id":"14nId2_mtj0v"}},{"cell_type":"markdown","source":"In this section we will try to fit models to our data so we could predict on new data, which of the 4 classes they are part of.","metadata":{"id":"qjlS_ke3u18x"}},{"cell_type":"markdown","source":"##**K-Mean Clustering**","metadata":{"id":"u9mkZM5D69GZ"}},{"cell_type":"markdown","source":"Our images are classifed to their classes, but we can try and check if the data is scattered around 4 clusters in the feature space, and if so, we can check if it corresponds to the 4 classes, we will try it by using the k-mean cluster method","metadata":{"id":"sK9Cu5qi69Ga"}},{"cell_type":"markdown","source":"First we will set the model","metadata":{"id":"HBZklwC3wPoP"}},{"cell_type":"code","source":"kmeans= sklearn.cluster.KMeans(n_clusters=4)","metadata":{"id":"hcmgCFOD69Gb","execution":{"iopub.status.busy":"2021-05-25T20:51:55.638919Z","iopub.status.idle":"2021-05-25T20:51:55.639487Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Mean Pixel Value**","metadata":{"id":"T9-ohnNt69Gh"}},{"cell_type":"markdown","source":"Firstly we will try to use a single value for each picture, its mean pixel value, it is one dimentioal and very simple.","metadata":{"id":"4oik4Vg369Gk"}},{"cell_type":"code","source":"avarage_pixel=images_array_all.mean(axis=1).reshape(-1, 1)\nkmeans.fit(avarage_pixel)\ndata_kmeans=kmeans.predict(avarage_pixel)","metadata":{"id":"g6_UHiRA69Gl","execution":{"iopub.status.busy":"2021-05-25T20:51:55.640748Z","iopub.status.idle":"2021-05-25T20:51:55.641325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The classification number will not correspond with the 4 clustering classes, so we will just check if the clustering classes are divided in a way that can be aligned with the classes","metadata":{"id":"hS4J2-jg69Gz"}},{"cell_type":"code","source":"labels=np.zeros_like(data_kmeans)\nfor i in range(4):\n  mask=(data_kmeans==i)\n  labels[mask]=(i)%4","metadata":{"id":"6KZQJFjJ69G1","execution":{"iopub.status.busy":"2021-05-25T20:51:55.642578Z","iopub.status.idle":"2021-05-25T20:51:55.643139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mat = sklearn.metrics.confusion_matrix(targets,labels)\nsns.heatmap(mat.T, square=True, annot=True, cbar=False)\nmatplotlib.pyplot.xlabel('true label')\nmatplotlib.pyplot.ylabel('predict label');","metadata":{"id":"i25c1tKR69HE","outputId":"7e6d92b6-84b4-40b1-e18b-41728cef943b","execution":{"iopub.status.busy":"2021-05-25T20:51:55.644141Z","iopub.status.idle":"2021-05-25T20:51:55.644738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen from the confusion matrix most of the predictions are all over the place and didn't converge into one classification for each class. Thus we can infer that the images are not separated into 4 clusters that corresponds to our classes in the discussed space","metadata":{"id":"fJIATO_l69HL"}},{"cell_type":"markdown","source":"### **Mean Channel Value**","metadata":{"id":"I-VS74-e69HM"}},{"cell_type":"markdown","source":"Now we will try to use a 3 channel mean value for image, a bit more complex space, a 3 dimentional one. Maybe this cold halp us identify some kind of pattern that has a connention to the classes","metadata":{"id":"I5OOFRA369HO"}},{"cell_type":"code","source":"avarage_pixel_channels=get_channels(images_array)\n","metadata":{"id":"cJ-Duvfp69HO","execution":{"iopub.status.busy":"2021-05-25T20:51:55.645894Z","iopub.status.idle":"2021-05-25T20:51:55.646454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kmeans.fit(avarage_pixel_channels)\ndata_kmeans=kmeans.predict(avarage_pixel_channels)\nprint(data_kmeans)","metadata":{"id":"551L5l2j69Hf","outputId":"c9528974-3cbd-4815-9077-5f3f3495413f","execution":{"iopub.status.busy":"2021-05-25T20:51:55.647681Z","iopub.status.idle":"2021-05-25T20:51:55.648243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The classification number will not correspond with the 4 clustering classes, so we will just check if the clustering classes are divided in a way that can be aligned with the classes","metadata":{"id":"fDAloVlVmGsY"}},{"cell_type":"code","source":"labels=np.zeros_like(data_kmeans)\nfor i in range(4):\n  mask=(data_kmeans==i)\n  labels[mask]=(i)%4","metadata":{"id":"OCQ8Os4W69Hv","execution":{"iopub.status.busy":"2021-05-25T20:51:55.649625Z","iopub.status.idle":"2021-05-25T20:51:55.650172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mat = sklearn.metrics.confusion_matrix(targets,labels)\nsns.heatmap(mat.T, square=True, annot=True, cbar=False)\nmatplotlib.pyplot.xlabel('true label')\nmatplotlib.pyplot.ylabel('predict label');","metadata":{"id":"4_R65Blb69H1","outputId":"660ee552-eb10-4be4-cafe-5170e271926c","execution":{"iopub.status.busy":"2021-05-25T20:51:55.651604Z","iopub.status.idle":"2021-05-25T20:51:55.652172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen from the confusion matrix most of the predictions are all over the place and didn't converge into one classification for each class. Thus we can infer that the images are not separated into 4 clusters that corresponds to our classes in the discussed space","metadata":{"id":"hbo7elfjmqtj"}},{"cell_type":"markdown","source":"### **All Channel Values**","metadata":{"id":"loIw_p4O69H-"}},{"cell_type":"markdown","source":"Now we will try to use all of the channel values, this is using all of the features ","metadata":{"id":"ZGGgVHjus7SV"}},{"cell_type":"code","source":"kmeans.fit(images_array_all)\ndata_kmeans=kmeans.predict(images_array_all)\nprint(data_kmeans)","metadata":{"id":"9MZIDzgP69IC","outputId":"4aea0fb0-bd73-4f4c-dd80-9fab409828be","execution":{"iopub.status.busy":"2021-05-25T20:51:55.653469Z","iopub.status.idle":"2021-05-25T20:51:55.654031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The classification number will not correspond with the 4 clustering classes, so we will just check if the clustering classes are divided in a way that can be aligned with the classes","metadata":{"id":"qzUcFBJSsv3P"}},{"cell_type":"code","source":"labels=np.zeros_like(data_kmeans)\nfor i in range(4):\n  mask=(data_kmeans==i)\n  labels[mask]=(i)%4","metadata":{"id":"FX0qJSFw69IK","execution":{"iopub.status.busy":"2021-05-25T20:51:55.655487Z","iopub.status.idle":"2021-05-25T20:51:55.6561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mat = sklearn.metrics.confusion_matrix(targets,labels)\nsns.heatmap(mat.T, square=True, annot=True, cbar=False)\nmatplotlib.pyplot.xlabel('true label')\nmatplotlib.pyplot.ylabel('predict label');","metadata":{"id":"y7qV2OlE69IP","outputId":"81519ed5-864b-476e-9cf3-944a70d7b5f0","execution":{"iopub.status.busy":"2021-05-25T20:51:55.657351Z","iopub.status.idle":"2021-05-25T20:51:55.657941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As can be seen from the confusion matrix most of the predictions are all over the place and didn't converge into one classification for each class. Thus we can infer that the images are not separated into 4 clusters that corresponds to our classes in the discussed space","metadata":{"id":"LdvTnqG0tRPB"}},{"cell_type":"markdown","source":"### **Conclusin Thus Far**","metadata":{"id":"dRvpWphQtYr1"}},{"cell_type":"markdown","source":"We saw that the data is not seperated into 4 clusters that corresponds to their class. Its seems that the data seperation does not corresponds to the classes, parhaps the fine details that separate the classes needs a more robust method to get unveiled, or that the background color interferes in the modeling","metadata":{"id":"T5ytHq3otZm9"}},{"cell_type":"markdown","source":"##**Logistic Regression**","metadata":{"id":"liVO59oW69IX"}},{"cell_type":"markdown","source":"As we already established, the data is not scatered in a way that can differ easily the classe. In our next attempt we would try to create a logistic regression that we classify images base of their values.","metadata":{"id":"c4zE53qR0M2w"}},{"cell_type":"markdown","source":"Firstly, we would separate the data to train and test, we chose to use a standard 80-20 split","metadata":{"id":"3JEEyhyU2bX5"}},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = sklearn.model_selection.train_test_split(images_array, targets, test_size=0.2, random_state=seed)\nX_train_2d=get_all_pixels(X_train)\nX_test_2d=get_all_pixels(X_test)","metadata":{"id":"LTcTmuqw69Ib","execution":{"iopub.status.busy":"2021-05-25T20:51:55.659302Z","iopub.status.idle":"2021-05-25T20:51:55.659853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(X_train_2d.shape, X_test_2d.shape, y_train.shape, y_test.shape)","metadata":{"id":"CViReZBN69Ij","outputId":"e6684a63-0ff3-4c4a-b9f9-6860c8354b94","execution":{"iopub.status.busy":"2021-05-25T20:51:55.661125Z","iopub.status.idle":"2021-05-25T20:51:55.661695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Mean Pixel Value**","metadata":{"id":"VfY4xz0P33vs"}},{"cell_type":"markdown","source":"As previously we will try first the simplest evaluation method, the mean pixel value ","metadata":{"id":"Ay5313GL69Ir"}},{"cell_type":"code","source":"avarage_pixel_train=X_train_2d.mean(axis=1).reshape(-1, 1)\navarage_pixel_test=X_test_2d.mean(axis=1).reshape(-1, 1)","metadata":{"id":"Z8rPzhfK69Is","execution":{"iopub.status.busy":"2021-05-25T20:51:55.662907Z","iopub.status.idle":"2021-05-25T20:51:55.663494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lrf=sklearn.linear_model.LogisticRegression()\nlrf.fit(avarage_pixel_train,y_train)","metadata":{"id":"7qugT2eB69Ix","outputId":"4bf28c84-c66e-4dbb-acc0-a707e6693d5f","execution":{"iopub.status.busy":"2021-05-25T20:51:55.6646Z","iopub.status.idle":"2021-05-25T20:51:55.665167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets print the histogram of probabilities for each class","metadata":{"id":"Q-NUC29i3xGL"}},{"cell_type":"code","source":"y_pred_prob = lrf.predict_proba(avarage_pixel_test)\n\ncols, rows = 2,2  \nfig, ax = matplotlib.pyplot.subplots(nrows=rows, ncols=cols, figsize=(15, rows*10/3))\nfor i in range(cols*rows):\n  ax[int(i/cols), int(i%rows)].hist(y_pred_prob[:, i])\n  matplotlib.pyplot.setp(ax[int(i/cols), int(i%rows)], xlabel=f'P(label(image) = {current_train_data.columns[i+1]}')\n  matplotlib.pyplot.setp(ax[int(i/cols), int(i%rows)], ylabel='Frequency')\nmatplotlib.pyplot.show()","metadata":{"id":"O-oaHwjp69I7","outputId":"2777d67e-c4b7-401e-fc4c-4a3f06410dd8","execution":{"iopub.status.busy":"2021-05-25T20:51:55.666281Z","iopub.status.idle":"2021-05-25T20:51:55.666913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The histogram above tells us that our predictions are not good, the confidence of the classifications is low, and not a single image is classified in a probability of more then 40%.","metadata":{"id":"d2igBF6m69JF"}},{"cell_type":"code","source":"y_pred = lrf.predict(avarage_pixel_test)\nmat = sklearn.metrics.confusion_matrix(y_test,y_pred)\nsns.heatmap(mat.T, square=True, annot=True, cbar=False)\nmatplotlib.pyplot.xlabel('true label')\nmatplotlib.pyplot.ylabel('predict label');\nprint(sklearn.metrics.classification_report(y_test, y_pred))","metadata":{"id":"zbAZ8k2a69JG","outputId":"fbd24bf9-ee8a-4ef0-e019-55e37525088e","execution":{"iopub.status.busy":"2021-05-25T20:51:55.668082Z","iopub.status.idle":"2021-05-25T20:51:55.66866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our model prediction accuracy is abysmal, the model failed to classify leaves as healthy or as having 2 diseases, or scabs. The 35% accuracy looks a bit better than guessing, but in considerring of our imbalanced data set the accuracy seems not better then guessing","metadata":{"id":"2t9L1v_H69JP"}},{"cell_type":"markdown","source":"### **Mean Channel Values**","metadata":{"id":"3MtLpJIW7Gm4"}},{"cell_type":"markdown","source":"Our next step is to try the next simplest method, the mean channel values for every image as input values for our prediction","metadata":{"id":"ST024KrF69JR"}},{"cell_type":"code","source":"X_train_avarage_pixel_channels=get_channels(X_train)\nX_test_avarage_pixel_channels=get_channels(X_test)\nprint(X_train_avarage_pixel_channels.shape,X_test_avarage_pixel_channels.shape)","metadata":{"id":"XIFitYRI69JR","outputId":"e74ac802-5c61-4b0e-bbab-d7a2f1523fa8","execution":{"iopub.status.busy":"2021-05-25T20:51:55.669862Z","iopub.status.idle":"2021-05-25T20:51:55.670432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lrf=sklearn.linear_model.LogisticRegression(max_iter=1000)\nlrf.fit(X_train_avarage_pixel_channels,y_train)","metadata":{"id":"9pS6JyvJ69JX","outputId":"c09ab651-e5d4-4b64-c8a0-8abfe5f32877","execution":{"iopub.status.busy":"2021-05-25T20:51:55.671632Z","iopub.status.idle":"2021-05-25T20:51:55.672178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = lrf.predict(X_test_avarage_pixel_channels)\nmat = sklearn.metrics.confusion_matrix(y_test,y_pred)\nsns.heatmap(mat.T, square=True, annot=True, cbar=False)\nmatplotlib.pyplot.xlabel('true label')\nmatplotlib.pyplot.ylabel('predict label');\n\nprint(sklearn.metrics.classification_report(y_test, y_pred))","metadata":{"id":"S30OcWKw69Jd","outputId":"47bdbeef-39a0-4c81-af60-3a86b31dae68","execution":{"iopub.status.busy":"2021-05-25T20:51:55.673553Z","iopub.status.idle":"2021-05-25T20:51:55.674105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our model prediction accuracy is better, the model failed to classify leaves as having 2 diseases, but that can be challenging because of our small number of images for this class. The 39% accuracy marks an improvment over previous attempts, but doesn't gives us a real efficient working tool","metadata":{"id":"P6kZ-OQr69Jj"}},{"cell_type":"markdown","source":"Lets print the histogram of probabilities for each class","metadata":{"id":"jFgR7a251Fnm"}},{"cell_type":"code","source":"y_pred_prob = lrf.predict_proba(X_test_avarage_pixel_channels)\n\ncols, rows = 2,2  \nfig, ax = matplotlib.pyplot.subplots(nrows=rows, ncols=cols, figsize=(15, rows*10/3))\nfor i in range(cols*rows):\n  ax[int(i/cols), int(i%rows)].hist(y_pred_prob[:, i])\n  matplotlib.pyplot.setp(ax[int(i/cols), int(i%rows)], xlabel=f'P(label(image) = {current_train_data.columns[i+1]}')\n  matplotlib.pyplot.setp(ax[int(i/cols), int(i%rows)], ylabel='Frequency')\nmatplotlib.pyplot.show()","metadata":{"id":"VvHNZB_qAWCY","outputId":"509e9f0f-b1bc-47ad-9e97-fa83edfd9f1c","execution":{"iopub.status.busy":"2021-05-25T20:51:55.675334Z","iopub.status.idle":"2021-05-25T20:51:55.675921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that again the model classifies images with low degree of certainty, which indicate that our model is poor","metadata":{"id":"9WI3_YpfAibc"}},{"cell_type":"markdown","source":"### **All Channel Values**","metadata":{"id":"kCLco1hx8vli"}},{"cell_type":"markdown","source":"Our last attempt at this method is to try it on all of the pixels channels values as our input, perhapes this can give us a better model for classification","metadata":{"id":"SSQBolye69Js"}},{"cell_type":"code","source":"lrf=sklearn.linear_model.LogisticRegression(max_iter=10000)\nlrf.fit(X_train_2d,y_train)","metadata":{"id":"tXPm10eR69Ju","outputId":"7a7cf6f2-056b-4e4c-96e8-2fce828efbdb","execution":{"iopub.status.busy":"2021-05-25T20:51:55.677035Z","iopub.status.idle":"2021-05-25T20:51:55.677618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = lrf.predict(X_test_2d)\nmat = sklearn.metrics.confusion_matrix(y_test,y_pred)\nsns.heatmap(mat.T, square=True, annot=True, cbar=False)\nmatplotlib.pyplot.xlabel('true label')\nmatplotlib.pyplot.ylabel('predict label');\n\nprint(sklearn.metrics.classification_report(y_test, y_pred))","metadata":{"id":"fJeeAdqv69Jy","outputId":"63b0f29a-9f74-4bb2-a3de-2634c5ace7eb","execution":{"iopub.status.busy":"2021-05-25T20:51:55.678981Z","iopub.status.idle":"2021-05-25T20:51:55.679932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our model on every channel value has retairn its accuracy leve. It again failed to classify correctly multi-diseased leaves and is just a bit better then blindly guessing","metadata":{"id":"KBMsV9PN69J5"}},{"cell_type":"markdown","source":"Lets print the histogram of probabilities for each class","metadata":{"id":"mf67Fvhb1He9"}},{"cell_type":"code","source":"y_pred_prob = lrf.predict_proba(X_test_2d)\n\ncols, rows = 2,2  \nfig, ax = matplotlib.pyplot.subplots(nrows=rows, ncols=cols, figsize=(15, rows*10/3))\nfor i in range(cols*rows):\n  ax[int(i/cols), int(i%rows)].hist(y_pred_prob[:, i])\n  matplotlib.pyplot.setp(ax[int(i/cols), int(i%rows)], xlabel=f'P(label(image) = {current_train_data.columns[i+1]}')\n  matplotlib.pyplot.setp(ax[int(i/cols), int(i%rows)], ylabel='Frequency')\nmatplotlib.pyplot.show()","metadata":{"id":"AuQfRe-x69J5","outputId":"c2b15eff-2e80-4072-dc5a-21a990f26747","execution":{"iopub.status.busy":"2021-05-25T20:51:55.68116Z","iopub.status.idle":"2021-05-25T20:51:55.681749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Even though our model had the same accurecy, we can see that its classification is much more decisive. most of the classifications are in the edges of the probability scale, which mean that it is a bit better the our last one\n","metadata":{"id":"RwyoAxYr69J_"}},{"cell_type":"markdown","source":"### **Conclusion Thus Far**","metadata":{"id":"lN53SifDCQYL"}},{"cell_type":"markdown","source":"Our models, again, performs badly. We can infer that the problem is not as easy to solve as we thought at the begining. The logistic regression produces a model that does not classify images currectly. We guess that the main problem is that our data set is not cohesive enough: the brightness of the images differ greatly between images, the background has different color schemes, and over all the features that are importent are small in comparison to the whole image. Another problem in our data set could be the imbalances in groups.\nIn our future models we will try to incorporate more techniques that can handel those problems","metadata":{"id":"kOySOXwhCWr8"}},{"cell_type":"markdown","source":"##  **Multilayer perceptron Neural Network**","metadata":{"id":"2qT0tWznWHV_"}},{"cell_type":"markdown","source":"In our next attempts we will use a much more robust model, a simple neural network. Neural networks is considered to be one of the most accurate model for image classification","metadata":{"id":"sh5uMzf8Mczo"}},{"cell_type":"markdown","source":"First we will splits the data into train and test data","metadata":{"id":"tZIwKl3tOsoI"}},{"cell_type":"code","source":"current_train_data=train_data.loc[train_images.index]\nimages_array=images_4d_array(train_images)\ntargets=get_target_array(current_train_data)\nX_train, X_test, y_train, y_test = sklearn.model_selection.train_test_split(images_array, targets, test_size=0.2, random_state=seed)\nX_train_2d=get_all_pixels(X_train)\nX_test_2d=get_all_pixels(X_test)","metadata":{"id":"cN1Hmko3WsNA","execution":{"iopub.status.busy":"2021-05-25T20:51:55.683067Z","iopub.status.idle":"2021-05-25T20:51:55.68365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train = X_train.astype('float32')\nX_test = X_test.astype('float32')\nX_train /= 255\nX_test /= 255\nX_train_2d=get_all_pixels(X_train)\nX_test_2d=get_all_pixels(X_test)","metadata":{"id":"vVjsiq2iXyFp","execution":{"iopub.status.busy":"2021-05-25T20:51:55.684765Z","iopub.status.idle":"2021-05-25T20:51:55.685336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_2d.shape","metadata":{"id":"BWtf2HbPYHTk","outputId":"cdbe1fa9-a13c-4cfb-9ee3-827094de69c5","execution":{"iopub.status.busy":"2021-05-25T20:51:55.686467Z","iopub.status.idle":"2021-05-25T20:51:55.687018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will build our multilayer perceptron, we chose to start with this model because it can produce good results for some problems and  doesn't take as much time as other models. Starting with less resource consuming option can be beneficial if this model would yields good results ","metadata":{"id":"5NT3nSvHTADw"}},{"cell_type":"code","source":"\n\nbatch_size = 128\nepochs = 10\n\nmodel = keras.models.Sequential()\nmodel.add(keras.layers.Dense(512, activation='relu', input_shape=(30000,)))\nmodel.add(keras.layers.Dropout(0.2))\nmodel.add(keras.layers.Dense(256, activation='relu'))\nmodel.add(keras.layers.Dropout(0.2))\nmodel.add(keras.layers.Dense(128, activation='relu'))\nmodel.add(keras.layers.Dropout(0.2))\nmodel.add(keras.layers.Dense(64, activation='relu'))\nmodel.add(keras.layers.Dropout(0.2))\nmodel.add(keras.layers.Dense(32, activation='relu'))\nmodel.add(keras.layers.Dropout(0.2))\nmodel.add(keras.layers.Dense(16, activation='relu'))\nmodel.add(keras.layers.Dropout(0.2))\nmodel.add(keras.layers.Dense(4, activation='softmax'))\n\nmodel.summary()","metadata":{"id":"1wt2etP5YAa_","outputId":"3c8d990f-d9e6-49b2-ff3b-189ad9e8072c","execution":{"iopub.status.busy":"2021-05-25T20:51:55.688122Z","iopub.status.idle":"2021-05-25T20:51:55.688725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.compile(loss='binary_crossentropy',\n              optimizer='adam',\n              metrics=['accuracy'])","metadata":{"id":"_sI9xiqeZFqH","execution":{"iopub.status.busy":"2021-05-25T20:51:55.689946Z","iopub.status.idle":"2021-05-25T20:51:55.690581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(X_train_2d, y_train,\n                    batch_size=batch_size,\n                    epochs=epochs,\n                    verbose=1,\n                    shuffle=True,\n                    validation_data=(X_test_2d, y_test))\n\nscore = model.evaluate(X_test_2d, y_test, verbose=1)\n\nprint('Test loss:', score[0])\nprint('Test accuracy:', score[1])","metadata":{"id":"giX5LBY1ZMoc","outputId":"0159ecfb-4a8e-4bef-c2dc-09fceb5abc33","execution":{"iopub.status.busy":"2021-05-25T20:51:55.69186Z","iopub.status.idle":"2021-05-25T20:51:55.692604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The model's accuracy is the worst that we saw thus far. Our results are bad and we didn't produced a valid model, we could change the architectur and train to produce a better model, but we shall proceed to a more robust model","metadata":{"id":"MLPLI8NnVZSK"}},{"cell_type":"markdown","source":"## **CNN-Convolution Neural Network**","metadata":{"id":"1J9t9fXSv17Z"}},{"cell_type":"markdown","source":"Convolution neural network is the fittest and most robust model that we know of to this task. It will serve as our main tool in this classification task, thus we will prepare the ground before applying it. A convolutional neural network consists of an input and an output layer, as well as multiple hidden layers. The hidden layers of a CNN typically consist of a series of convolutional layers that convolve with a multiplication or other dot product. The activation function is commonly a ReLU layer, and is subsequently followed by additional convolutions such as pooling layers, fully connected layers and normalization layers, referred to as hidden layers because their inputs and outputs are masked by the activation function and final convolution.","metadata":{"id":"DxwCYnuAs_Wj"}},{"cell_type":"markdown","source":"### **Set Early Stopping Parameters**","metadata":{"id":"hGylXNCFPLVj"}},{"cell_type":"markdown","source":"We have limited resources, and we want to make the most out of them. In order to do so, we set stopping parameters which determands that as the model reaching a plateau, the run will stop","metadata":{"id":"iLzGSUt1uuxb"}},{"cell_type":"code","source":"LR_reduce=keras.callbacks.ReduceLROnPlateau(monitor='val_accuracy',\n                            factor=.5,\n                            patience=10,\n                            min_lr=.000001,\n                            verbose=0)\n\nES_monitor=keras.callbacks.EarlyStopping(monitor='val_loss',\n                          patience=20)\n\n\nreg = .0005","metadata":{"id":"yhp6yAY46fg1","execution":{"iopub.status.busy":"2021-05-25T20:51:55.69384Z","iopub.status.idle":"2021-05-25T20:51:55.694416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Plot History Function**","metadata":{"id":"JhFyHZLUT7PV"}},{"cell_type":"markdown","source":"We will use this function to plot our model progression","metadata":{"id":"UceHTdLyyoQU"}},{"cell_type":"code","source":"def plot_history(history):\n\n  h = history.history\n\n  offset = 5\n  epochs = range(offset, len(h['loss']))\n\n  matplotlib.pyplot.figure(1, figsize=(20, 6))\n\n  matplotlib.pyplot.subplot(121)\n  matplotlib.pyplot.xlabel('epochs')\n  matplotlib.pyplot.ylabel('loss')\n  matplotlib.pyplot.plot(epochs, h['loss'][offset:], label='train')\n  matplotlib.pyplot.plot(epochs, h['val_loss'][offset:], label='val')\n  matplotlib.pyplot.legend()\n\n  matplotlib.pyplot.subplot(122)\n  matplotlib.pyplot.xlabel('epochs')\n  matplotlib.pyplot.ylabel('accuracy')\n  matplotlib.pyplot.plot(h[f'accuracy'], label='train')\n  matplotlib.pyplot.plot(h[f'val_accuracy'], label='val')\n  matplotlib.pyplot.legend()\n\n  matplotlib.pyplot.show()\n\n  pred_test = model.predict(x_val)\n  roc_sum = 0\n  classes=['healthy',\t'multiple_diseases',\t'rust',\t'scab']\n  for i in range(4):\n      score = sklearn.metrics.roc_auc_score(y_val.iloc[:,i].values.astype('int32'), pred_test[:,i])\n      roc_sum += score\n      print(f'AUC-ROC {classes[i]}  {score:.3f}')\n\n  roc_sum /= 4\n  print(f'totally roc score:{roc_sum:.3f}')","metadata":{"id":"99_kJiBAUC9-","execution":{"iopub.status.busy":"2021-05-25T20:51:55.695768Z","iopub.status.idle":"2021-05-25T20:51:55.696351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **CNN Model**","metadata":{"id":"LbjpiKvwQfLA"}},{"cell_type":"markdown","source":"Now we will create a CNN model to our use, it is based on several standards CNN we found and combined","metadata":{"id":"hW2qYbNmzHP5"}},{"cell_type":"code","source":"def creat_model(img_size):\n  model = keras.models.Sequential()\n\n  model.add(keras.layers.Conv2D(32, kernel_size=(5,5),activation='relu', input_shape=(img_size, img_size, 3), kernel_regularizer=keras.regularizers.l2(reg)))\n  model.add(keras.layers.BatchNormalization(axis=-1,center=True,scale=False))\n  model.add(keras.layers.Conv2D(128, kernel_size=(5,5),activation='relu', kernel_regularizer=keras.regularizers.l2(reg)))\n  model.add(keras.layers.BatchNormalization(axis=-1,center=True,scale=False))\n  model.add(keras.layers.MaxPooling2D(pool_size=(2,2), padding='SAME'))\n  model.add(keras.layers.Dropout(.25))\n\n  model.add(keras.layers.Conv2D(32, kernel_size=(3,3),activation='relu', kernel_regularizer=keras.regularizers.l2(reg)))\n  model.add(keras.layers.BatchNormalization(axis=-1,center=True,scale=False))\n  model.add(keras.layers.Conv2D(128, kernel_size=(3,3),activation='relu',kernel_regularizer=keras.regularizers.l2(reg)))\n  model.add(keras.layers.BatchNormalization(axis=-1,center=True,scale=False))\n  model.add(keras.layers.MaxPooling2D(pool_size=(2,2), padding='SAME'))\n  model.add(keras.layers.Dropout(.25))\n\n\n  model.add(keras.layers.Conv2D(128, kernel_size=(5,5),activation='relu', kernel_regularizer=keras.regularizers.l2(reg)))\n  model.add(keras.layers.BatchNormalization(axis=-1,center=True,scale=False))\n  model.add(keras.layers.Conv2D(512, kernel_size=(5,5),activation='relu',kernel_regularizer=keras.regularizers.l2(reg)))\n  model.add(keras.layers.BatchNormalization(axis=-1,center=True,scale=False))\n  model.add(keras.layers.MaxPooling2D(pool_size=(2,2), padding='SAME'))\n  model.add(keras.layers.Dropout(.25))\n\n  model.add(keras.layers.Conv2D(128, kernel_size=(3,3),activation='relu',kernel_regularizer=keras.regularizers.l2(reg)))\n  model.add(keras.layers.BatchNormalization(axis=-1,center=True,scale=False))\n  model.add(keras.layers.Conv2D(512, kernel_size=(3,3),activation='relu',kernel_regularizer=keras.regularizers.l2(reg)))\n  model.add(keras.layers.BatchNormalization(axis=-1,center=True,scale=False))\n  model.add(keras.layers.MaxPooling2D(pool_size=(2,2), padding='SAME'))\n  model.add(keras.layers.Dropout(.25))\n\n  model.add(keras.layers.Flatten())\n  model.add(keras.layers.Dense(300,activation='relu'))\n  model.add(keras.layers.BatchNormalization(axis=-1,center=True,scale=False))\n  model.add(keras.layers.Dropout(.25))\n  model.add(keras.layers.Dense(200,activation='relu'))\n  model.add(keras.layers.BatchNormalization(axis=-1,center=True,scale=False))\n  model.add(keras.layers.Dropout(.25))\n  model.add(keras.layers.Dense(100,activation='relu'))\n  model.add(keras.layers.BatchNormalization(axis=-1,center=True,scale=False))\n  model.add(keras.layers.Dropout(.25))\n  model.add(keras.layers.Dense(4,activation='softmax'))\n\n  model.summary()\n \n  \n  return model","metadata":{"id":"OUZOQjFN5_Nw","execution":{"iopub.status.busy":"2021-05-25T20:51:55.697524Z","iopub.status.idle":"2021-05-25T20:51:55.698081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **First Attempt- 100x100 Image Size**\n","metadata":{"id":"ctkxbSpXOkhN"}},{"cell_type":"markdown","source":"First we will try our model on images that have 100 pixels as their heigth and 100 pixels are their lenght","metadata":{"id":"KebThCP52IFo"}},{"cell_type":"markdown","source":"Lets load the images","metadata":{"id":"Q95boSdY4MLU"}},{"cell_type":"code","source":"train_100X100_images = train_data[\"image_id\"].progress_apply(load_image, args=((100,100),))\n","metadata":{"id":"uWQtv7UK0dI8","outputId":"d1255e3c-907d-418a-fd7c-9b792dfdc8c8","execution":{"iopub.status.busy":"2021-05-25T20:51:55.69959Z","iopub.status.idle":"2021-05-25T20:51:55.700145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will splits the data into train and test data","metadata":{"id":"6bgdj2--suGJ"}},{"cell_type":"code","source":"targets=train_data.drop('image_id',axis=1)","metadata":{"id":"ZMghZA7iW9Ez","execution":{"iopub.status.busy":"2021-05-25T20:51:55.701392Z","iopub.status.idle":"2021-05-25T20:51:55.701953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"current_train_data=train_data.loc[train_100X100_images.index]\nimages_array=images_4d_array(train_100X100_images)\ntargets=current_train_data.drop('image_id',axis=1)\nx_train, x_val, y_train, y_val = sklearn.model_selection.train_test_split(images_array, targets, test_size=0.2, random_state=seed)\nx_train = x_train.astype('float32')\nx_val = x_val.astype('float32')\nx_train /= 255\nx_val /= 255\nx_train.shape, x_val.shape, y_train.shape, y_val.shape","metadata":{"id":"9Z4CLxeMfN1P","outputId":"d52f4711-343b-42f2-e892-517514bd29a2","execution":{"iopub.status.busy":"2021-05-25T20:51:55.703396Z","iopub.status.idle":"2021-05-25T20:51:55.703961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Running the model on 100x100 images","metadata":{"id":"IXVRsDu-RGDe"}},{"cell_type":"code","source":"model = creat_model(100)\nmodel.compile(optimizer='rmsprop',\n              loss='categorical_crossentropy',\n              metrics=['accuracy']\n              )","metadata":{"id":"jj5qgXH-6lAt","outputId":"5957de52-9c6b-4b0b-8e2f-82a43c6b407f","execution":{"iopub.status.busy":"2021-05-25T20:51:55.705149Z","iopub.status.idle":"2021-05-25T20:51:55.705743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(x_train,\n                    y_train,\n                    batch_size=24,\n                    epochs=250,\n                    steps_per_epoch=x_train.shape[0] // 24,\n                    verbose=0,\n                    callbacks=[ES_monitor,LR_reduce],\n                    validation_data=(x_val, y_val),\n                    validation_steps=x_val.shape[0]//24\n                    )","metadata":{"id":"IkztNL9c68NO","execution":{"iopub.status.busy":"2021-05-25T20:51:55.706954Z","iopub.status.idle":"2021-05-25T20:51:55.707565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets print our results","metadata":{"id":"_J11wQ1id026"}},{"cell_type":"code","source":"score = model.evaluate(x_val, y_val, verbose=0)\n\nprint('Test loss:', score[0])\nprint('Test accuracy:', score[1])\nplot_history(history)","metadata":{"id":"a5P9LHV4-roq","outputId":"063ca2e2-093e-499e-8944-45e9cc727b1c","execution":{"iopub.status.busy":"2021-05-25T20:51:55.708661Z","iopub.status.idle":"2021-05-25T20:51:55.709221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our accuracy has improved drastically to 84%, but there are still some problems, we can see at the end that there is a gap between the validation and the train matrices, which can indicate that there is some kind of overfitting. We can also see that the model is having hard time classifing the \"multiple_diseases\" class, due to it lower AUC-ROC score.\n\n\n\n\n","metadata":{"id":"8XcuWIU89kvt"}},{"cell_type":"markdown","source":"Lets take a look at the confusin matrix for more insight","metadata":{"id":"P5jnehuQxW9S"}},{"cell_type":"code","source":"y_val_pred = model.predict_classes(x_val)\nmat = sklearn.metrics.confusion_matrix(get_target_array_no_im_id(y_val),y_val_pred)\nsns.heatmap(mat.T, square=True, annot=True, cbar=False)\nmatplotlib.pyplot.xlabel('true label')\nmatplotlib.pyplot.ylabel('predict label');","metadata":{"id":"R-XR0ZlfqJgL","outputId":"7cadff78-7ce4-494a-84c2-00e0633da6a3","execution":{"iopub.status.busy":"2021-05-25T20:51:55.710441Z","iopub.status.idle":"2021-05-25T20:51:55.711003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The classes are ordered in this manner: \"healthy\"  \"multiple_diseases\"  \"rust\"  \"scab\". We can see that our model is pretty good at identifing the \"rust\" class and that it is over all can identify the classes we a decent accuracy. We can see that the main issue of our model is it poor prediction of the \"multiple_diseases\" class, it could be that the under representation of that class is a big factor of that. We can that our model predicts some leaves with \"scab\" as \"healthy\", we think that it can mistake the rust to be the leaves veins. We think that maybe using better resolution of the images can help that problems.","metadata":{"id":"R7UCT0cj42YD"}},{"cell_type":"code","source":"model.save(\"..output/kaggle/working/model_100X100.h5\")","metadata":{"id":"EairoTO_9G_W","execution":{"iopub.status.busy":"2021-05-25T20:51:55.712072Z","iopub.status.idle":"2021-05-25T20:51:55.712654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Second Attempt- 224x224 Image Size**\n","metadata":{"id":"BemH-w6_SWTK"}},{"cell_type":"markdown","source":"We will try to use better resolution so the model could maybe classify some of the classes (\"scab\") better ","metadata":{"id":"f_dIjP7p3tlk"}},{"cell_type":"code","source":"img_size = 224","metadata":{"id":"xl6dk4AI-LTL","execution":{"iopub.status.busy":"2021-05-25T20:51:55.713811Z","iopub.status.idle":"2021-05-25T20:51:55.714379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_224X224_images = train_data[\"image_id\"].progress_apply(load_image, args=((img_size,img_size),))","metadata":{"id":"VcXKVr362pCv","outputId":"1c822f27-c3d5-4a7d-abfc-793dee63a81a","execution":{"iopub.status.busy":"2021-05-25T20:51:55.715568Z","iopub.status.idle":"2021-05-25T20:51:55.716146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will splits the data into train and test data","metadata":{"id":"cVTRP3xLdre4"}},{"cell_type":"code","source":"current_train_data=train_data.loc[train_224X224_images.index]\nimages_array=images_4d_array(train_224X224_images)\ntargets=current_train_data.drop('image_id',axis=1)\nx_train2, x_val, y_train2, y_val = sklearn.model_selection.train_test_split(images_array, targets, test_size=0.2, random_state=seed)\nx_train2 = x_train2.astype('float32')\nx_val = x_val.astype('float32')\nx_train2 /= 255\nx_val /= 255\nx_train2.shape, x_val.shape, y_train2.shape, y_val.shape","metadata":{"id":"zgSmngwb2oQT","outputId":"598dcdae-b254-46bc-8aa0-6ff1c295aaa7","execution":{"iopub.status.busy":"2021-05-25T20:51:55.717315Z","iopub.status.idle":"2021-05-25T20:51:55.717963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Running the model on 224x224 images","metadata":{"id":"gaP1Rk-Nzc9U"}},{"cell_type":"code","source":"model = creat_model(img_size)\nmodel.compile(optimizer='rmsprop',\n              loss='categorical_crossentropy',\n              metrics=['accuracy']\n              )","metadata":{"id":"wxTqXVOK-G10","outputId":"341b0b1b-0a9e-42a5-d8ce-18955ab05807","execution":{"iopub.status.busy":"2021-05-25T20:51:55.719229Z","iopub.status.idle":"2021-05-25T20:51:55.71979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_train2.shape","metadata":{"id":"dH8xwWCr1Bsr","outputId":"47d1afd0-e6a1-4156-ed82-43ebba8f26f8","execution":{"iopub.status.busy":"2021-05-25T20:51:55.721031Z","iopub.status.idle":"2021-05-25T20:51:55.72162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(x_train2,\n                    y_train2,\n                    batch_size=24,\n                    epochs=250,\n                    steps_per_epoch=x_train2.shape[0] // 24,\n                    verbose=0,\n                    callbacks=[ES_monitor,LR_reduce],\n                    validation_data=(x_val, y_val),\n                    validation_steps=x_val.shape[0]//24\n                    )","metadata":{"id":"TJUhMLkW-aGZ","execution":{"iopub.status.busy":"2021-05-25T20:51:55.722916Z","iopub.status.idle":"2021-05-25T20:51:55.723496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets print our results","metadata":{"id":"2mZ2pB12d67i"}},{"cell_type":"code","source":"score = model.evaluate(x_val, y_val, verbose=0)\n\nprint('Test loss:', score[0])\nprint('Test accuracy:', score[1])","metadata":{"id":"_G_etafUOolK","outputId":"6bc64b5a-ded0-45c1-b9f3-d6f68c956978","execution":{"iopub.status.busy":"2021-05-25T20:51:55.724803Z","iopub.status.idle":"2021-05-25T20:51:55.72538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_history(history)","metadata":{"id":"oclTTF30UYoP","outputId":"f15677f0-f679-42ed-cf0a-241e212b425d","execution":{"iopub.status.busy":"2021-05-25T20:51:55.726674Z","iopub.status.idle":"2021-05-25T20:51:55.727251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our accuracy has improved to 87%, but there are still some problems, we can see at the end that there is still a gap between the validation and the train matrices, which can indicate that there is some kind of overfitting. We can also see that the model is still having hard time classifing the multi-diseased class, due to it lower AUC-ROC score.\nSo where did the improvment came from?\n\n\n\n","metadata":{"id":"tZttLfQF0sXM"}},{"cell_type":"markdown","source":"Lets take a look at the confusin matrix for more insight","metadata":{"id":"S8uU5zWq0sXN"}},{"cell_type":"code","source":"y_val_pred = model.predict_classes(x_val)\nmat = sklearn.metrics.confusion_matrix(get_target_array_no_im_id(y_val),y_val_pred)\nsns.heatmap(mat.T, square=True, annot=True, cbar=False)\nmatplotlib.pyplot.xlabel('true label')\nmatplotlib.pyplot.ylabel('predict label');","metadata":{"id":"UTZqXwoK0UqO","outputId":"eee9f506-098e-49d7-d89c-754738e960f7","execution":{"iopub.status.busy":"2021-05-25T20:51:55.728637Z","iopub.status.idle":"2021-05-25T20:51:55.729214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The classes are ordered in this manner: \"healthy\"  \"multiple_diseases\"  \"rust\"  \"scab\". We can see that our  new model is pretty good at identifing the \"rust\" class and that it is over all can identify the classes we a decent accuracy as did the last one, but here we can see that using a better rsoluton made the model to identify the \"scab\" leaves better, as it can easily differ the leaf veins from the disease symptoms. We can see that the main issue of our new model has retied it poor prediction of the \"multiple_diseases\" class, it classifies those images as having one of the diseases and not both. It could be that the underrepresentation of that class is a big factor of that. ","metadata":{"id":"P2WKJ_VD2TUJ"}},{"cell_type":"code","source":"model.save(\"..output/kaggle/working/model_224X224.h5\")","metadata":{"id":"g0wsNPKiQ8Le","execution":{"iopub.status.busy":"2021-05-25T20:51:55.730408Z","iopub.status.idle":"2021-05-25T20:51:55.730982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Data Transformations and Augmentation**","metadata":{"id":"YFLkLO66UfTz"}},{"cell_type":"markdown","source":"Image transformations and augmentation is an efficient way to diversify data and to generalize a model, we will try to use it to improve our model predictions and to maybe aid it to classify the \"multiple_diseases\" class better, as generating more images could give it a greater number of relevant images ","metadata":{"id":"EKqn5_bf5DYD"}},{"cell_type":"markdown","source":"We will use a built in function to do so, and we will change images rotation, orientation, brightness and scale\n\n ","metadata":{"id":"HCc2JDphWOoW"}},{"cell_type":"code","source":"datagen = keras.preprocessing.image.ImageDataGenerator(rotation_range=45,\n                             shear_range=.25,\n                              zoom_range=.25,\n                              width_shift_range=.25,\n                              height_shift_range=.25,\n                              rescale=1/255,\n                              brightness_range=[.5,1.5],\n                              horizontal_flip=True,\n                              vertical_flip=True,\n                              fill_mode='nearest'\n#                              featurewise_center=True,\n#                              samplewise_center=True,\n#                              featurewise_std_normalization=True,\n#                              samplewise_std_normalization=True,\n#                              zca_whitening=True\n                              )","metadata":{"id":"5U0bMaR2U1rV","execution":{"iopub.status.busy":"2021-05-25T20:51:55.732003Z","iopub.status.idle":"2021-05-25T20:51:55.732583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Running the model on 224x224 images with the transformations and augmentation","metadata":{"id":"fBZmdQ0TWWG_"}},{"cell_type":"code","source":"model = creat_model(img_size)\nmodel.compile(optimizer='rmsprop',\n              loss='categorical_crossentropy',\n              metrics=['accuracy']\n              )","metadata":{"id":"Z_IDAnYXU9gF","outputId":"753d2c55-cfa3-40be-ef5d-301b244581c5","execution":{"iopub.status.busy":"2021-05-25T20:51:55.733817Z","iopub.status.idle":"2021-05-25T20:51:55.734391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit_generator(datagen.flow(x_train2, y_train2, batch_size=24),\n                              epochs=250,\n                              steps_per_epoch=x_train2.shape[0] // 24,\n                              verbose=0,\n                              callbacks=[ES_monitor,LR_reduce],\n                              validation_data=datagen.flow(x_val, y_val,batch_size=24),\n                              validation_steps=x_val.shape[0]//24\n                              )","metadata":{"id":"i5Nn5W9oVO-w","outputId":"7e121db2-a31f-498a-adca-8bfb5e02bd6b","execution":{"iopub.status.busy":"2021-05-25T20:51:55.735602Z","iopub.status.idle":"2021-05-25T20:51:55.736179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets print our results","metadata":{"id":"QMTkNsyId_Cn"}},{"cell_type":"code","source":"score = model.evaluate(x_val, y_val, verbose=0)\n\nprint('Test loss:', score[0])\nprint('Test accuracy:', score[1])","metadata":{"id":"TzbUbeCtU-2j","outputId":"bfe59be3-674a-4a52-e657-5d8d316f17bb","execution":{"iopub.status.busy":"2021-05-25T20:51:55.73733Z","iopub.status.idle":"2021-05-25T20:51:55.737891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_history(history)","metadata":{"id":"RImEOWUTVIIr","outputId":"0c29e7d6-82b7-4f2f-8a93-9cad76108ba4","execution":{"iopub.status.busy":"2021-05-25T20:51:55.739053Z","iopub.status.idle":"2021-05-25T20:51:55.739635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our accuracy has improved to 92%, a great accomplishment. we can see that the vlidation and test graph have no gap which can show that the model didn't overpitted to the test. We can see that the model is having a better time classifing the \"multiple_diseases\" class from the better AUC-ROC score. ","metadata":{"id":"Yx2s4V4P74t-"}},{"cell_type":"markdown","source":"Lets take a look at the confusin matrix for more insight","metadata":{"id":"qXdex2RiIk4A"}},{"cell_type":"code","source":"y_val_pred = model.predict_classes(x_val)\nmat = sklearn.metrics.confusion_matrix(get_target_array_no_im_id(y_val),y_val_pred)\nsns.heatmap(mat.T, square=True, annot=True, cbar=False)\nmatplotlib.pyplot.xlabel('true label')\nmatplotlib.pyplot.ylabel('predict label');","metadata":{"id":"V10MoqZN6EbW","outputId":"245a3838-1bb9-40de-b8a1-7752590666a1","execution":{"iopub.status.busy":"2021-05-25T20:51:55.740842Z","iopub.status.idle":"2021-05-25T20:51:55.741413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The classes are ordered in this manner: \"healthy\" \"multiple_diseases\" \"rust\" \"scab\". We can see that our new model is better at idintifing the \"healthy\" \"rust\" \"scab\" classes, as it classified corretly almost all of the images of those classes, but still even though is classified more images of that class correctly, the classification seems random.  Again it could be that the underrepresentation of that class is a big factor of that.","metadata":{"id":"u4-u8MkrKseL"}},{"cell_type":"code","source":"model.save(\"..output/kaggle/working/model_224X224_image_generate.h5\")","metadata":{"id":"m83xcesC9Yu6","execution":{"iopub.status.busy":"2021-05-25T20:51:55.742652Z","iopub.status.idle":"2021-05-25T20:51:55.743216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Handling Imbalanced in Our Dataset -SMOTE**\n","metadata":{"id":"b26IjCgJcM8g"}},{"cell_type":"markdown","source":"We'll be applying Synthetic Minority Oversampling Technique (SMOTE). SMOTE works by selecting examples that are close in the feature space, drawing a line between the examples in the feature space and drawing a new sample at a point along that line, and by that balance a minority class. [SMOTE](https://machinelearningmastery.com/smote-oversampling-for-imbalanced-classification/)\n","metadata":{"id":"SB7LS85Rc1Fj"}},{"cell_type":"markdown","source":" Lets check its affect on our dataset. The size of our dataset before running the function","metadata":{"id":"N3k6WMWkdKkE"}},{"cell_type":"code","source":"print(x_train2.shape,y_train2.shape)","metadata":{"id":"lGMU74rZxS5V","outputId":"a917be7c-173b-4f7e-fed1-fb6e3bff3d9f","execution":{"iopub.status.busy":"2021-05-25T20:51:55.74428Z","iopub.status.idle":"2021-05-25T20:51:55.74484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Applying SMOTE on our dataset","metadata":{"id":"eD6c6tmcEL71"}},{"cell_type":"code","source":" sm = imblearn.over_sampling.SMOTE(random_state = 115) \n \nx_train2, y_train2 = sm.fit_resample(get_all_pixels(x_train2),y_train2.to_numpy())\nx_train2 = x_train2.reshape((-1, img_size, img_size, 3))\nx_train2.shape, y_train2.sum(axis=0)","metadata":{"id":"ryG6MKNOvP3l","outputId":"1b7fc9ed-28a7-4b60-befd-19b347cae1c8","execution":{"iopub.status.busy":"2021-05-25T20:51:55.745879Z","iopub.status.idle":"2021-05-25T20:51:55.746455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We see that we have added about 500 examples,most of them to the minority class, and our data is now balanced","metadata":{"id":"uzLB1eEDx9dN"}},{"cell_type":"markdown","source":"Let's try to fit the CNN to the new dataset","metadata":{"id":"H1XCryy2FjP4"}},{"cell_type":"code","source":"model = creat_model(img_size)\nmodel.compile(optimizer='rmsprop',\n              loss='categorical_crossentropy',\n              metrics=['accuracy']\n              )","metadata":{"id":"XDqEBJvOwfdA","outputId":"151e7e2d-634a-4f54-e70c-229743a0aca4","execution":{"iopub.status.busy":"2021-05-25T20:51:55.747631Z","iopub.status.idle":"2021-05-25T20:51:55.748177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit_generator(datagen.flow(x_train2, y_train2, batch_size=24),\n                              epochs=250,\n                              steps_per_epoch=x_train2.shape[0] // 24,\n                              verbose=0,\n                              callbacks=[ES_monitor,LR_reduce],\n                              validation_data=datagen.flow(x_val, y_val,batch_size=24),\n                              validation_steps=x_val.shape[0]//24\n                              )","metadata":{"id":"0joLyHtJwkc0","execution":{"iopub.status.busy":"2021-05-25T20:51:55.749267Z","iopub.status.idle":"2021-05-25T20:51:55.749827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets print our results","metadata":{"id":"nYeIVYr_eLA6"}},{"cell_type":"code","source":"score = model.evaluate(x_val, y_val, verbose=0)\n\nprint('Test loss:', score[0])\nprint('Test accuracy:', score[1])","metadata":{"id":"J67RP9KswoLb","outputId":"d01f0c47-91a3-42ec-e816-2e13d4ac9da1","execution":{"iopub.status.busy":"2021-05-25T20:51:55.751051Z","iopub.status.idle":"2021-05-25T20:51:55.751632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_history(history)","metadata":{"id":"BGKGtXCgvWNB","outputId":"a0dbc66f-8cea-4b7c-bbe9-9bc23d2ed4b2","execution":{"iopub.status.busy":"2021-05-25T20:51:55.752709Z","iopub.status.idle":"2021-05-25T20:51:55.753279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our accuracy has decreased to 74%, Worst then our first CNN. It hard to understand what went wrong; so lets take a look at the confusin matrix for more insight","metadata":{"id":"HxZ6U6yeF3nA"}},{"cell_type":"code","source":"y_val_pred = model.predict_classes(x_val)\nmat = sklearn.metrics.confusion_matrix(get_target_array_no_im_id(y_val),y_val_pred)\nsns.heatmap(mat.T, square=True, annot=True, cbar=False)\nmatplotlib.pyplot.xlabel('true label')\nmatplotlib.pyplot.ylabel('predict label');","metadata":{"id":"0n9qbfmS_KyL","outputId":"bc1d46c6-0a2c-415c-aa55-447f918db45a","execution":{"iopub.status.busy":"2021-05-25T20:51:55.75454Z","iopub.status.idle":"2021-05-25T20:51:55.755303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The classes are ordered in this manner: \"healthy\" \"multiple_diseases\" \"rust\" \"scab\". We can see that our new model classified most of the \"multiple_diseases\" class, but by doing it, it also classified leaves that are in the \"rust\" class as \"multiple_diseases\", and that has made it worse. It seems that the fact that the \"multiple_diseases\" has similar features of the other classes made the \"synthetic\" samlpes have similar atribute to this classes which made the model to get worse. ","metadata":{"id":"CEVx27PpUYhY"}},{"cell_type":"code","source":"model.save(\"..output/kaggle/working/model_224X224_smoth.h5\")","metadata":{"id":"GXWoiqwk9j40","execution":{"iopub.status.busy":"2021-05-25T20:51:55.75669Z","iopub.status.idle":"2021-05-25T20:51:55.757268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Transfer Learning**","metadata":{"id":"7vUVJErXZm7t"}},{"cell_type":"markdown","source":"The general idea of transfer learning is to use knowledge learned from tasks for which a lot of labelled data is available in settings where only little labelled data is available. We will use the pretrained model DenseNet121 inorder to classify","metadata":{"id":"aaO_j-UBZ6p3"}},{"cell_type":"markdown","source":"### **Densenet Model**","metadata":{"id":"qkpYXpFAlN1M"}},{"cell_type":"markdown","source":"We chose DenseNet because it have some advantages over CNN when they are deep. This is because the path for information from the input layer until the output layer (and for the gradient in the opposite direction) becomes so big, that they can get vanished before reaching the other side.\nDenseNets simplify the connectivity pattern between layers introduced in other architectures:\nHighway Networks [2]\nResidual Networks [3]\nFractal Networks [4]\nThis solve the problem ensuring maximum information (and gradient) flow. To do it, the model is simply connected from every layer directly to the other ones. Instead of drawing representational power from extremely deep or wide architectures, DenseNets exploit the potential of the network through feature reuse. Counter-intuitively, by connecting this way DenseNets require fewer parameters than an equivalent traditional CNN, as there is no need to learn redundant feature maps","metadata":{"id":"G0be_-M0oZmJ"}},{"cell_type":"markdown","source":"We will start by setting a TPU","metadata":{"id":"mpnUtBSedNqS"}},{"cell_type":"code","source":"AUTO = tf.data.experimental.AUTOTUNE\ntpu = tf.distribute.cluster_resolver.TPUClusterResolver()\ntf.config.experimental_connect_to_cluster(tpu)\ntf.tpu.experimental.initialize_tpu_system(tpu)\nstrategy = tf.distribute.experimental.TPUStrategy(tpu)\nwarnings.filterwarnings(\"ignore\")","metadata":{"id":"3TLl23G4CwAD","outputId":"0ef3d876-3954-4c78-ad58-1aea07e15f5f","execution":{"iopub.status.busy":"2021-05-25T20:51:55.75854Z","iopub.status.idle":"2021-05-25T20:51:55.759102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will load the images and split them to train and validation","metadata":{"id":"qPW-NaH9dYJu"}},{"cell_type":"code","source":"img_size = 224","metadata":{"id":"g7q61qYWNMKk","execution":{"iopub.status.busy":"2021-05-25T20:51:55.760214Z","iopub.status.idle":"2021-05-25T20:51:55.76078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_224X224_images = train_data[\"image_id\"].progress_apply(load_image, args=((img_size,img_size),))","metadata":{"id":"2yynxPAsIWHl","outputId":"153923d4-83ee-4953-9e19-6ea30df66bb2","execution":{"iopub.status.busy":"2021-05-25T20:51:55.76198Z","iopub.status.idle":"2021-05-25T20:51:55.762562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"current_train_data=train_data.loc[train_224X224_images.index]\nimages_array=images_4d_array(train_224X224_images)\ntargets=current_train_data.drop('image_id',axis=1)\nx_train2, x_val, y_train2, y_val = sklearn.model_selection.train_test_split(images_array, targets, test_size=0.2, random_state=seed)\nx_train2 = x_train2.astype('float32')\nx_val = x_val.astype('float32')\nx_train2 /= 255\nx_val /= 255\nx_train2.shape, x_val.shape, y_train2.shape, y_val.shape","metadata":{"id":"NJS6BqsHH8IO","outputId":"d2637eb3-5e7a-453b-90d1-eb213131ff20","execution":{"iopub.status.busy":"2021-05-25T20:51:55.763637Z","iopub.status.idle":"2021-05-25T20:51:55.764207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will define the model and try to fit it","metadata":{"id":"3r9iRKv_eYg5"}},{"cell_type":"code","source":"with strategy.scope():\n    model = tf.keras.Sequential([tf.keras.applications.DenseNet121(input_shape=(224, 224, 3),\n                                             weights='imagenet',\n                                             include_top=False),\n                                 L.GlobalAveragePooling2D(),\n                                 L.Dense(y_train2.shape[1],\n                                         activation='softmax')])\n        \n    model.compile(optimizer='rmsprop',\n              loss='categorical_crossentropy',\n              metrics=['accuracy']\n              )\n    model.summary()","metadata":{"id":"7Hrh6aT319TT","outputId":"68f5b3b2-9b84-4561-f91e-bbcf78c0d9d0","execution":{"iopub.status.busy":"2021-05-25T20:51:55.765487Z","iopub.status.idle":"2021-05-25T20:51:55.766047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(x_train2,y_train2,\n                    epochs=250,\n                    callbacks=[ES_monitor,LR_reduce],\n                    steps_per_epoch=15,\n                    verbose=0,\n                    validation_data=(x_val, y_val))","metadata":{"id":"GOXKQzAzNenp","execution":{"iopub.status.busy":"2021-05-25T20:51:55.767287Z","iopub.status.idle":"2021-05-25T20:51:55.767851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets print our results","metadata":{"id":"h2VxAQxEi7a_"}},{"cell_type":"code","source":"score = model.evaluate(x_val, y_val, verbose=0)\n\nprint('Test loss:', score[0])\nprint('Test accuracy:', score[1])","metadata":{"id":"Wfd75UW5riNG","outputId":"ac4ca416-065f-45de-df40-36eec3ec8157","execution":{"iopub.status.busy":"2021-05-25T20:51:55.769023Z","iopub.status.idle":"2021-05-25T20:51:55.769611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_history(history)","metadata":{"id":"geDgZmfm_epj","outputId":"cbd23e13-72d6-4932-cd54-8142fb886985","execution":{"iopub.status.busy":"2021-05-25T20:51:55.77075Z","iopub.status.idle":"2021-05-25T20:51:55.771336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our accuracy has stayed 92%, and the fitting was much faster, but there are still some problems, we can see at the end that there i a gap between the validation and the train matrices, which can indicate that there is some kind of overfittin, the accuracy of the train data is almos perferct, which can show again over fitting, but we can see that this is the case from almot the beggining which indicate that it could have been because of the pre trained model and it's proficiency  ","metadata":{"id":"9_pcU_W0nV8g"}},{"cell_type":"markdown","source":"Lets take a look at the confusin matrix for more insight","metadata":{"id":"JfMCWA_XpK7a"}},{"cell_type":"code","source":"y_val_pred = model.predict_classes(x_val)\nmat = sklearn.metrics.confusion_matrix(get_target_array_no_im_id(y_val),y_val_pred)\nsns.heatmap(mat.T, square=True, annot=True, cbar=False)\nmatplotlib.pyplot.xlabel('true label')\nmatplotlib.pyplot.ylabel('predict label');","metadata":{"id":"HQ6pPm351Z6S","outputId":"c55e6000-2c27-486c-eb9a-b32013e0a76e","execution":{"iopub.status.busy":"2021-05-25T20:51:55.772451Z","iopub.status.idle":"2021-05-25T20:51:55.773009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The classes are ordered in this manner: \"healthy\" \"multiple_diseases\" \"rust\" \"scab\". We can see that our model is pretty accurate at idintifing the \"healthy\" \"rust\" \"scab\" classes, as it classified corretly almost all of the images of those classes. Again the \"multiple_diseases\" class images' classification is not better then random. It could be that the underrepresentation of that class is a big factor of that and we haven't found a model that whould help us with that.","metadata":{"id":"ZIo97uNFpiVE"}},{"cell_type":"code","source":"model.save(\"..output/kaggle/working/model_densnet121.h5\")\n","metadata":{"id":"JeFehfilCKq8","execution":{"iopub.status.busy":"2021-05-25T20:51:55.774046Z","iopub.status.idle":"2021-05-25T20:51:55.774761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **DenseNet With SMOTE**","metadata":{"id":"ssH6oFZI-TB7"}},{"cell_type":"markdown","source":"Our last attempt with SMOTE left us with mixed feeling about it, but it did had an impact on the \"multiple_diseases\" class, was the wannted one. So we will try it once more but with DenseNet this time ","metadata":{"id":"-hDgcAvhrljq"}},{"cell_type":"markdown","source":"Applying SMOTE on our dataset","metadata":{"id":"VbSUXSQIuFC-"}},{"cell_type":"code","source":"sm = imblearn.over_sampling.SMOTE(random_state = 115) \n \nx_train2, y_train2 = sm.fit_resample(get_all_pixels(x_train2),y_train2.to_numpy())\nx_train2 = x_train2.reshape((-1, img_size, img_size, 3))\nx_train2.shape, y_train2.sum(axis=0)","metadata":{"id":"TH285jk5-R14","outputId":"7bba1af7-4cef-40fa-dde7-2ec9d7d45466","execution":{"iopub.status.busy":"2021-05-25T20:51:55.776114Z","iopub.status.idle":"2021-05-25T20:51:55.776721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will define the model and try to fit it","metadata":{"id":"XRWmpHt6uZr2"}},{"cell_type":"code","source":"with strategy.scope():\n    model = tf.keras.Sequential([tf.keras.applications.DenseNet121(input_shape=(224, 224, 3),\n                                             weights='imagenet',\n                                             include_top=False),\n                                 L.GlobalAveragePooling2D(),\n                                 L.Dense(y_train2.shape[1],\n                                         activation='softmax')])\n        \n    model.compile(optimizer='rmsprop',\n              loss='categorical_crossentropy',\n              metrics=['accuracy']\n              )\n    model.summary()","metadata":{"id":"7txeetgR-8jq","outputId":"58143125-2962-4ee4-e4c8-4d95df829dd9","execution":{"iopub.status.busy":"2021-05-25T20:51:55.777951Z","iopub.status.idle":"2021-05-25T20:51:55.77862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(x_train2,y_train2,\n                    epochs=250,\n                    callbacks=[ES_monitor,LR_reduce],\n                    steps_per_epoch=15,\n                    verbose=0,\n                    validation_data=(x_val, y_val))","metadata":{"id":"GQjE9E-z-8jr","execution":{"iopub.status.busy":"2021-05-25T20:51:55.779802Z","iopub.status.idle":"2021-05-25T20:51:55.780382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets print our results","metadata":{"id":"hnsvZAaJudZY"}},{"cell_type":"code","source":"score = model.evaluate(x_val, y_val, verbose=0)\n\nprint('Test loss:', score[0])\nprint('Test accuracy:', score[1])","metadata":{"id":"-AkzwsG--8jr","outputId":"a957576f-e36c-4d61-a474-073969ab8ad4","execution":{"iopub.status.busy":"2021-05-25T20:51:55.781609Z","iopub.status.idle":"2021-05-25T20:51:55.782158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_history(history)","metadata":{"id":"eU4SX-lQ-8jr","outputId":"023ae301-9eee-4e0b-d7ab-7a67b96ba094","execution":{"iopub.status.busy":"2021-05-25T20:51:55.783422Z","iopub.status.idle":"2021-05-25T20:51:55.783985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our accuracy has stayed 92%, but there are the same problems, we can see at the end that there is a gap between the validation and the train matrices, which can indicate that there is some kind of overfittin, the accuracy of the train data is almos perferct, which can show again over fitting, but we can see that this is the case from almot the beggining which indicate that it could have been because of the pre trained model and it's proficiency","metadata":{"id":"pQdQCorBup5Z"}},{"cell_type":"markdown","source":"Lets take a look at the confusin matrix for more insight","metadata":{"id":"Svqf6A_lvFdG"}},{"cell_type":"code","source":"y_val_pred = model.predict_classes(x_val)\nmat = sklearn.metrics.confusion_matrix(get_target_array_no_im_id(y_val),y_val_pred)\nsns.heatmap(mat.T, square=True, annot=True, cbar=False)\nmatplotlib.pyplot.xlabel('true label')\nmatplotlib.pyplot.ylabel('predict label');","metadata":{"id":"p4hmWV-x-8jr","outputId":"ab3a2f34-d32c-4314-c412-c26d0fb704d3","execution":{"iopub.status.busy":"2021-05-25T20:51:55.785764Z","iopub.status.idle":"2021-05-25T20:51:55.786398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The classes are ordered in this manner: \"healthy\" \"multiple_diseases\" \"rust\" \"scab\". We can see that there were minimal change to the classification and that SMOTE did't help us classify the \"multiple_diseases\" better","metadata":{"id":"MMu43XoSvG04"}},{"cell_type":"code","source":"model.save(\"..output/kaggle/working/model_densnet121_SMOTE.h5\")","metadata":{"id":"vBuJ8_b9-8jr","execution":{"iopub.status.busy":"2021-05-25T20:51:55.787894Z","iopub.status.idle":"2021-05-25T20:51:55.788469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **DenseNet With Data Transformations and Augmentation**","metadata":{"id":"9P0xFybvyfcx"}},{"cell_type":"markdown","source":"As using DenseNet with SMOTE didn't produced a better model, we will use Data Transformations and Augmentation as well. Maybe that could produce a better model. We Didn't use it in at first place because using TPU and ImageDataGenerator is no possible, and using a GPU take considerebly longer.","metadata":{"id":"mHgmK_okzPXj"}},{"cell_type":"markdown","source":"Loading the images and spliting them to train and test","metadata":{"id":"fptobmiG0MDR"}},{"cell_type":"code","source":"img_size = 224","metadata":{"id":"IYYBUMjWzPro","execution":{"iopub.status.busy":"2021-05-25T20:51:55.790108Z","iopub.status.idle":"2021-05-25T20:51:55.790788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_224X224_images = train_data[\"image_id\"].progress_apply(load_image, args=((img_size,img_size),))","metadata":{"id":"Blnl9Eo1zPro","outputId":"db3ccbe7-4232-4916-8b89-b9d9470b3e6b","execution":{"iopub.status.busy":"2021-05-25T20:51:55.79202Z","iopub.status.idle":"2021-05-25T20:51:55.792611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"current_train_data=train_data.loc[train_224X224_images.index]\nimages_array=images_4d_array(train_224X224_images)\ntargets=current_train_data.drop('image_id',axis=1)\nx_train2, x_val, y_train2, y_val = sklearn.model_selection.train_test_split(images_array, targets, test_size=0.2, random_state=seed)\nx_train2 = x_train2.astype('float32')\nx_val = x_val.astype('float32')\nx_train2 /= 255\nx_val /= 255\nx_train2.shape, x_val.shape, y_train2.shape, y_val.shape","metadata":{"id":"QUn0fyBazPrp","outputId":"9f5cfdfa-0a40-407e-9da0-4d4648734ab9","execution":{"iopub.status.busy":"2021-05-25T20:51:55.793844Z","iopub.status.idle":"2021-05-25T20:51:55.79442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will define the model and try to fit it","metadata":{"id":"hng6IfVKzPrq"}},{"cell_type":"code","source":"\nmodel = tf.keras.Sequential([tf.keras.applications.DenseNet121(input_shape=(224, 224, 3),\n                                          weights='imagenet',\n                                          include_top=False),\n                              L.GlobalAveragePooling2D(),\n                              L.Dense(y_train2.shape[1],\n                                      activation='softmax')])\n    \nmodel.compile(optimizer='rmsprop',\n          loss='categorical_crossentropy',\n          metrics=['accuracy']\n          )\nmodel.summary()","metadata":{"id":"MMnp0ix1zPrq","outputId":"59cb6b3c-b18f-43eb-bb27-df1e5bb5c57d","execution":{"iopub.status.busy":"2021-05-25T20:51:55.795487Z","iopub.status.idle":"2021-05-25T20:51:55.796143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(datagen.flow(x_train2,y_train2,batch_size=24),\n                    epochs=250,\n                    callbacks=[ES_monitor,LR_reduce],\n                    steps_per_epoch=15,\n                    verbose=0,\n                    validation_data=(x_val, y_val))","metadata":{"id":"6qKdNefOzPrq","execution":{"iopub.status.busy":"2021-05-25T20:51:55.797462Z","iopub.status.idle":"2021-05-25T20:51:55.798025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets print our results","metadata":{"id":"n2YxQvmu0Ygm"}},{"cell_type":"code","source":"score = model.evaluate(x_val, y_val, verbose=0)\n\nprint('Test loss:', score[0])\nprint('Test accuracy:', score[1])","metadata":{"id":"AxJKfj-YzPrq","outputId":"ca44df2b-e2cb-4624-a089-5bb37837a753","execution":{"iopub.status.busy":"2021-05-25T20:51:55.799239Z","iopub.status.idle":"2021-05-25T20:51:55.799795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_history(history)","metadata":{"id":"6dKbl1i_zPrq","outputId":"4b838126-e051-496e-fbc2-b458570ef1b6","execution":{"iopub.status.busy":"2021-05-25T20:51:55.801102Z","iopub.status.idle":"2021-05-25T20:51:55.801686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our accuracy has improved to 94%, the best accuracy thus far we can see that the vlidation and test graph have no gap which can show that the model didn't overpitted to the test. We can see that the model is having a better time classifing the \"multiple_diseases\" class then ever before from the better AUC-ROC score.","metadata":{"id":"xv-1_MVQ0vwc"}},{"cell_type":"markdown","source":"Lets take a look at the confusin matrix for more insight","metadata":{"id":"GmoV3E0T1j-9"}},{"cell_type":"code","source":"y_val_pred = model.predict_classes(x_val)\nmat = sklearn.metrics.confusion_matrix(get_target_array_no_im_id(y_val),y_val_pred)\nsns.heatmap(mat.T, square=True, annot=True, cbar=False)\nmatplotlib.pyplot.xlabel('true label')\nmatplotlib.pyplot.ylabel('predict label');","metadata":{"id":"NZl91ns-zPrq","outputId":"8ed202c2-caae-4e55-b15b-f77829da27f2","execution":{"iopub.status.busy":"2021-05-25T20:51:55.803281Z","iopub.status.idle":"2021-05-25T20:51:55.803829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The classes are ordered in this manner: \"healthy\" \"multiple_diseases\" \"rust\" \"scab\". We can see that our new model is better at idintifing the \"multiple_diseases\" class while classifing the other classes, then any other model we produced. It main problem is still this class but it does much better job at classifing it.","metadata":{"id":"JBDgKDnp2YWg"}},{"cell_type":"code","source":"model.save(\"..output/kaggle/working/model_densnet121_DG.h5\")\n","metadata":{"id":"fr7JFThKzPrr","execution":{"iopub.status.busy":"2021-05-25T20:51:55.805008Z","iopub.status.idle":"2021-05-25T20:51:55.805631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **DenseNet With Data Transformations and Augmentation and With  SMOTE**","metadata":{"id":"_dMd6Aq93Gpn"}},{"cell_type":"markdown","source":"Our last model did great, but still can have difficulties in classifing the \"multiple_diseases\". Even though SMOTE has faild us one, and had no effect the other time, e will try to use SMOTE to better our model, we don't have great hopes for it to work but maybe it wiil prove us wrong ","metadata":{"id":"q31pkMwS33Ai"}},{"cell_type":"markdown","source":"Again we will use the same procider as before but we will use SMOTE as well","metadata":{"id":"rUBRXCeK5O1-"}},{"cell_type":"markdown","source":"Applying SMOTE on our dataset","metadata":{"id":"Z3OMFYeh5eE_"}},{"cell_type":"code","source":"sm = imblearn.over_sampling.SMOTE(random_state = 115) \n \nx_train2, y_train2 = sm.fit_resample(get_all_pixels(x_train2),y_train2.to_numpy())\nx_train2 = x_train2.reshape((-1, img_size, img_size, 3))\nx_train2.shape, y_train2.sum(axis=0)","metadata":{"id":"W_-bbOWH3Gpo","outputId":"df67216e-1d15-486e-d414-ab95054cbc5d","execution":{"iopub.status.busy":"2021-05-25T20:51:55.807744Z","iopub.status.idle":"2021-05-25T20:51:55.808331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will define the model and try to fit it","metadata":{"id":"O1BXmV0K5mF5"}},{"cell_type":"code","source":"\nmodel = tf.keras.Sequential([tf.keras.applications.DenseNet121(input_shape=(224, 224, 3),\n                                          weights='imagenet',\n                                          include_top=False),\n                              L.GlobalAveragePooling2D(),\n                              L.Dense(y_train2.shape[1],\n                                      activation='softmax')])\n    \nmodel.compile(optimizer='rmsprop',\n          loss='categorical_crossentropy',\n          metrics=['accuracy']\n          )\nmodel.summary()","metadata":{"id":"cdWgfxHw3Gpp","outputId":"07d64474-ebce-4327-882c-4ba1c9fc4f99","execution":{"iopub.status.busy":"2021-05-25T20:51:55.809721Z","iopub.status.idle":"2021-05-25T20:51:55.810312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(datagen.flow(x_train2,y_train2,batch_size=24),\n                    epochs=250,\n                    callbacks=[ES_monitor,LR_reduce],\n                    steps_per_epoch=15,\n                    verbose=0,\n                    validation_data=(x_val, y_val))","metadata":{"id":"oldG8R0Q3Gpp","execution":{"iopub.status.busy":"2021-05-25T20:51:55.811555Z","iopub.status.idle":"2021-05-25T20:51:55.812117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets print our results","metadata":{"id":"aRyBmIc55XLZ"}},{"cell_type":"code","source":"score = model.evaluate(x_val, y_val, verbose=0)\n\nprint('Test loss:', score[0])\nprint('Test accuracy:', score[1])","metadata":{"id":"fIP7s1Ji3Gpp","outputId":"5d6a1226-ac24-41a0-e81b-31e2501c921d","execution":{"iopub.status.busy":"2021-05-25T20:51:55.813412Z","iopub.status.idle":"2021-05-25T20:51:55.814009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_history(history)","metadata":{"id":"tySfs7IV3Gpq","outputId":"3c2aa91b-8f31-4106-88e4-60b837aa1f8b","execution":{"iopub.status.busy":"2021-05-25T20:51:55.815437Z","iopub.status.idle":"2021-05-25T20:51:55.816215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"SMOTE didn't fail us and didn't help us. our accuracy has decreased to 88%, not the worst but a downgarde from our privious model. lets take a look at the confusin matrix for more insight about what went wrong","metadata":{"id":"s9IISxF25xQS"}},{"cell_type":"code","source":"y_val_pred = model.predict_classes(x_val)\nmat = sklearn.metrics.confusion_matrix(get_target_array_no_im_id(y_val),y_val_pred)\nsns.heatmap(mat.T, square=True, annot=True, cbar=False)\nmatplotlib.pyplot.xlabel('true label')\nmatplotlib.pyplot.ylabel('predict label');","metadata":{"id":"DPEMM-Ey3Gpq","outputId":"6470c374-4994-4420-ddf5-2abae728f9db","execution":{"iopub.status.busy":"2021-05-25T20:51:55.817515Z","iopub.status.idle":"2021-05-25T20:51:55.817882Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The classes are ordered in this manner: \"healthy\" \"multiple_diseases\" \"rust\" \"scab\". We can see that our new model is worse then the one before at any classification. We can see that the SMOTE didnt had the effect that we wanted and it had all the negative effects from the first time we used it","metadata":{"id":"A95w83MO6qcW"}},{"cell_type":"code","source":"model.save(\"..output/kaggle/working/model_densnet121_SMOTE_DG.h5\")\n","metadata":{"id":"Zm2QjDxg3Gpq","execution":{"iopub.status.busy":"2021-05-25T20:51:55.818581Z","iopub.status.idle":"2021-05-25T20:51:55.818902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Discussion**","metadata":{"id":"fQOn6EM0wj1l"}},{"cell_type":"markdown","source":"Our goal in this project was to produce a model that can classify leaves images to 4 classes of health condition. The data set we had wes imbalance, one of the classes had drastically less images. Throughout the project that fact had great influence on our models and success. At first we explored our data and noticed that there is color difference between classes, and thought that it could heip us, but as the PCA showed us, the color of the image is greatly infuenced by the background and it has a greater effect then the symptoms of the diseases on the leaves. We tried to negate this  effect with 2 methods, cropping the leaves out ot the image and by blacking out the background, non of those methods were successful. Per image we can use the methods to great extent, but automatin those methods prove too sofisticated for us, because of the color of the background compare to the leaves, because of the curved shape of the leave which made their edges lose focus in some instances, which meant that their veins were more prominent as an edge.\n\nAs of next we tried some classic prediction models, which didn't go good at all, and gave prediction that are good as guessing.\n\nNext we tried the the most common method for classifing images, neural network. At first we tried a basic model that didn't performd well at all, next we continued to convolutional neural network, a great method for image classifing, that wes the first model that gave us good results. We tried at first to use it on images with the size of 100X100 pixels, it had good results but it misclassified some diseased infected leave as healthy, but more noticeably it didn't classified the \"multiple_diseases\" class well at all, due to it being underrepresented and having features of other classes. Using bigger images, 224X244 pixels to be specific, we managed to reduce it's misclassification of the diseased leaves as healthy, but didn't change the \"multiple_diseases\" problem, we saw as well some signs of overfiting to the train set. \n\nAs we saw the problems in our previous model, we use data augmentatinons, we changed images rotation, orientation, brightness and scale, and fed them to the CNN. This method improved our model and msde his be more presice when classifing diseased leaves, but once again we colud not classify the \"multiple_diseases\" class well.\n\nWe tried next a method that could maybe help us with this problem, we used SMOTE, it's a method that balances our data by generating samples to the minor class, but as it turn out this method just made the model worse. The problem was that our \"multiple_diseases\" class has similar featurs as the other two diseases calsses, and by generating the samples it made those features ones that assosiate greatly with the \"multiple_diseases\" class, and not the other diseased classes.\n\nThus far our best attempt ,data augmentation, was pretty good but had some difficulties and could not tackle the main issue of the data set, it's imbalance. \n\nNext we tried to incorporate transfer learning to our attempts. This methos is using an allready traind model on general case as our base model, and in addition we used a different neural network, denstenet which has different kind of architecture the our CNN. We tried using it with or without SMOTE and with or without data augmentation. The best model was produced as a result of the the one with data augmentation and without SMOTE. It can tackle the \"multiple_diseases\" class much better but still has a bit of dificult with this class, at the rest it does great work.\n\n\nOur data main challenges were it's imbalance and the fact that the leaves were filmed on a dinamic background. Further more, another challenge is the resemblance of some of its classes features. All of the above made our work interesting and challenging. We could not find an efficient way to reduce the impact that the back ground has, and could not find a way that could heip us with the imbalances without reducing model accurecy. The first problem could be solved by creating a model that can by himself black out the background, but that problem is another big one that would required a project by itself. The secound problem could be tackled by gathering more samples of this class, a mission for plant disease expert that can identify the correct disease. \n\nOverall we produced a competant model that can in great success predict the disease an apple leaf has.\n\n","metadata":{"id":"rO18hcJyxbvn"}},{"cell_type":"markdown","source":"# **Conclusion**","metadata":{"id":"s2H_Fh9No9og"}},{"cell_type":"markdown","source":"This project was a great challenge but a great fun to work on. We've learned a lot and were able to dive into the world of data science. We learned that some task that look trivial like identifing a leaf and differentiating it from the background can be convoluted if the background has certain features. We also learnd that in the case that we have two classes with shared features using sample genaration would probebly wont work. We saw first-handed how promenent of a  solution neural networks for classifing images, especially over classic methods. We found out that managing resources is a big part of a data science project and can have big impect on our time and effort.","metadata":{"id":"Q_YMnbXaphqA"}}]}