{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Cassava Leaf Disease Classification\n\nAs the second-largest provider of carbohydrates in Africa, cassava is a key food security crop grown by smallholder farmers because it can withstand harsh conditions. At least 80% of household farms in Sub-Saharan Africa grow this starchy root, but viral diseases are major sources of poor yields. With the help of data science, it may be possible to identify common diseases so they can be treated.\n\nExisting methods of disease detection require farmers to solicit the help of government-funded agricultural experts to visually inspect and diagnose the plants. This suffers from being labor-intensive, low-supply and costly. As an added challenge, effective solutions for farmers must perform well under significant constraints, since African farmers may only have access to mobile-quality cameras with low-bandwidth.\n\nThe dataset is of 21,367 labeled images collected during a regular survey in Uganda. Most images were crowdsourced from farmers taking photos of their gardens, and annotated by experts at the National Crops Resources Research Institute (NaCRRI) in collaboration with the AI lab at Makerere University, Kampala.\n\nSo in this competition we have classify each of the image into four disease categories or a fifth category indicating a healthy leaf through Image Processing Techniques and Deep Lerning.\n\n<img src=\"https://www.rural21.com/fileadmin/_processed_/5/3/csm_Science_57_cassava_shutterstock_83105608_Kopie_01_9817f8d733.jpg\" style = 'width:750px;height:450px;'>"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"# import the required libraries\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom tensorflow import keras\nimport seaborn as sns\nimport os\nimport random\nimport cv2","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Load the dataset\nFollowing is the mapping corresponding to the labels -\n\n    0 - Cassava Bacterial Blight (CBB)\n    1 - Cassava Brown Streak Disease (CBSD)\n    2 - Cassava Green Mottle (CGM)\n    3 - Cassava Mosaic Disease (CMD)\n    4 - Healthy"},{"metadata":{"trusted":true},"cell_type":"code","source":"labels = {0: 'CBB (Cassava Bacterial Blight)', 1: 'CBSD (Cassava Brown Streak Disease)', 2: 'CGM (Cassava Green Mottle)', 3: 'CMD (Cassava Mosaic Disease)', 4: 'Healthy'}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"DIR = '../input/cassava-leaf-disease-classification/train_images/'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# read the output labels from the csv file\ndf = pd.read_csv('../input/cassava-leaf-disease-classification/train.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"Y = df['label'].values","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"Y shape:\" + str(Y.shape))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Data Visualization"},{"metadata":{"trusted":true},"cell_type":"code","source":"# count of all the diseases in our dataset\nsns.countplot(x = 'label', data = df);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# function to plot images of different classes\ndef plotImages(class_):\n    print('Some images of the class {}'.format(labels[class_]))\n    \n    images = df[df.label == class_].sample(8)\n    \n    fig, ax = plt.subplots(2, 4, figsize = (16,6))\n    for i in range(2):\n        for j in range(4):\n            path = os.path.join(DIR, images['image_id'].iloc[i * 4 + j])\n            img = cv2.imread(path, cv2.IMREAD_COLOR)\n            ax[i, j].imshow(img)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# CBB (Cassava Bacterial Blight)\nplotImages(0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# CBSD (Cassava Brown Streak Disease)\nplotImages(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# CGM (Cassava Green Mottle)\nplotImages(2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# CMD (Cassava Mosaic Disease)\nplotImages(3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Healthy leaves\nplotImages(4)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Processing of data"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Main paramaters\nBATCH_SIZE = 64\nIMG_SIZE = 300\nEPOCHS = 10","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from keras.preprocessing.image import ImageDataGenerator","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.label = df.label.astype('str')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# define the data denerators for training and vlaidation data\ntrain_datagen = ImageDataGenerator(validation_split = 0.2,\n                                horizontal_flip = True,\n                                vertical_flip = True,\n                                rotation_range = 30,\n                                width_shift_range = 0.1,\n                                height_shift_range = 0.1,\n                                fill_mode = 'nearest',\n                                zoom_range = 0.2)\n\ntrain_generator = train_datagen.flow_from_dataframe(df, \n                                                    directory = DIR, \n                                                    x_col = 'image_id', \n                                                    y_col = 'label', \n                                                    subset = \"training\",\n                                                    batch_size = BATCH_SIZE,\n                                                    target_size = (IMG_SIZE, IMG_SIZE))\n\n\nval_datagen = ImageDataGenerator(validation_split = 0.2)\n\nval_generator = val_datagen.flow_from_dataframe(df, \n                                                directory = DIR,\n                                                x_col = 'image_id',\n                                                y_col = 'label', \n                                                subset = \"validation\",\n                                                batch_size = BATCH_SIZE,\n                                                target_size = (IMG_SIZE, IMG_SIZE))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# reshape the data\nY = Y.reshape(-1, 1)\nY = keras.utils.to_categorical(Y, 5)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Model implementation"},{"metadata":{},"cell_type":"markdown","source":"### 1. CNN Model from scratch"},{"metadata":{"trusted":true},"cell_type":"code","source":"# import the required libraries\nfrom keras.models import Sequential\nfrom keras.layers import Conv2D, MaxPool2D, Dropout, Flatten\nfrom keras.layers import Dense","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model = Sequential(\n    [\n        Conv2D(filters = 32, kernel_size = (3, 3), padding = 'same', activation = 'relu', input_shape = (IMG_SIZE, IMG_SIZE, 3)),\n        MaxPool2D(pool_size = (2, 2), strides = (2, 2)),\n        \n        Conv2D(filters = 64, kernel_size = (3, 3), padding = 'same', activation = 'relu'),\n        MaxPool2D(pool_size = (2, 2), strides = (2, 2)),\n        \n        Conv2D(filters = 128, kernel_size = (3, 3), padding = 'same', activation = 'relu'),\n        MaxPool2D(pool_size = (2, 2), strides = (2, 2)),\n        \n        Conv2D(filters = 256, kernel_size = (3, 3), padding = 'same', activation = 'relu'),\n        MaxPool2D(pool_size = (2, 2), strides = (2, 2)),\n        \n        Conv2D(filters = 512, kernel_size = (3, 3), padding = 'same', activation = 'relu'),\n        MaxPool2D(pool_size = (2, 2), strides = (2, 2)),\n        \n        Flatten(),\n        Dense(1024, activation = 'relu'),\n        Dropout(0.5),\n        Dense(5, activation = 'softmax')\n    ]\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# compile the model\nmodel.compile(optimizer = 'adam', loss = 'categorical_crossentropy', metrics = ['accuracy'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"history = model.fit_generator(train_generator, epochs = 5, validation_data = val_generator)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# plot training and validation accuracy values\nplt.plot(history.history['accuracy'])\nplt.plot(history.history['val_accuracy'])\nplt.title('Model accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend(['Train', 'Val'], loc = 'upper left')\nplt.show()\n\n# plot training and validation loss values\nplt.plot(history.history['loss'])\nplt.plot(history.history['val_loss'])\nplt.title('Model loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend(['Train', 'Val'], loc = 'upper left')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"loss = model.evaluate_generator(val_generator, steps=24)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## 2. Using EfficientNetB3 model"},{"metadata":{"trusted":true},"cell_type":"code","source":"# import the required libraries\nfrom tensorflow.keras.applications import EfficientNetB3\nfrom keras.layers import GlobalAveragePooling2D, Dropout\nfrom tensorflow.keras.callbacks import ModelCheckpoint, EarlyStopping, ReduceLROnPlateau","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"base_model = EfficientNetB3(include_top = False, input_tensor = None, input_shape = (IMG_SIZE, IMG_SIZE, 3))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"len(base_model.layers)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# To reduce the runtime we don't train the initial 126 layers\nfor layer in base_model.layers[:126]:\n    layer.trainable = False","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model = Sequential(\n    [\n        base_model,\n        GlobalAveragePooling2D(),\n        Dense(256, activation = 'relu'),\n        Dropout(0.4),\n        Dense(5, activation = 'softmax')\n    ]\n)\n\n# compile the model\nmodel.compile(optimizer = 'adam', loss = 'categorical_crossentropy', metrics = ['accuracy'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model.summary()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Callbacks"},{"metadata":{"trusted":true},"cell_type":"code","source":"weights_path = './eff_net.h5'\n\n# Model Checkpont for saving the best model\nsave_model = ModelCheckpoint(weights_path, \n                             monitor = 'val_loss',\n                             save_best_only = True, \n                             mode = 'min', \n                             save_weights_only = False,\n                             verbose = 1)\n\n# Early Stopping\nearly_stop = EarlyStopping(monitor = 'val_loss', \n                      mode = 'min',\n                      restore_best_weights = True,\n                      patience = 5,\n                      min_delta = 0.001,\n                      verbose = 1)\n\n# Reducing the learning rate on Plateau\nreduceLR = ReduceLROnPlateau(monitor='val_loss', \n                                   factor = 0.8, \n                                   patience = 4, \n                                   mode = 'min', \n                                   min_delta = 0.0001, \n                                   cooldown = 5)\n\ncallbacks = [save_model, early_stop, reduceLR]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"history = model.fit(train_generator, epochs = EPOCHS, validation_data = val_generator, callbacks = callbacks)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# plot training and validation accuracy values\nplt.plot(history.history['accuracy'])\nplt.plot(history.history['val_accuracy'])\nplt.title('Model accuracy')\nplt.xlabel('Epoch')\nplt.ylabel('Accuracy')\nplt.legend(['Train', 'Val'], loc = 'upper left')\nplt.show()\n\n# plot training and validation loss values\nplt.plot(history.history['loss'])\nplt.plot(history.history['val_loss'])\nplt.title('Model loss')\nplt.xlabel('Epoch')\nplt.ylabel('Loss')\nplt.legend(['Train', 'Val'], loc = 'upper left')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Getting the outputs"},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df = pd.read_csv('../input/cassava-leaf-disease-classification/sample_submission.csv')\ntest_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# get the predicted output of all the images in the test_images directory\npath = '../input/cassava-leaf-disease-classification/test_images/'\n\npreds = []\n\nfor image_id in test_df.image_id:\n    image = cv2.imread(os.path.join(path, image_id))\n    image = cv2.resize(image, (IMG_SIZE, IMG_SIZE))\n    image = np.expand_dims(image, axis = 0)\n    preds.append(np.argmax(model.predict(image)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"preds","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}