{"cells":[{"metadata":{},"cell_type":"markdown","source":"Overview:\n\n***As the second-largest provider of carbohydrates in Africa, cassava is a key food security crop grown by smallholder farmers because it can withstand harsh conditions. At least 80% of household farms in Sub-Saharan Africa grow this starchy root, but viral diseases are major sources of poor yields. With the help of data science, it may be possible to identify common diseases so they can be treated.\n\nExisting methods of disease detection require farmers to solicit the help of government-funded agricultural experts to visually inspect and diagnose the plants. This suffers from being labor-intensive, low-supply and costly. As an added challenge, effective solutions for farmers must perform well under significant constraints, since African farmers may only have access to mobile-quality cameras with low-bandwidth.\n\nIn this competition, we introduce a dataset of 21,367 labeled images collected during a regular survey in Uganda. Most images were crowdsourced from farmers taking photos of their gardens, and annotated by experts at the National Crops Resources Research Institute (NaCRRI) in collaboration with the AI lab at Makerere University, Kampala. This is in a format that most realistically represents what farmers would need to diagnose in real life.\n\nYour task is to classify each cassava image into four disease categories or a fifth category indicating a healthy leaf. With your help, farmers may be able to quickly identify diseased plants, potentially saving their crops before they inflict irreparable damage.***"},{"metadata":{"trusted":true},"cell_type":"code","source":"# Lets check the GPU provided\n!nvidia-smi ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# from kaggle_secrets import UserSecretsClient\n# user_secrets = UserSecretsClient()\n# user_credential = user_secrets.get_gcloud_credential()\n# user_secrets.set_tensorflow_credential(user_credential)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!ls /kaggle/input","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# Import all the directories\nimport os\nimport numpy as np \nimport pandas as pd \nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nfrom tensorflow.keras import applications\nfrom tensorflow.keras import layers\nfrom tensorflow.keras import models\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nimport warnings\nimport json\nimport cv2\n\nwarnings.filterwarnings('ignore')\n%matplotlib inline\n\nIMAGE_SIZE=128","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Lets check the tensorflow version\ntf.__version__","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# GPU Initialize\ndevice_name = tf.test.gpu_device_name()\nif device_name!='/device:GPU:0':\n    raise SystemError('GPU Device not found')\nprint('Found GPU at:{}'.format(device_name))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"# Lets initialize the parent dir\nPARENT_DIR = '../input/cassava-leaf-disease-classification'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# List folders are files\nprint(os.listdir(PARENT_DIR))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Import train and sample csv\ntrain_df = pd.read_csv(os.path.join(PARENT_DIR,'train.csv'))\nsample_df = pd.read_csv(os.path.join(PARENT_DIR,'sample_submission.csv'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Reading Json file and check the mapping of labels\nwith open(os.path.join(PARENT_DIR, \"label_num_to_disease_map.json\")) as jfile:\n    map_classes = json.loads(jfile.read())\n\nprint(map_classes)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Take a look on training csv\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Lets take a look into the images\ntrain_df_0 = train_df[train_df['label']==0].head(10).image_id\ntrain_df_1 = train_df[train_df['label']==1].head(10).image_id\ntrain_df_2 = train_df[train_df['label']==2].head(10).image_id\ntrain_df_3 = train_df[train_df['label']==3].head(10).image_id\ntrain_df_4 = train_df[train_df['label']==4].head(10).image_id","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def show_image(img_dir):\n    i_dir = img_dir\n    train_dir= PARENT_DIR +'/'+'train_images'\n    i = 1\n    plt.figure(figsize=(20,10))\n    for img in i_dir:\n        img = cv2.imread(os.path.join(train_dir,img),cv2.COLOR_BGR2RGB)\n        img = cv2.resize(img, (IMAGE_SIZE, IMAGE_SIZE),interpolation = cv2.INTER_NEAREST)\n        plt.subplot(2,5,i)\n        plt.imshow(img)\n        i+=1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Images for --Cassava Bacterial Blight (CBB)\nshow_image(train_df_0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Images for --Cassava Brown Streak Disease (CBSD)\nshow_image(train_df_1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Images for --Cassava Green Mottle (CGM)\nshow_image(train_df_2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Images for --Cassava Mosaic Disease (CMD)\nshow_image(train_df_3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Images for --Healthy\nshow_image(train_df_4)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Mapping numbers with labels\ntrain_df['label'] = train_df.label.map({0: 'Cassava Bacterial Blight (CBB)',\n                    1: 'Cassava Brown Streak Disease (CBSD)',\n                    2: 'Cassava Green Mottle (CGM)',\n                    3: 'Cassava Mosaic Disease (CMD)', \n                    4: 'Healthy'})","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Check for class imbalence will handle it using agumentation in ImageDataGenerator\nplt.figure(figsize=(20,10))\nsns.countplot(train_df['label'])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Creating training validation and test generator\ndatagen = ImageDataGenerator(\n                    rotation_range = 40,\n                    width_shift_range = 0.2,\n                    height_shift_range = 0.2,\n                    shear_range = 0.2,\n                    zoom_range = 0.2,\n                    horizontal_flip = True,\n                    vertical_flip = True,\n                    fill_mode = 'nearest',\n                    validation_split=0.25\n                    )\n\ntrain_generator=datagen.flow_from_dataframe(\n                    dataframe=train_df,\n                    directory=\"../input/cassava-leaf-disease-classification/train_images/\",\n                    x_col=\"image_id\",\n                    y_col=\"label\",\n                    subset=\"training\",\n                    batch_size=32,\n                    seed=42,\n                    shuffle=True,\n                    class_mode = 'categorical',\n                    color_mode='rgb',\n                    target_size=(IMAGE_SIZE,IMAGE_SIZE)\n                    )\n\n\nval_generator=datagen.flow_from_dataframe(\n                    dataframe=train_df,\n                    directory=\"../input/cassava-leaf-disease-classification/train_images/\",\n                    x_col=\"image_id\",\n                    y_col=\"label\",\n                    subset=\"validation\",\n                    batch_size=32,\n                    seed=42,\n                    shuffle=True,\n                    class_mode=\"categorical\",\n                    color_mode='rgb',\n                    target_size=(IMAGE_SIZE,IMAGE_SIZE)\n                    )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Mapping numbers with test labels\nsample_df['label'] = sample_df.label.map({0: 'Cassava Bacterial Blight (CBB)',\n                    1: 'Cassava Brown Streak Disease (CBSD)',\n                    2: 'Cassava Green Mottle (CGM)',\n                    3: 'Cassava Mosaic Disease (CMD)', \n                    4: 'Healthy'})","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"'''\ntest_datagen=tf.keras.preprocessing.image.ImageDataGenerator(rescale=1./255.)\ntest_generator=test_datagen.flow_from_dataframe(\n    dataframe=sample_df,\n    directory='../input/cassava-leaf-disease-classification/test_images/',\n    x_col=\"image_id\",\n    y_col=None,\n    batch_size=32,\n    seed=42,\n    shuffle=False,\n    class_mode=None,\n    target_size=(IMAGE_SIZE,IMAGE_SIZE)\n)\n'''","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Lets check the size of one batch\nfor img,lab in train_generator:\n#     print(lab)\n    print(img.shape)\n    break","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Define CallBacks\ncheckpoint_cb = tf.keras.callbacks.ModelCheckpoint(\n    \"InceptionV3_Model.h5\",\n    save_best_only=True,\n    monitor = 'val_loss',\n    mode='min'\n)\nreduce_lr = tf.keras.callbacks.ReduceLROnPlateau(monitor = 'val_loss',\n                                  factor = 0.3,\n                                  patience = 3,\n                                  min_lr = 1e-5,\n                                  mode = 'min',\n                                  verbose = 1)\n\nearly_stopping_cb = tf.keras.callbacks.EarlyStopping(\n    monitor='val_loss',\n    mode='min', \n    patience=5,\n    restore_best_weights=True, \n    verbose=1\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Perform training \nBATCH_SIZE = 32\nwith tf.device('/gpu:0'):\n    model = tf.keras.Sequential([\n        tf.keras.applications.EfficientNetB7(\n            input_shape=(IMAGE_SIZE, IMAGE_SIZE, 3),\n            weights=None,\n            include_top=False\n        #    drop_connect_rate=0.7\n        ),\n        tf.keras.layers.GlobalAveragePooling2D(),\n        tf.keras.layers.Flatten(),\n        tf.keras.layers.Dense(512,\n                              activation = 'relu', \n                              bias_regularizer=tf.keras.regularizers.l1_l2(l1=0.01,\n                                                                           l2=0.001)),\n        tf.keras.layers.Dropout(0.7),\n        tf.keras.layers.Dense(5, activation='softmax')\n    ])\n    model.add_weight('../input/tfkerasefficientnetimagenetnotop/efficientnetb7_notop.h5')\n    model.compile(\n        optimizer=tf.keras.optimizers.Adam(learning_rate = 1e-3),\n        loss='categorical_crossentropy',\n        metrics=['categorical_accuracy'])\n    history = model.fit(\n            train_generator,\n            steps_per_epoch = train_generator.n/BATCH_SIZE,\n            epochs=20,\n            batch_size = BATCH_SIZE,\n            validation_data=val_generator,\n            validation_steps = val_generator.n/BATCH_SIZE,\n            callbacks=[checkpoint_cb,reduce_lr,early_stopping_cb])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Plotting accuracy history\nplt.figure(figsize= (15,10))\nplt.plot(history.history['categorical_accuracy'])\nplt.plot(history.history['val_categorical_accuracy'])\nplt.title('Accuracy Tracker', fontsize=15)\nplt.xlabel('Epochs', fontsize=15)\nplt.ylabel('Accuracy', fontsize=15)\nplt.legend(['training', 'validation'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Plotting loss history\nplt.figure(figsize= (15,10))\nplt.plot(history.history['loss'])\nplt.plot(history.history['val_loss'])\nplt.title('Loss Tracker', fontsize=15)\nplt.xlabel('Epochs', fontsize=15)\nplt.ylabel('Loss', fontsize=15)\nplt.legend(['training', 'validation'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission = pd.DataFrame(columns=['image_id','label'])\nfor image_name in os.listdir(PARENT_DIR + '/test_images'):\n    image_path = os.path.join(PARENT_DIR + '/test_images', image_name)\n    image = tf.keras.preprocessing.image.load_img(image_path)\n    resized_image = image.resize((IMAGE_SIZE, IMAGE_SIZE))\n    numpied_image = np.expand_dims(resized_image, 0)\n    tensored_image = tf.cast(numpied_image, tf.float32)\n    submission = submission.append(pd.DataFrame({'image_id': image_name,\n                                                 'label': model.predict_classes(tensored_image)}))\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Saving CSV to output folder\nsubmission.to_csv('submission.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}