{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceType":"competition","sourceId":13836,"databundleVersionId":1718836}],"dockerImageVersionId":31329,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Cassava Leaf Disease Image Classification-","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://miro.medium.com/v2/resize:fit:1100/format:webp/1*JMMAN8_qmtIUakWEg-ouzw.jpeg\" width=\"400\">","metadata":{}},{"cell_type":"markdown","source":"## Project Aim","metadata":{}},{"cell_type":"markdown","source":"This study aims to automate the diagnosis of viral and bacterial pathogens in Manihot esculenta (Cassava) plants, a staple crop for global food security, by leveraging Convolutional Neural Network (CNN) architectures. The primary objective is to develop a high-generalization classification model capable of accurately distinguishing critical diseases such as Cassava Mosaic Disease (CMD) and Cassava Brown Streak Disease (CBSD) from low-resolution, noisy, and field-acquired images. By employing deep learning-based feature extraction techniques, the project intends to establish a robust digital decision-support mechanism to mitigate crop yield losses and enhance agricultural sustainability.","metadata":{}},{"cell_type":"markdown","source":"## Reading Dataset","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport warnings\nwarnings.filterwarnings('ignore')\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom PIL import Image\nimport cv2\n\nfrom sklearn.model_selection import train_test_split\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nimport tensorflow as tf\nfrom tensorflow.keras.models import Sequential, Model\nfrom tensorflow.keras.layers import (Dense, Flatten, Dropout, BatchNormalization, \n                                     GlobalAveragePooling2D, Input)\nfrom tensorflow.keras import layers, models, optimizers\nfrom tensorflow.keras.applications import EfficientNetB0\nfrom tensorflow.keras.callbacks import ModelCheckpoint, ReduceLROnPlateau, EarlyStopping\nfrom sklearn.metrics import classification_report, confusion_matrix, accuracy_score\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T15:50:39.606815Z","iopub.execute_input":"2026-04-07T15:50:39.607623Z","iopub.status.idle":"2026-04-07T15:50:39.613748Z","shell.execute_reply.started":"2026-04-07T15:50:39.607589Z","shell.execute_reply":"2026-04-07T15:50:39.612795Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Reading Dataset","metadata":{}},{"cell_type":"code","source":"path = '/kaggle/input/competitions/cassava-leaf-disease-classification/'\ndf = pd.read_csv(path + 'train.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T15:50:41.726477Z","iopub.execute_input":"2026-04-07T15:50:41.727209Z","iopub.status.idle":"2026-04-07T15:50:41.748281Z","shell.execute_reply.started":"2026-04-07T15:50:41.727176Z","shell.execute_reply":"2026-04-07T15:50:41.747371Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df['image_id'] = path + 'train_images/' + df['image_id']\ndf['label'] = df['label'].astype(str) ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T15:50:43.203811Z","iopub.execute_input":"2026-04-07T15:50:43.204572Z","iopub.status.idle":"2026-04-07T15:50:43.218385Z","shell.execute_reply.started":"2026-04-07T15:50:43.204537Z","shell.execute_reply":"2026-04-07T15:50:43.217455Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## EDA","metadata":{}},{"cell_type":"code","source":"df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T15:50:44.776771Z","iopub.execute_input":"2026-04-07T15:50:44.777681Z","iopub.status.idle":"2026-04-07T15:50:44.786015Z","shell.execute_reply.started":"2026-04-07T15:50:44.777636Z","shell.execute_reply":"2026-04-07T15:50:44.785179Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"dic = {'0': '0: CBB (Bacterial Blight)','1': '1: CBSD (Brown Streak)','2': '2: CGM (Green Mite)',\n    '3': '3: CMD (Mosaic Disease)','4': '4: Healthy'}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T15:50:46.720306Z","iopub.execute_input":"2026-04-07T15:50:46.721224Z","iopub.status.idle":"2026-04-07T15:50:46.725350Z","shell.execute_reply.started":"2026-04-07T15:50:46.721193Z","shell.execute_reply":"2026-04-07T15:50:46.724333Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(8, 4))\nsns.countplot(y=df['label'].map(dic), palette='viridis', order=dic.values())\nplt.title('Distribution of Cassava Leaf Diseases');","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T15:50:48.386559Z","iopub.execute_input":"2026-04-07T15:50:48.387300Z","iopub.status.idle":"2026-04-07T15:50:48.612440Z","shell.execute_reply.started":"2026-04-07T15:50:48.387266Z","shell.execute_reply":"2026-04-07T15:50:48.611723Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df['label'].value_counts()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T15:50:51.227188Z","iopub.execute_input":"2026-04-07T15:50:51.227503Z","iopub.status.idle":"2026-04-07T15:50:51.236710Z","shell.execute_reply.started":"2026-04-07T15:50:51.227478Z","shell.execute_reply":"2026-04-07T15:50:51.235667Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(20, 4))\nfor i, name in dic.items():\n    path = df[df['label'] == i].sample(1).image_id.values[0]\n    plt.subplot(1, 5, int(i)+1)\n    plt.imshow(plt.imread(path)) \n    plt.title(name, fontsize=9)\n    plt.axis('off')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T15:50:52.934870Z","iopub.execute_input":"2026-04-07T15:50:52.935477Z","iopub.status.idle":"2026-04-07T15:50:53.575197Z","shell.execute_reply.started":"2026-04-07T15:50:52.935448Z","shell.execute_reply":"2026-04-07T15:50:53.574240Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df, val_df = train_test_split(df, test_size=0.2, random_state=42, stratify=df['label'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T15:53:30.986087Z","iopub.execute_input":"2026-04-07T15:53:30.986823Z","iopub.status.idle":"2026-04-07T15:53:31.017081Z","shell.execute_reply.started":"2026-04-07T15:53:30.986794Z","shell.execute_reply":"2026-04-07T15:53:31.016468Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.utils import class_weight\n\nweights = class_weight.compute_class_weight(class_weight='balanced',classes=np.unique(train_df['label']),y=train_df['label'])\nclass_weights = dict(enumerate(weights))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T15:53:40.024586Z","iopub.execute_input":"2026-04-07T15:53:40.025439Z","iopub.status.idle":"2026-04-07T15:53:40.046115Z","shell.execute_reply.started":"2026-04-07T15:53:40.025371Z","shell.execute_reply":"2026-04-07T15:53:40.045368Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_datagen = ImageDataGenerator(rescale=1./255, rotation_range=30, horizontal_flip=True)\nval_datagen = ImageDataGenerator(rescale=1./255)\n\ntrain_gen = train_datagen.flow_from_dataframe(train_df, x_col='image_id', y_col='label',target_size=(224, 224), batch_size=32, class_mode='sparse')\n\nval_gen = val_datagen.flow_from_dataframe(val_df, x_col='image_id', y_col='label',target_size=(224, 224), batch_size=32, class_mode='sparse')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T15:54:17.669823Z","iopub.execute_input":"2026-04-07T15:54:17.670110Z","iopub.status.idle":"2026-04-07T15:54:50.724508Z","shell.execute_reply.started":"2026-04-07T15:54:17.670087Z","shell.execute_reply":"2026-04-07T15:54:50.723509Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"base_model = tf.keras.applications.EfficientNetB0(include_top=False, weights='imagenet',input_shape=(224, 224, 3))\n\nmodel = models.Sequential([\n    base_model,\n    layers.GlobalAveragePooling2D(), # Görseli yoğunlaştırır\n    layers.BatchNormalization(),      # Eğitimi hızlandırır ve stabilize eder\n    layers.Dropout(0.3),              # Ezberlemeyi (overfitting) önler\n    layers.Dense(256, activation='relu'),\n    layers.Dropout(0.2),\n    layers.Dense(5, activation='softmax') ])\n\nmodel.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=1e-3),loss='sparse_categorical_crossentropy',metrics=['accuracy'])\nmodel.summary() # Modelin yapısını kontrol et","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T15:56:05.894066Z","iopub.execute_input":"2026-04-07T15:56:05.894895Z","iopub.status.idle":"2026-04-07T15:56:08.235920Z","shell.execute_reply.started":"2026-04-07T15:56:05.894865Z","shell.execute_reply":"2026-04-07T15:56:08.234963Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"callbacks = [\n    tf.keras.callbacks.ModelCheckpoint(\"best_model.h5\", save_best_only=True),\n    tf.keras.callbacks.ReduceLROnPlateau(monitor='val_loss', factor=0.2, patience=2, min_lr=1e-6)]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T15:56:13.100542Z","iopub.execute_input":"2026-04-07T15:56:13.101373Z","iopub.status.idle":"2026-04-07T15:56:13.105943Z","shell.execute_reply.started":"2026-04-07T15:56:13.101339Z","shell.execute_reply":"2026-04-07T15:56:13.104888Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"history = model.fit(train_gen,validation_data=val_gen,epochs=10,class_weight=class_weights,callbacks=callbacks )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T15:56:20.850972Z","iopub.execute_input":"2026-04-07T15:56:20.851611Z","iopub.status.idle":"2026-04-07T16:47:01.562244Z","shell.execute_reply.started":"2026-04-07T15:56:20.851581Z","shell.execute_reply":"2026-04-07T16:47:01.561530Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"base_model.trainable = True \nmodel.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=1e-5),loss='sparse_categorical_crossentropy',metrics=['accuracy'])\n\nhistory_fine = model.fit(train_gen,validation_data=val_gen,epochs=10, class_weight=class_weights,callbacks=callbacks)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T17:32:42.968950Z","iopub.execute_input":"2026-04-07T17:32:42.969587Z","iopub.status.idle":"2026-04-07T18:18:03.868245Z","shell.execute_reply.started":"2026-04-07T17:32:42.969543Z","shell.execute_reply":"2026-04-07T18:18:03.867224Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"val_loss, val_acc = model.evaluate(val_gen, verbose=0)\nprint(f\"Overall Accuracy Score (Genel Başarı Skoru): %{val_acc * 100:.2f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T18:59:20.440028Z","iopub.execute_input":"2026-04-07T18:59:20.440752Z","iopub.status.idle":"2026-04-07T18:59:43.550574Z","shell.execute_reply.started":"2026-04-07T18:59:20.440715Z","shell.execute_reply":"2026-04-07T18:59:43.549495Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Model Performance Visualization","metadata":{}},{"cell_type":"code","source":"images, labels = next(val_gen) \npreds = model.predict(images, verbose=0)\n\nplt.figure(figsize=(12, 12))\nfor i in range(9):\n    plt.subplot(3, 3, i+1)\n    plt.imshow(images[i])\n    true_idx = np.argmax(labels[i]) if len(labels[i].shape) > 0 else int(labels[i])\n    pred_idx = np.argmax(preds[i])\n    color = 'green' if true_idx == pred_idx else 'red'\n    plt.title(f\": {pred_idx}\\nTrue: {true_idx}\", color=color)\n    plt.axis('off')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T18:55:59.278175Z","iopub.execute_input":"2026-04-07T18:55:59.279311Z","iopub.status.idle":"2026-04-07T18:56:00.492423Z","shell.execute_reply.started":"2026-04-07T18:55:59.279245Z","shell.execute_reply":"2026-04-07T18:56:00.490825Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.load_weights('best_model.h5')\nmodel.save('cassava_model_v1.keras')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-07T18:46:17.946601Z","iopub.execute_input":"2026-04-07T18:46:17.946917Z","iopub.status.idle":"2026-04-07T18:46:19.256942Z","shell.execute_reply.started":"2026-04-07T18:46:17.946892Z","shell.execute_reply":"2026-04-07T18:46:19.256192Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\"In this study, Transfer Learning was employed to classify Cassava leaf diseases. The initial training with frozen layers achieved a validation accuracy of 72.20%. After performing 'Fine-Tuning' by unfreezing the base model, the training accuracy climbed to 81.33%, while the validation accuracy settled at 69.72%. This discrepancy indicates a degree of 'overfitting,' where the model memorizes specific training patterns rather than generalizing. However, the 81% training performance demonstrates a high potential for feature extraction. Future improvements focusing on intensive 'Data Augmentation' and regularization could mitigate overfitting and consistently push the performance beyond the 80% threshold.\"","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Modelling","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}