{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":30886,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 0. Introduction","metadata":{}},{"cell_type":"markdown","source":"<div style = 'border : 3px solid non; background-color:#f2f2f2 ; ;padding:10px'>\n    \nRecent advancements in slide scanning technology have made it easier to conduct microscopic evaluations of **histopathologically** stained tissues and to digitize the results. As a result, digital analysis using deep learning has emerged as a promising tool for diagnosis. \n\nEvaluating the extent of cancer spread through the histopathological analysis of **sentinel axillary lymph nodes (SLNs)** is a vital part of the breast cancer staging process.\n\nThis document discusses strategies to mitigate **overfitting** in models and explores practical techniques for **optimal modeling**. It emphasizes the importance of balancing model complexity with generalization to achieve the best performance on unseen data.","metadata":{}},{"cell_type":"markdown","source":"## - Dataset Used","metadata":{}},{"cell_type":"markdown","source":"<div style = 'border : 3px solid non; background-color:#f2f2f2 ; ;padding:10px'>\n    \nThis dataset contains **several small pathology images** (histopathology images) that need classification. Each file is named with a unique image ID. The `train_labels.csv` file provides the ground truth labels for the images in the training folder. I will be predicting the labels for the images located in the `test` folder.\n\nA positive label indicates that the **central 32x32 pixel** area of a patch contains at least one pixel of tumor tissue. Note that tumor tissue in the outer region of the patch does not affect the label. The inclusion of this outer region aims to support fully convolutional models that do not use zero-padding, ensuring consistent behavior when applied to whole-slide images.\n\nFor **access to the data**, please use the following link:\nhttps://www.kaggle.com/c/histopathologic-cancer-detection/data","metadata":{}},{"cell_type":"markdown","source":"## - Which Structures?","metadata":{}},{"cell_type":"markdown","source":"<div style = 'border : 3px solid non; background-color:#f2f2f2 ; ;padding:10px'>\n    \nIn this project; First, I will utilize three types of Convolutional Neural Network (CNN) models: **AlexNet**, **VGGNet**, and **ResNet**. These models are some of the most popular CNN architectures used for image analysis tasks. Each CNN structure has unique attributes and advantages, and we will evaluate which models are best suited for our images by comparing their results.\n\nAdditionally, I will use the **EfficientNet** models, which are popular for their various versions ranging from B0 to B7. Depending on factors such as the number of inputs, image sizes, model layers, and computational power, one can select the simplest model, EfficientNetB0, or move up to the more complex EfficientNetB7. Given the small image size in this dataset, I prefer to use **EfficientNetB0**, **EfficientNetB1**, and **EfficientNetB2**. Ultimately, I will compare their results to determine the best version for **histopathology images**!\n\nIt’s important to note that the actual size of the images is **96x96 pixels**; however, to avoid overfitting, you could resize the **input images to 128x128 pixels** but this will cause an increase in training time and likely **time-out error** in kaggle run.","metadata":{}},{"cell_type":"markdown","source":"> ### ***NOTE*:**\n>\n> I adjusted the code to run on images, yielding the best results with **modifying each version**. First, I addressed the **overfitting** issue. Because accuracy decreased when overfitting occurred, I then focused on **optimizing accuracy**. In the end, to resolve the **time-out problem**, I employed the **GPU T4X2**, which reduces training time through parallel processing on two GPUs.\n> \n> Finally, in **the last two versions**, to address the time-out error in Kaggle runs and reduce training time, I used GPU T4X2 and adjusted the total number of epochs from twenty to **two runs of ten epochs each**. Additionally, I applied functions to improve code readability.\n\n> *I will explain these modifications in the introduction following sections. Moreover, I have provided detailed explanations throughout the code to enhance clarity.*\n","metadata":{}},{"cell_type":"markdown","source":"> ### ***Please:***\n>\n>   *Feel free to **share your thoughts in the *discussion*** section—let’s challenge this run together!*","metadata":{}},{"cell_type":"markdown","source":"## - First Modification: Overcoming \"Overfitting\" in Models","metadata":{}},{"cell_type":"markdown","source":"<div style = 'border : 3px solid non; background-color:#f2f2f2 ; ;padding:10px'>\nIn the first versions, the model tended to overfit due to its general hypotheses. To address this issue, I implemented the following solutions:\n\n- Applied **Dropout** to the model and increased the dropout rate from 0.5 to 0.6.\n- Added **L2 Regularization** to both the Conv2D and Dense layers.\n- Reduced the **learning rate** to 0.0001 (while I could use a non-constant rate, it would incur a higher computational cost).\n- Utilized **data augmentation** (rotation, shift, zoom, horizontal flip) to diversify the images.\n- **Frozen the initial layers** of the EfficientNet architecture.","metadata":{}},{"cell_type":"markdown","source":"## - Second Modification: Optimizing Models to \"Enhance Accuracy\"","metadata":{}},{"cell_type":"markdown","source":"<div style = 'border : 3px solid non; background-color:#f2f2f2 ; ;padding:10px'>\n    \n**After addressing overfitting**, the learning of the training dataset decreases slightly due to data augmentation and nearest result learning to actual outcomes. I implemented these changes in the next versions to **enhance accuracy**:\n- Enable **Early Stopping** and **increase the number of epochs**. **Early Stopping** will halt training when there is no improvement in performance, helping to prevent overfitting and reducing overall training time.\n\n- Enable **ReduceLROnPlateau** to optimize the learning rate. This technique dynamically **lowers the learning rate**, providing the model with more opportunities for effective learning. I started with a **learning rate of 0.001**, then modified it from the previous rate of 0.0001.\n\n- The **Dropout** value has been reduced to **0.5** for the first three models and to **0.4** for the EfficientNet models, as the EfficientNet architecture typically requires less Dropout.\n\n- The **L2 regularization** value has been **lowered to 0.0001** to maintain weight stability.\n\n- The intensity of **data augmentation** has been balanced with the following settings: **Rotation set** to 15 degrees, **Shifts** set to 0.1, and **Zoom** set to 0.1.","metadata":{}},{"cell_type":"markdown","source":"## - Third Modification: GPU Used; GPU T4X2 ","metadata":{}},{"cell_type":"markdown","source":"<div style = 'border : 3px solid non; background-color:#f2f2f2 ; ;padding:10px'>\n\nRunning the code in Kaggle is likely to result in a **timeout error**, as it may take more than 12 hours to complete. To address this issue, I chose to use the **`GPU T4x2`** setup, utilizing **two GPUs**. It's important to remember that we need to implement a **MirroredStrategy** for compiling the models. \n\nTo effectively distribute the data, **strategy.experimental_distribute_dataset()** was utilized. This enables a more efficient **distribution of data across GPUs**.\n\nThese solutions significantly **shorten the training duration** and solve the likely **timeout error**. ","metadata":{}},{"cell_type":"markdown","source":">\n>***Size of input images** and **BatchSize** :*\n>\n>The input images are considered to have an **actual size of 96x96 pixels**. I will test the **BatchSize values of 16, 32, and 64**.","metadata":{}},{"cell_type":"markdown","source":"## - Fourth Modification: Split the Training into Two Runs","metadata":{}},{"cell_type":"markdown","source":"<div style = 'border : 3px solid non; background-color:#f2f2f2 ; ;padding:10px'>\n    \nThe main challenge remains **reducing training time** to fit within Kaggle's execution limits. Using two GPUs alone was not enough to solve this issue.\n\n*I suggest the following approach:*\n- Use `ModelCheckpoint` to **automatically save** the best model at each epoch.\n- Split the training process into **two runs**, ensuring results are saved in **history** and reloaded when available using custom save and load functions.\n \n*Two Strategies Considered:*\n1. **Splitting the models:** Since we are training six models, the first run will train AlexNet, VGGNet, and ResNet, while the second run will train EfficientNetB0, EfficientNetB1, and EfficientNetB2.\n2. **Splitting the epochs (Recommended):** Train all models for 10 epochs in the first run, then continue from epoch 10 to 20 in the second run.\n   ","metadata":{}},{"cell_type":"markdown","source":"> ***Note:***\n> \n> These modifications are implemented in the last two versions. **Let me know your thoughts in the discussion section!**","metadata":{"execution":{"iopub.status.busy":"2025-04-20T21:30:45.143235Z","iopub.execute_input":"2025-04-20T21:30:45.143556Z","iopub.status.idle":"2025-04-20T21:30:45.148987Z","shell.execute_reply.started":"2025-04-20T21:30:45.143527Z","shell.execute_reply":"2025-04-20T21:30:45.147509Z"}}},{"cell_type":"markdown","source":"## - Final Adjustments","metadata":{}},{"cell_type":"markdown","source":"<div style = 'border : 3px solid non; background-color:#f2f2f2 ; ;padding:10px'>\n\n*After adjusting the settings to reduce training time and optimize results, the final configurations are as follows:*\n\n- **Batch Size: 64** (for two T4 GPUs) [Batch size of 64 produced the best results.]\n- Disabling **Mixed Precision**. It caused **NaN** in the loss function.\n- Implement `ReduceLROnPlateau` with a **factor of 0.5**, a **patience of 2**, and a **minimum learning rate of 1e-5**.\n- Use `Early Stopping` with a **patience of 3**.\n- Set the initial **Learning Rate to 0.001**.\n- Set **Dropout between 0.4 and 0.6**.","metadata":{}},{"cell_type":"markdown","source":"### - Always keep the following points:","metadata":{}},{"cell_type":"markdown","source":"<div style = 'border : 3px solid non; background-color:#f2f2f2 ; ;padding:10px'>\n\n- The number of images in all three datasets- `training`, `validation`, and `testing`- must be **divisible by the batch size**. Otherwise, the last incomplete batch may be divided among GPUs, resulting in errors.\n- It is advisable to **specify `steps`** when evaluating or predicting on the dataset. This is important because when we use `tf.data.Dataset`, TensorFlow may not know when to stop, especially if repeat() is used, or if the dataset is endless.\n- Ensure the **existence of the model** before `loading` it, as verifying file existence helps maintain control and stability during the load, evaluation, and prediction steps.\n- Utilize **`drop_remainder=True`** in `tf.data.Dataset.from_generator` when employing **MultiGPU** (MirroredStrategy) to drop the incomplete batch. Otherwise, the second GPU may not receive data in the last batch.\n- Using `prefetch` and **AUTOTUNE** can enhance the model's speed, particularly in the parallel loading of data. This is crucial because GPUs should not wait for samples to be fed into the data pipeline, ultimately increasing the overall speed of training and prediction.\n- A common and significant issue with **MirroredStrategy** is the uneven distribution of data across two GPUs, particularly when running different models (for instance, six models were used for comparison in this case). To address this problem, follow these steps:\n\n1. Call `K.clear_session()`, `gc.collect()`, and `reset_default_graph()` before defining the strategy.\n2. The `train_model()` function is invoked multiple times during the code execution, but the previous model remains loaded on the GPU. To resolve this, *define MirroredStrategy only once*, outside the `train_model` function, and use it within the function.","metadata":{}},{"cell_type":"markdown","source":"## **Top 10 Benefits of This Notebook:**","metadata":{}},{"cell_type":"markdown","source":"<div style = 'border : 3px solid non; background-color:#f2f2f2 ; ;padding:10px'>\n\n1. Proper data partitioning for **batch sizes** (`Batch Size = 64`) across `training`, `validation`, and `test` datasets.\n2. Conversion of data into **DataFrame** format and transformation of *labels* into **strings** to ensure compatibility with `flow_from_dataframe`. This includes appropriate usage of `flow_from_dataframe` with `class_mode='binary'` and proper implementation of shuffling.\n3. Effective utilization of `ImageDataGenerator` for **data augmentation** and **normalization**.\n4. Implementation of `MirroredStrategy` to enhance **GPUs** utilization.\n5. Conversion of the generator to `tf.data.Dataset` with *unbatching* and *batching* while respecting `drop_remainder=True` to **avoid incomplete batches** from being distributed between two GPUs.\n6. Use of **callbacks** such as `ReduceLROnPlateau`, `EarlyStopping`, and `ModelCheckpoint`.\n7. Conversion of the generator to `tf.data.Dataset` to leverage features like `prefetch(tf.data.AUTOTUNE)` for **optimized processing**.\n8. Precise definition of `steps_per_epoch`, `validation_steps`, and `test_steps` based on the *length of the generator*.\n9. **Saving** and **loading** the training history in `.npz` format, along with saving results for continuous model execution.\n10. Saving **outputs** and **predictions** in a *CSV file* with columns for ID and label.","metadata":{}},{"cell_type":"markdown","source":"# 1. Import Library","metadata":{}},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings(\"ignore\")\nfrom numpy import asarray\nimport numpy as np\nimport pandas as pd\nfrom PIL import Image\nimport cv2\nimport glob\nimport os \nimport random\nimport subprocess\n\nimport matplotlib.pyplot as plt\nfrom PIL import Image, ImageDraw\nfrom skimage.io import imread\nfrom matplotlib.patches import Rectangle\n\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers, models, Input, Model\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.losses import BinaryCrossentropy\nfrom tensorflow.keras.losses import CategoricalCrossentropy\nfrom tensorflow.keras.regularizers import l2\nfrom tensorflow.keras.callbacks import ReduceLROnPlateau, EarlyStopping, ModelCheckpoint\nfrom tensorflow.keras.applications.vgg16 import VGG16\nfrom tensorflow.keras.applications import EfficientNetB0, EfficientNetB1, EfficientNetB2\n\nimport gc\nfrom keras import backend as K # to clear the previous session and memory\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## -  Enable Mixed Precision Processing","metadata":{}},{"cell_type":"markdown","source":"Enable `mixed precision processing` to **optimize GPU usage** and reduce memory consumption.","metadata":{}},{"cell_type":"markdown","source":"> ***Note:***\n>\n> This may produce **NaN** or **Inf** in some numeric operations, such as `Loss` function.","metadata":{}},{"cell_type":"code","source":"# tf.keras.mixed_precision.set_global_policy('mixed_float16')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - Test GPUs","metadata":{"execution":{"iopub.status.busy":"2025-03-23T10:38:00.498039Z","iopub.execute_input":"2025-03-23T10:38:00.498341Z","iopub.status.idle":"2025-03-23T10:38:00.501907Z","shell.execute_reply.started":"2025-03-23T10:38:00.498314Z","shell.execute_reply":"2025-03-23T10:38:00.501152Z"}}},{"cell_type":"code","source":"# GPU or CPU is using\n\ngpu_devices = tf.config.list_physical_devices('GPU')\nif len(gpu_devices) > 1:\n    print('Two GPUs are using') \nelif len(gpu_devices) == 1:\n    print('GPU is using') \nelse: \n    print('CPU is using')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 2. Import Dataset","metadata":{}},{"cell_type":"code","source":"train_path = '/kaggle/input/histopathologic-cancer-detection/train'\ntrain_label_path = '/kaggle/input/histopathologic-cancer-detection/train_labels.csv'\ntest_path = '/kaggle/input/histopathologic-cancer-detection/test.csv'\nsample_path = '/kaggle/input/histopathologic-cancer-detection/sample_submission.csv'","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df = pd.read_csv(train_label_path)\nprint(df.head().to_markdown())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.isnull().sum()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_sample = pd.read_csv(sample_path)\nprint(df_sample.head().to_markdown())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_sample.isnull().sum()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 3. Explore the Dataset","metadata":{}},{"cell_type":"markdown","source":"## - Basic Information","metadata":{}},{"cell_type":"code","source":"df.shape","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.info()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df['label'].value_counts()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('Number of image : ', len(df))\nprint('Ratio labels : ', sum(df['label'].values)/len(df))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Malignant (positive) and Normal (negative) Images\n\npositive = df.loc[df['label']==1]['id'].values    # the ids of positive (malignant) cases\nnegative = df.loc[df['label']==0]['id'].values      \nlabel_percent = df.value_counts(normalize=True)\n\nprint('Malignant (positive):')\nprint ('    Number of images:', len(positive),'=',round(len(positive)/len(df),2),'of all images')\nprint(positive[0:3],'\\n')\n      \nprint('Normal (negative):')\nprint ('    Number of images:', len(negative),'=',round(len(negative)/len(df),2),'of all images')\nprint(negative[0:3],'\\n')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - Plot of Images","metadata":{}},{"cell_type":"code","source":"def plot_fig(ids,title,nrows=5,ncols=15):\n\n    fig,ax = plt.subplots(nrows,ncols,figsize=(18,6))\n    plt.subplots_adjust(wspace=0, hspace=0) \n    for i,j in enumerate(ids[:nrows*ncols]):\n        fname = os.path.join(train_path ,j +'.tif')\n        #fname = os.path.join(train_path ,j)\n        img = Image.open(fname)\n        idcol = ImageDraw.Draw(img)\n        idcol.rectangle(((0,0),(95,95)),outline='white')\n        plt.subplot(nrows, ncols, i+1) \n        plt.imshow(np.array(img))\n        plt.axis('off')\n\n    plt.suptitle(title, y=0.94)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_fig(positive,'Malignant (positive)')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_fig(negative,'Normal (negative):')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - Make Lists of images","metadata":{}},{"cell_type":"code","source":"# Input Images into the List\n\ntrain_file_path = '/kaggle/input/histopathologic-cancer-detection/train'\ntest_file_path = '/kaggle/input/histopathologic-cancer-detection/test'\nlabels_csv_path = '/kaggle/input/histopathologic-cancer-detection/train_labels.csv'# labels in CSV file\n\n# Read the CSV file include labels\ndf_labels = pd.read_csv(labels_csv_path)\n\n# Convert the dataframe to a dictionary for faster access to label\nlabel_dict = dict(zip(df_labels['id'], df_labels['label']))\n\n# Two-dimensional list for train\nimage_list_train = []\nfor train in os.listdir(train_file_path): \n    if train.endswith(\".tif\"): \n        file_path = os.path.join(train_file_path, train)\n        image_id = train.split(\".\")[0]  # Remove the .tif extension to find the label\n        label = label_dict.get(image_id, \"Unknown\")  # Get the label from the dictionary (Unknown if doesn't exist)\n        image_list_train.append([file_path, label])\n\n# Two-dimensional list for test (No labels)\nimage_list_test = []\nfor test in os.listdir(test_file_path): \n    if test.endswith(\".tif\"): \n        file_path = os.path.join(test_file_path, test)\n        # label = \"Test\"  # Replace the label\n        # image_list_test.append([file_path, label])\n        image_list_test.append([file_path])\n\n\nprint(image_list_train[:5])\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# The Number of Images\ntrain_len = len(image_list_train)\nprint (\"The Number of Histopathologic images in the train dataset: \", train_len)\ntest_len = len(image_list_test)\nprint (\"The Number of Histopathologic images in the test dataset: \", test_len)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - Calculate the Ratio of Images","metadata":{}},{"cell_type":"code","source":"# Show Some Image\nfor i in range(5):\n    img = plt.imread('/kaggle/input/histopathologic-cancer-detection/train/'+df.iloc[i]['id']+'.tif')\n    print(df.iloc[i]['label'])\n    plt.imshow(img)\n    plt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"random_img = random.choice(df['id'])\nrandom_img","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Calculate the Ratio of Images\n\n# img = '/kaggle/input/histopathologic-cancer-detection/train/'+ df.iloc[0]['id']+'.tif'\nimg = '/kaggle/input/histopathologic-cancer-detection/train/'+ random_img +'.tif'\n\nimage= cv2.imread(img)\nheight, width= image.shape[:2]\nprint(\"The height is \", height)\nprint(\"The width is \", width)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 4. Pre Modeling","metadata":{}},{"cell_type":"markdown","source":"## - Define Hyperparameters","metadata":{"execution":{"iopub.status.busy":"2025-02-11T16:35:47.594707Z","iopub.execute_input":"2025-02-11T16:35:47.595502Z","iopub.status.idle":"2025-02-11T16:35:47.601057Z","shell.execute_reply.started":"2025-02-11T16:35:47.595433Z","shell.execute_reply":"2025-02-11T16:35:47.599644Z"}}},{"cell_type":"markdown","source":"> ***Note:***\n> - I regarded **height and width** as a **96*96 format**, since overfitting alters it, you can change it to 128 * 128. But it may occur the **time-out error in Kaggle**.\n>\n> - Additionally, since I used the **GPU T4x2**, which has less RAM than the TP100, and applied **ReduceLR** to dynamically adjust the learning rate to prevent fluctuations, I will test optimal batch size setting the **BatchSize values of 8, 16, and 32**., and finally change the **Batch size to 32**. [In GPU P100 the BatchSize = 64]","metadata":{}},{"cell_type":"code","source":"Img_height = 96\nImg_width = 96\nBatch_size = 64 # Since overfitting occurs, and use GPU T4x2 and ReduceLR to adjust the dynamic learning rate.\n\nepochs=10","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - Train_test Split Data","metadata":{}},{"cell_type":"markdown","source":"> Split train dataset to: **1.Train = 85%**,  **2.Validation = 15%**","metadata":{}},{"cell_type":"code","source":"LenTrain = 0.85\nLenValid = 0.15","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":">***Note:***\n>\n> In ***training*** and ***validation***, the amount of data must be **divisible by the batch size** (i.e., 32). This is crucial because, during model fitting, each batch needs to have exactly the expected number of samples; otherwise, TensorFlow may produce an error in the last batch.\nHowever, this requirement **does not apply** during ***testing***, as the evaluate() and predict() functions do not necessitate batches of uniform size and can handle an incomplete final batch without any issues.","metadata":{}},{"cell_type":"code","source":"train_size = (int(LenTrain * len(image_list_train)) // Batch_size) * Batch_size\nvalid_size = (int(LenValid * len(image_list_train)) // Batch_size) * Batch_size\ntest_size = (int(len(image_list_test) // Batch_size) * Batch_size) # Because I'll use \"drop_remainder=True\"!\nprint(\"Split images in train to: \", \"Train =\", train_size , \" Valid =\", valid_size)\nprint(\"Also images in test: \", \"Test =\", test_size)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_train = image_list_train[:train_size]\ndata_valid = image_list_train[train_size:train_size + valid_size]\ndata_test = image_list_test[:test_size]","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - Image Augmentation and Rescale","metadata":{}},{"cell_type":"markdown","source":"> I perform **image augmentation** on the training dataset (rotation, shift, zoom, horizontal flip); however, since the histopathology images are small and it's crucial to preserve information, these modifications are minimal.\n> Image augmentations enhance accuracy in the training dataset and improve the validation dataset.","metadata":{}},{"cell_type":"code","source":"# Image Augmentation and Rescale the \"Train Dataset\"\ntrainGenerator = ImageDataGenerator(\n    rescale=1./255.,\n    rotation_range=15,       # Low rotation (15 degrees)\n    width_shift_range=0.1,   # shift (up to 10% of the image)\n    height_shift_range=0.1,  \n    zoom_range=0.1,          # Low zoom (max 10% resizing)\n    horizontal_flip=True,    # Horizontal mirroring\n    fill_mode='nearest'      # How to do new pixels\n)\n\n# Only Rescale the \"Validation Dataset\" and \"Test Dataset\"\nvalGenerator = ImageDataGenerator(rescale=1./255.)\ntestGenerator = ImageDataGenerator(rescale=1./255.)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - Create Dataframes for Train, Validation and Test","metadata":{}},{"cell_type":"code","source":"data_train[:10]","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_valid[:10]","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"data_test[:10]","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create dataframes for train\ndf_train = pd.DataFrame(data_train,columns = ['id', 'label'])\n\n# Create dataframes for validation\ndf_valid = pd.DataFrame(data_valid,columns = ['id', 'label'])\n\n# Create dataframes for test\ndf_test = pd.DataFrame(data_test,columns = ['id'])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(type(df_valid))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_valid.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train.shape","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_valid.shape","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_test.shape","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_test.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train.info()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### - Convert numeric values of label column to string","metadata":{}},{"cell_type":"code","source":"# Convert numeric values of label column to string\ndf_train['label'] = df_train['label'].astype(str)\ndf_valid['label'] = df_valid['label'].astype(str)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_train.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df_valid.info()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - Input Model","metadata":{}},{"cell_type":"markdown","source":">***Note:***\n>\n> Since the model has a scalar (0 or 1) and an output node with a **sigmoid** function, the **class_mode** is set to **'binary'** in **train** and **validation** dataset.","metadata":{}},{"cell_type":"code","source":"trainDataset = trainGenerator.flow_from_dataframe(\n  dataframe = df_train,\n  class_mode= \"binary\",  # class_mode set to 'binary'\n  x_col = \"id\",\n  y_col = \"label\",\n  batch_size = Batch_size,\n  seed = 42,\n  shuffle = True,\n  target_size = (Img_height,Img_width), #set the height and width of the images\n  #drop_remainder=True  # The flow_from_dataframe or flow_from_directory functions do not support this parameter.\n)\n\nvalDataset = valGenerator.flow_from_dataframe(\n  dataframe = df_valid,\n  class_mode = \"binary\", # class_mode set to 'binary'\n  x_col = \"id\",\n  y_col = \"label\",\n  batch_size = Batch_size,\n  seed = 42,\n  shuffle = True,\n  target_size = (Img_height,Img_width),\n  #drop_remainder=True\n)\n\ntestDataset = testGenerator.flow_from_dataframe(\n  dataframe = df_test,\n  class_mode = None, # Because it doesn't have any label\n  x_col = \"id\",\n  y_col = None, # Because it doesn't have any label\n  batch_size = Batch_size,\n  seed = 42,\n  shuffle = False, # No need to shuffle test data\n  target_size = (Img_height,Img_width),\n)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - Data Distribution Across GPUs","metadata":{}},{"cell_type":"markdown","source":"To effectively distribute the data, **strategy.experimental_distribute_dataset()** was utilized. This enables a more efficient **distribution of data across GPUs**.","metadata":{}},{"cell_type":"markdown","source":"> When using the **GPU T4x2** in Kaggle with **MirroredStrategy** enabled to utilize the model on two GPUs, it is necessary to use `tf.data.Dataset.from_generator` for **trainDataset** and **valDataset**, as flow_from_dataframe cannot be directly used with MirroredStrategy. However, this is not required for test data, where prediction and evaluation are conducted.","metadata":{}},{"cell_type":"code","source":"# A function that converts the Generator to a tf.data.Dataset [Usable for training, validation, and testing datasets]\ndef dataset_from_generator(generator, has_labels=True):\n    if has_labels:\n        dataset = tf.data.Dataset.from_generator(\n            lambda: generator,\n            output_signature=(\n                tf.TensorSpec(shape=(None, Img_height, Img_width, 3), dtype=tf.float32),\n                tf.TensorSpec(shape=(None,), dtype=tf.float32)\n            )\n        )\n    else:\n        dataset = tf.data.Dataset.from_generator(\n            lambda: generator,\n            output_signature=tf.TensorSpec(shape=(None, Img_height, Img_width, 3), dtype=tf.float32)\n        )\n    # Since we are using multiple GPUs, never forget 'drop_remainder=True'! Otherwise the second GPU may not receive data in the last batch.\n    # Safely unbatch and re-batch\n    return dataset.unbatch().batch(Batch_size, drop_remainder=True)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 5) Define Functions: Save and Load, Train Model, and Visualization Result","metadata":{}},{"cell_type":"markdown","source":"## - Define Save and Load Functions","metadata":{}},{"cell_type":"markdown","source":"We need to define the `save` and `load` **functions** to preserve the model and its results for use in subsequent runs and to retrieve them.","metadata":{}},{"cell_type":"code","source":"# Load\ndef load_training_history(history_file):\n    if os.path.exists(history_file):\n        try:\n            data = np.load(history_file)\n            print(\">>> Load Previous history ...\")\n            return (list(data['acc']) if 'acc' in data else [], \n                    list(data['val_acc']) if 'val_acc' in data else [],\n                    list(data['loss']) if 'loss' in data else [],\n                    list(data['val_loss']) if 'val_loss' in data else [])\n        except Exception as e:\n            print(f\"⚠️ Error loading history: {e}\")\n            return [], [], [], []\n    else:\n        print(\"*** First run: Model training starts from scratch ... \")\n        return [], [], [], []\n# Save\ndef save_training_history(history_file, full_acc, full_val_acc, full_loss, full_val_loss):\n    try:\n        np.savez(history_file, acc=full_acc, val_acc=full_val_acc, loss=full_loss, val_loss=full_val_loss)\n        print(\"Training history saved!\")\n    except Exception as e:\n        print(f\"⚠️ Error save history: {e}\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - Define Train Model Function","metadata":{}},{"cell_type":"markdown","source":"In the **train_model** function, I utilize `reduce_lr`, `early_stopping`, and `checkpoint` in **callbacks**, and also `save` and `load` the history of the training model (epochs). \nThe function creates a new model or loads the model (if it has already been trained and the results have been saved) and continues training. \n\nThe Function parameters:\n\n- model_name: Model file name to save and load\n- model_fn: Function that creates the new model\n- trainDataset: Training dataset\n- valDataset: Validation dataset\n- epochs: Number of new epochs to train\n\n> ***Note:* Keras no longer supports the \".h5\" format for ModelCheckpoint!**\n> \n> In the **latest versions of TensorFlow/Keras**, the `ModelCheckpoint` function only saves models in the **\".keras\"** format.","metadata":{}},{"cell_type":"code","source":"# Define MirroredStrategy only once, outside the `train_model` function\n\ngpus = tf.config.experimental.list_physical_devices('GPU')\nif gpus:\n    for gpu in gpus:\n        tf.config.experimental.set_memory_growth(gpu, True)\n    print(f\"✅ GPUs detected: {[gpu.name for gpu in gpus]}\")\n    strategy = tf.distribute.MirroredStrategy()\nelse:\n    print(\"⚠ No GPU detected! Running on CPU.\")\n    strategy = tf.distribute.OneDeviceStrategy(device=\"/CPU:0\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def train_model(model_name, model_fn, trainDataset, valDataset, testDataset, epochs):\n    \n    # 0) Clear previous session to avoid shape mismatch or residual graph issues\n    K.clear_session()\n    gc.collect()\n    tf.compat.v1.reset_default_graph()\n    \n    # 1) History management\n    history_file = f\"{model_name}_history.npz\"\n    prev_acc, prev_val_acc, prev_loss, prev_val_loss = load_training_history(history_file)\n\n    # 2) Using MirroredStrategy: Build or load model inside strategy scope\n        # - Strategy already defined outside; use it here\n    with strategy.scope():\n        if os.path.exists(model_name):\n            print(f\">>> Loading the model from '{model_name}' to continue training...\")\n            model = tf.keras.models.load_model(model_name)  \n        else:\n            print(f\"*** Building a new model and starting training ...\")\n            model = model_fn()  \n            model.compile(loss=BinaryCrossentropy(),\n                          optimizer=Adam(learning_rate=0.001), # The Best Result [Change from learning_rate=0.002 and 0.0005]\n                          metrics=['accuracy'])        \n\n    if not os.path.exists(model_name):\n        model.summary()\n\n    print(f\"[INFO] Dataset sizes -> train: {len(trainDataset)}, valid: {len(valDataset)}, test: {len(testDataset)}\")\n    assert len(trainDataset) % strategy.num_replicas_in_sync == 0, \"Train dataset not divisible by number of GPUs\"\n\n    # 3) Convert DataGenerator to tf.data.Dataset and preprocess\n        # - Optimizing processing and preventing slowdowns can be achieved by \"adding prefetch\".\n    trainDataset_tf = dataset_from_generator(trainDataset, has_labels=True).prefetch(tf.data.AUTOTUNE)\n    valDataset_tf   = dataset_from_generator(valDataset, has_labels=True).prefetch(tf.data.AUTOTUNE)\n    testDataset_tf  = dataset_from_generator(testDataset, has_labels=False).prefetch(tf.data.AUTOTUNE)\n\n\n    # 4) Set Callbacks\n    reduce_lr = ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=2, min_lr=1e-5, verbose=1)\n    early_stopping = EarlyStopping(monitor='val_loss', patience=3, restore_best_weights=True, verbose=1)\n    checkpoint = ModelCheckpoint(model_name, monitor='val_loss', save_best_only=True, mode='min', verbose=1)\n    callbacks = [reduce_lr, early_stopping, checkpoint]\n\n    # 5) Check batch shape of input data before training\n    \n    for sample_batch in trainDataset_tf.take(1):\n        if isinstance(sample_batch, tuple):\n            print(\"Batch shape (X):\", sample_batch[0].shape)\n            print(\"Batch shape (Y):\", sample_batch[1].shape)\n        else:\n            print(\"Batch shape:\", sample_batch.shape)\n\n\n    # 6) Model training\n        # - We must define steps to limit the number of execution steps; otherwise, execution may continue indefinitely.\n        # - Note: Since flow_from_dataframe has been used before, its output works in batches, and the length (len(generator)) is the number of batches. \n          # [Therefore, there is no need for \"// Batch_size\" on trainDataset or valDataset unless we want to consider a smaller amount of data!]\n    steps_per_epoch = len(trainDataset) ##// Batch_size # [Batch_size = 32]\n    print(\"steps_per_epoch =\", steps_per_epoch )\n    validation_steps = len(valDataset) ##// Batch_size\n    test_steps = df_test.shape[0] // Batch_size  # Only the full batches will be used  \n    valid_test_count = test_steps * Batch_size\n    \n    history = model.fit(\n        trainDataset_tf,\n        epochs=epochs,\n        validation_data=valDataset_tf,\n        callbacks=callbacks,\n        steps_per_epoch=steps_per_epoch,\n        validation_steps=validation_steps\n    )\n\n\n    # 7) Merge previous history with new data\n    full_acc = prev_acc + history.history['accuracy']\n    full_val_acc = prev_val_acc + history.history['val_accuracy']\n    full_loss = prev_loss + history.history['loss']\n    full_val_loss = prev_val_loss + history.history['val_loss']\n    print('Merged')\n        # Save new history\n    save_training_history(history_file, full_acc, full_val_acc, full_loss, full_val_loss)\n    print('saved')\n\n    # 8) Evaluate and predict the saved model\n        # Verifying that the model has been successfully loaded [Note: If this code is not written, loading and execution may stop]\n    if os.path.exists(model_name):\n        print(f\"✅ Model '{model_name}' exists. Loading...\")\n        best_model = tf.keras.models.load_model(model_name)\n    else:\n        print(f\"❌ Model '{model_name}' NOT found!\")\n        return \n        # - Evaluating the validation dataset\n    print('Evaluated')\n    loss, acc = best_model.evaluate(valDataset, verbose=1, steps=validation_steps) ## valDataset_tf\n    print(f\"✅ Model '{model_name}' -> Loss: {loss:.4f}, Accuracy: {acc:.4f}\")\n    \n        # - Predicting the test dataset\n    print(\"🔍 Predicting on test dataset...\")\n    # testDataset_tf = testDataset_tf.cache()\n    predictions = best_model.predict(testDataset, verbose=1, steps=test_steps) ## testDataset_tf\n    binary_predictions = (predictions > 0.5).astype(int).flatten() # Convert [0,1] predictions to 0 or 1\n    print(binary_predictions[:20])\n\n    # 9) Create a DataFrame from the model's prediction results and save it to a CSV file.\n    ids = df_test['id'].values[:valid_test_count] # Extracting ids from df_test [Match with only the IDs that correspond to used test samples]\n    \n    output_df = pd.DataFrame({\n        'id': ids,\n        'label': binary_predictions\n    })\n    output_filename = f\"{model_name.replace('.keras', '')}_predictions.csv\"\n    output_df.to_csv(output_filename, index=False)\n    print(f\"📁 Predictions saved to '{output_filename}'\")\n\n    return best_model \n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - Define Visualization Functions","metadata":{}},{"cell_type":"markdown","source":"I used four functions to display the training history and comparison models. Two functions to display the history of **accuracy** and **loss** during model training, and two functions to compare the accuracy and loss of the models on validation data together.","metadata":{}},{"cell_type":"code","source":"# History loading function and visualization \"Accuracy\" plot\ndef plot_accuracy(model_name):\n    history_file = f\"{model_name}_history.npz\"\n    if not os.path.exists(history_file):\n        print(\"⚠️No history was found for this model.\")\n        return\n\n    data = np.load(history_file)\n    acc = data['acc']\n    val_acc = data['val_acc']\n\n    plt.figure(figsize=(8, 6))\n    plt.plot(acc, label='Training Accuracy')\n    plt.plot(val_acc, label='Validation Accuracy')\n    plt.xlabel('Epochs')\n    plt.ylabel('Accuracy')\n    plt.title(f'Number of epochs & Accuracy in - {model_name}')\n    plt.legend()\n    plt.grid()\n    plt.show()\n\n# History loading function and visualization \"Loss\" plot\ndef plot_loss(model_name):\n    history_file = f\"{model_name}_history.npz\"\n    if not os.path.exists(history_file):\n        print(\"⚠️No history was found for this model.\")\n        return\n\n    data = np.load(history_file)\n    loss = data['loss']\n    val_loss = data['val_loss']\n\n    plt.figure(figsize=(8, 6))\n    plt.plot(loss, label='Training Loss')\n    plt.plot(val_loss, label='Validation Loss')\n    plt.xlabel('Epochs')\n    plt.ylabel('Loss')\n    plt.title(f'Number of epochs & Loss in - {model_name}')\n    plt.legend()\n    plt.grid()\n    plt.show()\n\n# Function to compare the \"Accuracy\" of multiple models together\ndef compare_models_accuracy(models):\n    plt.figure(figsize=(10, 6))\n    \n    for model_name in models:\n        history_file = f\"{model_name}_history.npz\"\n        if os.path.exists(history_file):\n            data = np.load(history_file)\n            plt.plot(data['acc'], label=f\"{model_name} - Train\")\n            plt.plot(data['val_acc'], linestyle='dashed', label=f\"{model_name} - Val\")\n    \n    plt.xlabel('Epochs')\n    plt.ylabel('Accuracy')\n    plt.title('Comparison of Models - Accuracy')\n    plt.legend()\n    plt.grid()\n    plt.show()\n\n# Function to compare the \"Loss\" of multiple models together\ndef compare_models_loss(models):\n    plt.figure(figsize=(10, 6))\n    \n    for model_name in models:\n        history_file = f\"{model_name}_history.npz\"\n        if os.path.exists(history_file):\n            data = np.load(history_file)\n            plt.plot(data['loss'], label=f\"{model_name} - Train\")\n            plt.plot(data['val_loss'], linestyle='dashed', label=f\"{model_name} - Val\")\n    \n    plt.xlabel('Epochs')\n    plt.ylabel('Loss')\n    plt.title('Comparison of Models - Loss')\n    plt.legend()\n    plt.grid()\n    plt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 6) Modeling: AlexNet, VGGNet, and ResNet","metadata":{}},{"cell_type":"code","source":"print ('Remeber That;')\nprint ('The Size of input Images in the Modle are: ',Img_height,'*', Img_width)\n\nprint ('The Batch Size is: ', Batch_size)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"##\ngpus = tf.config.experimental.list_physical_devices('GPU')\nif gpus:\n    for gpu in gpus:\n        tf.config.experimental.set_memory_growth(gpu, True)\n    print(\"GPUs are ready:\", gpus)\nelse:\n    print(\"No GPU detected!\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - I) AlexNet","metadata":{}},{"cell_type":"markdown","source":"**AlexNet**; One of the first deep convolutional neural networks (CNNs) that popularized deep learning in computer vision. It introduced ReLU activation, dropout, and efficient GPU training.","metadata":{}},{"cell_type":"code","source":"def AlexNet():\n    inp = layers.Input((Img_height,Img_width, 3)) \n\n    x = layers.Conv2D(96, 7, 2, activation='relu', padding='same', kernel_regularizer=l2(0.0001))(inp)\n    x = layers.BatchNormalization()(x)\n    x = layers.MaxPooling2D(3, 2, padding='same')(x)\n\n    x = layers.Conv2D(256, 5, 1, activation='relu', padding='same', kernel_regularizer=l2(0.0001))(x)\n    x = layers.BatchNormalization()(x)\n    x = layers.MaxPooling2D(3, 2, padding='same')(x)\n\n    x = layers.Conv2D(384, 3, 1, activation='relu', padding='same', kernel_regularizer=l2(0.0001))(x)\n    x = layers.Conv2D(384, 3, 1, activation='relu', padding='same', kernel_regularizer=l2(0.0001))(x)\n    x = layers.Conv2D(256, 3, 1, activation='relu', padding='same', kernel_regularizer=l2(0.0001))(x)\n    x = layers.MaxPooling2D(3, 2, padding='same')(x)\n\n    x = layers.Flatten()(x)\n    x = layers.Dense(4096, activation='relu', kernel_regularizer=l2(0.0001))(x) # Since overfitting occurs (Add L2) \n    x = layers.Dropout(0.5)(x) # Since overfitting occurs\n    x = layers.Dense(4096, activation='relu', kernel_regularizer=l2(0.0001))(x)\n    x = layers.Dropout(0.5)(x) # Since overfitting occurs\n\n    x = layers.Dense(1, activation='sigmoid')(x)  # One-dimensional output for binary classification\n   \n    model_Alex = models.Model(inputs=inp, outputs=x)\n\n    return model_Alex\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ُShape of a random Input of images\nsample_img, _ = next(iter(trainDataset)) \nprint(\"Image shape from dataset:\", sample_img.shape)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Utilize train_model Function to Model (epochs=10)\n\ntrain_model(\"AlexNet.keras\", AlexNet, trainDataset, valDataset, testDataset, epochs)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### - Visualization Output","metadata":{}},{"cell_type":"code","source":"# Visualize Accuracy history\nplot_accuracy(\"AlexNet.keras\")\n\n# Visualize loss history\nplot_loss(\"AlexNet.keras\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - II) VGGNet","metadata":{}},{"cell_type":"markdown","source":"**VGGNet**; Known for its deep but simple architecture with small 3×3 convolutional filters. It is widely used for feature extraction and transfer learning.","metadata":{}},{"cell_type":"markdown","source":"> In the ordinary VGG model, for 224x224 inputs, five MaxPoolings are considered. But **four** times is enough for our small images (96x96). Also, to avoid possible errors in the \"Flatten layer,\" I **reduced the number of Dense filters** (1024 or 2048).","metadata":{}},{"cell_type":"code","source":"#VGGNet\n\ndef VGGNet():\n    inp = layers.Input((Img_height,Img_width, 3))\n    \n    x = layers.Conv2D(64, 3, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(inp) # Since overfitting occurs (Add L2)\n    x = layers.Conv2D(64, 3, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(x)\n    x = layers.BatchNormalization()(x)\n    x = layers.MaxPooling2D(2, 2)(x)  # (48, 48, 64)\n\n    x = layers.Conv2D(128, 3, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(x)\n    x = layers.Conv2D(128, 3, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(x)\n    x = layers.BatchNormalization()(x)\n    x = layers.MaxPooling2D(2, 2)(x)  # (24, 24, 128)\n\n    x = layers.Conv2D(256, 3, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(x)\n    x = layers.Conv2D(256, 3, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(x)\n    x = layers.Conv2D(256, 3, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(x)\n    x = layers.BatchNormalization()(x)\n    x = layers.MaxPooling2D(2, 2)(x)  # (12, 12, 256)\n\n    x = layers.Conv2D(512, 3, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(x)\n    x = layers.Conv2D(512, 3, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(x)\n    x = layers.Conv2D(512, 3, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(x)\n    x = layers.BatchNormalization()(x)\n    x = layers.MaxPooling2D(2, 2)(x)  # (6, 6, 512)\n\n    \"\"\"\n    x = layers.Flatten()(x)\n    x = layers.Dense(1024, activation='relu', kernel_regularizer=l2(0.0005))(x)  # Reduce the number of nodes from 2048\n    x = layers.Dropout(0.5)(x) # Since overfitting occurs\n    x = layers.Dense(512, activation='relu', kernel_regularizer=l2(0.0005))(x)  # Reduce the number of nodes from 1024\n    x = layers.Dropout(0.5)(x) # Since overfitting occurs\n    x = layers.Dense(1, activation='sigmoid')(x)  # Binary output\n\n    model_VGG = models.Model(inputs=inp, outputs=x)\n    \"\"\"\n\n    # Classifier\n    x = layers.Flatten()(x)\n    x = layers.Dense(512, activation='relu', kernel_regularizer=l2(0.0005))(x) # Reduce the number of nodes\n    x = layers.Dropout(0.4)(x) # Since overfitting occurs\n    x = layers.Dense(1, activation='sigmoid')(x) # Binary output\n\n    model_VGG = models.Model(inputs=inp, outputs=x)\n\n    return model_VGG\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Utilize train_model Function to Model (epochs=10)\n\ntrain_model(\"VGGNet.keras\", VGGNet, trainDataset, valDataset, testDataset, epochs)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### - Visualization Output","metadata":{}},{"cell_type":"code","source":"# Visualize Accuracy history\nplot_accuracy(\"VGGNet.keras\")\n\n# Visualize Loss history\nplot_loss(\"VGGNet.keras\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - III) ResNet","metadata":{}},{"cell_type":"markdown","source":"**ResNet**; Introduced residual connections (skip connections) to solve the vanishing gradient problem, allowing the training of very deep networks efficiently.","metadata":{}},{"cell_type":"markdown","source":"For 96x96 images,at the **first layers** 3x3 kernel with stride one was used. Stride tow was also used in the **middle layers** to prevent feature dimensions from shrinking too quickly. And adding **Skip Connections** for ResNet model.","metadata":{}},{"cell_type":"code","source":"# ResNet\n\n# Skip Connection\ndef residual_block(x, filters):\n    shortcut = x \n    \n    x = layers.Conv2D(filters, 3, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(x) # Since overfitting occurs (Add L2)\n    x = layers.BatchNormalization()(x)\n    x = layers.Conv2D(filters, 3, padding='same', activation=None, kernel_regularizer=l2(0.0001))(x)\n    x = layers.BatchNormalization()(x)\n    \n    # Add Skip Connection\n    x = layers.Add()([x, shortcut])\n    x = layers.Activation('relu')(x)\n    \n    return x\n    \n\ndef ResNet34():\n    inp = layers.Input((Img_height,Img_width, 3))\n    \n    # The initial layer\n    x = layers.Conv2D(64, 3, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(inp)\n    x = layers.BatchNormalization()(x)\n    x = layers.MaxPooling2D(pool_size=2, strides=2)(x)\n    \n    # ResNet blocks\n    for _ in range(3):   # 3 blocks with 64 filters\n        x = residual_block(x, 64)\n    \n    x = layers.Conv2D(128, 3, strides=2, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(x) \n    x = layers.BatchNormalization()(x)\n    \n    for _ in range(4):\n        x = residual_block(x, 128)\n    \n    x = layers.Conv2D(256, 3, strides=2, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(x)\n    x = layers.BatchNormalization()(x)\n    \n    for _ in range(6):\n        x = residual_block(x, 256)\n    \n    x = layers.Conv2D(512, 3, strides=2, padding='same', activation='relu', kernel_regularizer=l2(0.0001))(x)\n    x = layers.BatchNormalization()(x)\n    \n    for _ in range(3):\n        x = residual_block(x, 512)\n    \n    # The final layers\n    x = layers.GlobalAveragePooling2D()(x)\n    x = layers.Dense(512, activation='relu', kernel_regularizer=l2(0.0001))(x)\n    x = layers.Dropout(0.5)(x) # Since overfitting occurs\n    x = layers.Dense(1, activation='sigmoid')(x)  \n    \n    model = models.Model(inputs=inp, outputs=x)\n    return model\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Utilize train_model Function to Model (epochs=10)\n\ntrain_model(\"ResNet34.keras\", ResNet34, trainDataset, valDataset, testDataset, epochs)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### - Visualization Output","metadata":{}},{"cell_type":"code","source":"# Visualize Accuracy history\nplot_accuracy(\"ResNet34.keras\")\n\n# Visualize Loss history\nplot_loss(\"ResNet34.keras\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - Result","metadata":{}},{"cell_type":"code","source":"# Comparison of \"AlexNet\", \"VGGNet\", \"ResNet34\" models [Accuracy & Loss]\ncompare_models_accuracy([\"AlexNet.keras\", \"VGGNet.keras\", \"ResNet34.keras\"])\ncompare_models_loss([\"AlexNet.keras\", \"VGGNet.keras\", \"ResNet34.keras\"])\ncompare_models_accuracy([\"VGGNet.keras\", \"ResNet34.keras\"])\ncompare_models_loss([\"VGGNet.keras\", \"ResNet34.keras\"])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"# 7. Modeling: EfficientNet","metadata":{}},{"cell_type":"markdown","source":"**EfficientNet**; A family of models that optimize accuracy and efficiency using compound scaling, balancing network width, depth, and resolution for better performance with fewer parameters.","metadata":{}},{"cell_type":"markdown","source":"\n> I use three versions of EfficientNet: **EfficientNetB0**, **EfficientNetB1** and **EfficientNetB2**. I believe that more powerful versions, like EfficientNetB4, may be overly complex for input dimensions of **96*96 pixels**. Also I utilized **pre-trained weights** to fine-tune only the outer layers of the new dataset. This approach helps to **reduce computational** costs and allows the model to achieve higher accuracy more quickly, even with less data. [ I change the settings by updating **\"weights=None\"** to **\"weights='imagenet'\"** and setting **\"base_model.trainable=True\"** to **\"base_model.trainable=False**.\" ]\n","metadata":{}},{"cell_type":"code","source":"# Using weights=\"imagenet\" for pre-trained weights\ndef efficientnet(model_type, input_shape=(Img_height,Img_width, 3), num_classes=1):\n    if model_type == \"B0\":\n        base_model = EfficientNetB0(weights=\"imagenet\", include_top=False, input_shape=input_shape)\n    elif model_type == \"B1\":\n        base_model = EfficientNetB1(weights=\"imagenet\", include_top=False, input_shape=input_shape)\n    elif model_type == \"B2\":\n        base_model = EfficientNetB2(weights=\"imagenet\", include_top=False, input_shape=input_shape)\n    else:\n        raise ValueError(\"Invalid model type. Choose 'B0', 'B1' or 'B2'\")\n    \n    # Freeze base model\n    base_model.trainable = False  # Set False to use pre-trained weights\n    \n    x = layers.GlobalAveragePooling2D()(base_model.output)\n    x = layers.Dense(512, activation='relu', kernel_regularizer=l2(0.0001))(x) # Since overfitting occurs (Add L2)\n    x = layers.Dropout(0.4)(x) # Since overfitting occurs\n    output = layers.Dense(num_classes, activation='sigmoid')(x)  # Binary classification\n    \n    model = models.Model(inputs=base_model.input, outputs=output)\n    \n    return model","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - I) EfficientNetB0","metadata":{}},{"cell_type":"markdown","source":"> **Note:** If we pass **efficientnet(\"B0\")** directly, the model will be built immediately, but we want it to be built only at runtime.\nI'll use \"**lambda: efficientnet(\"B0\")**\", which gives train_model just one function and executes it later.","metadata":{}},{"cell_type":"code","source":"# Utilize train_model Function to Model (epochs=10)\n\ntrain_model(\"EfficientNetB0.keras\", lambda: efficientnet(\"B0\"), trainDataset, valDataset, testDataset, epochs)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### - Visualization Output","metadata":{}},{"cell_type":"code","source":"# Visualize Accuracy history\nplot_accuracy(\"EfficientNetB0.keras\")\n\n# Visualize Loss history\nplot_loss(\"EfficientNetB0.keras\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - II) EfficientNetB1","metadata":{}},{"cell_type":"code","source":"# Utilize train_model Function to Model (epochs=10)\n\ntrain_model(\"EfficientNetB1.keras\", lambda: efficientnet(\"B1\"), trainDataset, valDataset, testDataset, epochs)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### - Visualization Output","metadata":{}},{"cell_type":"code","source":"# Visualize Accuracy history\nplot_accuracy(\"EfficientNetB1.keras\")\n\n# Visualize Loss history\nplot_loss(\"EfficientNetB1.keras\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - III) EfficientNetB2","metadata":{}},{"cell_type":"code","source":"# Utilize train_model Function to Model (epochs=10)\n\ntrain_model(\"EfficientNetB2.keras\", lambda: efficientnet(\"B2\"), trainDataset, valDataset, testDataset, epochs)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### - Visualization Output","metadata":{}},{"cell_type":"code","source":"# Visualize Accuracy history\nplot_accuracy(\"EfficientNetB2.keras\")\n\n# Visualize Loss history\nplot_loss(\"EfficientNetB2.keras\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - Result","metadata":{}},{"cell_type":"code","source":"# Comparison of \"EfficientNetB2\", \"EfficientNetB1\", \"EfficientNetB0\" models [Accuracy & Loss]\ncompare_models_accuracy([\"EfficientNetB2.keras\", \"EfficientNetB1.keras\", \"EfficientNetB0.keras\"])\ncompare_models_loss([\"EfficientNetB2.keras\", \"EfficientNetB1.keras\", \"EfficientNetB0.keras\"])\ncompare_models_accuracy([\"EfficientNetB1.keras\", \"EfficientNetB0.keras\"])\ncompare_models_loss([\"EfficientNetB1.keras\", \"EfficientNetB0.keras\"])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 8. Result","metadata":{}},{"cell_type":"markdown","source":"## - Comparing All Models Visually","metadata":{}},{"cell_type":"code","source":"# Comparison of All models [Accuracy & Loss]\ncompare_models_accuracy([\"AlexNet.keras\", \"VGGNet.keras\", \"ResNet34.keras\",\"EfficientNetB2.keras\", \"EfficientNetB1.keras\", \"EfficientNetB0.keras\"])\ncompare_models_loss([\"AlexNet.keras\", \"VGGNet.keras\", \"ResNet34.keras\",\"EfficientNetB2.keras\", \"EfficientNetB1.keras\", \"EfficientNetB0.keras\"])\ncompare_models_accuracy([\"VGGNet.keras\", \"ResNet34.keras\",\"EfficientNetB1.keras\", \"EfficientNetB0.keras\"])\ncompare_models_loss([\"VGGNet.keras\", \"ResNet34.keras\",\"EfficientNetB1.keras\", \"EfficientNetB0.keras\"])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## - Prediction and Evaluation of Models","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"> ### ***Please:***\n>\n>   *Feel free to **share your thoughts in the *discussion*** section—let’s challenge this run together!*","metadata":{}}]}