{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":11848,"databundleVersionId":862157,"sourceType":"competition"}],"dockerImageVersionId":31153,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-10-31T17:03:44.964078Z","iopub.execute_input":"2025-10-31T17:03:44.964330Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Import Required Libraries for CNN Model","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom PIL import Image\nimport os\nimport shutil\nfrom pathlib import Path \nimport keras\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Conv2D, MaxPooling2D, Flatten, BatchNormalization\nfrom tensorflow.keras import regularizers\nimport random\nimport PIL\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\nfrom tensorflow.keras.utils import load_img, img_to_array","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:18:26.336779Z","iopub.execute_input":"2025-11-05T15:18:26.336943Z","iopub.status.idle":"2025-11-05T15:18:41.392867Z","shell.execute_reply.started":"2025-11-05T15:18:26.336928Z","shell.execute_reply":"2025-11-05T15:18:41.392237Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Problem Description","metadata":{}},{"cell_type":"markdown","source":"The dataset contains small image patches taken from larger digital pathology scans. The data for this competition is a slightly modified version of the PatchCamelyon (PCam) benchmark dataset (the original PCam dataset). It consists of 220025 color images (96 x 96 px) extracted from histopathologic scans. Each image has id for unique identification and binary label (0 ,1) associated with it called label. The train dataset contains id and label features. The image is in the form of tif (Tagged Image File Format). A TIF image is a high-quality raster image file that is popular in professional printing and photography. The problem is to classify the images in the test set either they are cancerous or non-cancerous image that is with label 0 or 1. Convolutional Neural Network models are trained for binary image classification. The problem is to build a CNN model that accurately predict or classify the binary label for given images. ","metadata":{}},{"cell_type":"markdown","source":"## Exploratory Data Analysis","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('/kaggle/input/histopathologic-cancer-detection/train_labels.csv')\ndf.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:18:50.817296Z","iopub.execute_input":"2025-11-05T15:18:50.818057Z","iopub.status.idle":"2025-11-05T15:18:51.413628Z","shell.execute_reply.started":"2025-11-05T15:18:50.818034Z","shell.execute_reply":"2025-11-05T15:18:51.413016Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T15:46:14.228525Z","iopub.execute_input":"2025-11-04T15:46:14.228869Z","iopub.status.idle":"2025-11-04T15:46:14.273742Z","shell.execute_reply.started":"2025-11-04T15:46:14.228844Z","shell.execute_reply":"2025-11-04T15:46:14.272223Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T15:46:56.201369Z","iopub.execute_input":"2025-11-04T15:46:56.202523Z","iopub.status.idle":"2025-11-04T15:46:56.209688Z","shell.execute_reply.started":"2025-11-04T15:46:56.202481Z","shell.execute_reply":"2025-11-04T15:46:56.208438Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T15:46:58.629744Z","iopub.execute_input":"2025-11-04T15:46:58.630135Z","iopub.status.idle":"2025-11-04T15:46:58.657391Z","shell.execute_reply.started":"2025-11-04T15:46:58.630089Z","shell.execute_reply":"2025-11-04T15:46:58.656207Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Cleaning","metadata":{}},{"cell_type":"code","source":"df.isna().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T15:56:48.571274Z","iopub.execute_input":"2025-11-04T15:56:48.571664Z","iopub.status.idle":"2025-11-04T15:56:48.596390Z","shell.execute_reply.started":"2025-11-04T15:56:48.571635Z","shell.execute_reply":"2025-11-04T15:56:48.595184Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"df.isnull().sum()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T15:56:59.402704Z","iopub.execute_input":"2025-11-04T15:56:59.403675Z","iopub.status.idle":"2025-11-04T15:56:59.435348Z","shell.execute_reply.started":"2025-11-04T15:56:59.403629Z","shell.execute_reply":"2025-11-04T15:56:59.433183Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Description\n","metadata":{}},{"cell_type":"markdown","source":"## Train Images","metadata":{}},{"cell_type":"markdown","source":"On inspection the dataset contains train_labels dataframe with 220025 rows and 2 columns unique ids with labels (0,1). The shape of the dataframe (220025, 2). The datatype of id feature is object type and label is (int64) integer datatype. The 220025 train images are in the folder '/kaggle/input/histopathologic-cancer-detection/train' in tif format. The size of images is (96 x 96 px).The statistics of label feature shows count as 220025 and min and max are 0 and 1 respectively. There are 130908 class 0 labels and 89117 class 1 labels. ","metadata":{}},{"cell_type":"markdown","source":"## Test Images","metadata":{}},{"cell_type":"markdown","source":"The 57458 test images are in the folder '/kaggle/input/histopathologic-cancer-detection/test' folder and in the tif format. The size of images is (96 x 96 px). There is a sample_submission csv folder shows sample 57458 RGB images and their labels for submission format. Initial inspection shows no null or NA or duplicate values. There are no duplicates in image files.","metadata":{}},{"cell_type":"markdown","source":"## Feature Visualization","metadata":{}},{"cell_type":"markdown","source":"## Histogram of Label Feature","metadata":{}},{"cell_type":"code","source":"sns.countplot(df, x=\"label\", order=df['label'].value_counts().index)\n# Add a title\nplt.title('Label Histogram')\nproportions = list(df['label'].value_counts(normalize=True))\nprint(proportions)\n# Display the chart\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:04:58.079405Z","iopub.execute_input":"2025-11-04T16:04:58.079881Z","iopub.status.idle":"2025-11-04T16:04:58.418462Z","shell.execute_reply.started":"2025-11-04T16:04:58.079850Z","shell.execute_reply":"2025-11-04T16:04:58.416582Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Label histogram reveals that there are **130908 class 0 labels and 89117 class 1 labels**.  In the train set , 59.5% of images with label = 0 (non-cancer) and 40.5% with label = 1 (cancer cells). It shows that there is an imbalance in the class labels and not equally distributed.","metadata":{}},{"cell_type":"markdown","source":"## Pie Chart - Distribution of Labels","metadata":{}},{"cell_type":"code","source":"proportions = list(df['label'].value_counts(normalize=True))\nprint(proportions)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:05:19.904769Z","iopub.execute_input":"2025-11-04T16:05:19.905867Z","iopub.status.idle":"2025-11-04T16:05:19.917006Z","shell.execute_reply.started":"2025-11-04T16:05:19.905816Z","shell.execute_reply":"2025-11-04T16:05:19.915518Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.pie(proportions,labels = df['label'].unique(),\n        autopct='%1.1f%%',  # Display percentages on slices\n        shadow=True,\n        startangle=140)\n\n# circular pie chart\nplt.axis('equal')\nplt.title('Label Distribution')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:05:22.518048Z","iopub.execute_input":"2025-11-04T16:05:22.518439Z","iopub.status.idle":"2025-11-04T16:05:22.668329Z","shell.execute_reply.started":"2025-11-04T16:05:22.518414Z","shell.execute_reply":"2025-11-04T16:05:22.667349Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Pie chart reveals that in the train set , 59.5% of images with label = 0 (non-cancer) and 40.5% with label = 1 (cancer cells). It shows that there is an imbalance in the class labels and not equally distributed.","metadata":{}},{"cell_type":"code","source":"dfs = pd.read_csv('/kaggle/input/histopathologic-cancer-detection/sample_submission.csv')\ndfs.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:01:49.125602Z","iopub.execute_input":"2025-11-04T16:01:49.125978Z","iopub.status.idle":"2025-11-04T16:01:49.205625Z","shell.execute_reply.started":"2025-11-04T16:01:49.125956Z","shell.execute_reply":"2025-11-04T16:01:49.204638Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Train Image Files","metadata":{}},{"cell_type":"code","source":"image_directory = '/kaggle/input/histopathologic-cancer-detection/train/'\n    \n# list of image file paths\nimage_files = [os.path.join(image_directory, f) for f in os.listdir(image_directory) if f.endswith(('.jpg', '.png', '.jpeg', '.tif'))]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:06:17.047344Z","iopub.execute_input":"2025-11-04T16:06:17.048366Z","iopub.status.idle":"2025-11-04T16:06:24.824534Z","shell.execute_reply.started":"2025-11-04T16:06:17.048337Z","shell.execute_reply":"2025-11-04T16:06:24.823051Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Number of Train Images","metadata":{}},{"cell_type":"code","source":"len(image_files)   # 220025","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:06:28.282736Z","iopub.execute_input":"2025-11-04T16:06:28.283031Z","iopub.status.idle":"2025-11-04T16:06:28.292465Z","shell.execute_reply.started":"2025-11-04T16:06:28.283013Z","shell.execute_reply":"2025-11-04T16:06:28.290819Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Check for Duplicates in Train Images","metadata":{}},{"cell_type":"code","source":"from collections import Counter\n\ndef find_duplicates_with_counter(input_list):\n    counts = Counter(input_list)\n    duplicates = [item for item, count in counts.items() if count > 1]\n    return duplicates\n\nduplicate_values = find_duplicates_with_counter(image_files)\nprint(f\"Duplicate values: {duplicate_values}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:06:43.251327Z","iopub.execute_input":"2025-11-04T16:06:43.252388Z","iopub.status.idle":"2025-11-04T16:06:43.416759Z","shell.execute_reply.started":"2025-11-04T16:06:43.252348Z","shell.execute_reply":"2025-11-04T16:06:43.414633Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There are no duplicate images in train image set.","metadata":{}},{"cell_type":"markdown","source":"## Test Image Files","metadata":{}},{"cell_type":"code","source":"image_directory_test = '/kaggle/input/histopathologic-cancer-detection/test/'\n    \n# list of image file paths\nimage_files_test = [os.path.join(image_directory_test, f) for f in os.listdir(image_directory_test) if f.endswith(('.jpg', '.png', '.jpeg', '.tif'))]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:03:50.450460Z","iopub.execute_input":"2025-11-04T16:03:50.450825Z","iopub.status.idle":"2025-11-04T16:03:50.553470Z","shell.execute_reply.started":"2025-11-04T16:03:50.450799Z","shell.execute_reply":"2025-11-04T16:03:50.551451Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Number of Test Images","metadata":{}},{"cell_type":"code","source":"len(image_files_test)   ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:03:53.077982Z","iopub.execute_input":"2025-11-04T16:03:53.078417Z","iopub.status.idle":"2025-11-04T16:03:53.087658Z","shell.execute_reply.started":"2025-11-04T16:03:53.078390Z","shell.execute_reply":"2025-11-04T16:03:53.086184Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Check for Duplicates","metadata":{}},{"cell_type":"code","source":"duplicate_values = find_duplicates_with_counter(image_files_test)\nprint(f\"Duplicate values: {duplicate_values}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-30T19:05:59.589318Z","iopub.execute_input":"2025-10-30T19:05:59.589736Z","iopub.status.idle":"2025-10-30T19:05:59.615374Z","shell.execute_reply.started":"2025-10-30T19:05:59.589711Z","shell.execute_reply":"2025-10-30T19:05:59.613940Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There are no duplicates in test images too.","metadata":{}},{"cell_type":"markdown","source":"## Image Visualizations","metadata":{}},{"cell_type":"markdown","source":"## Plots of Class Label 0 and 1 Train Images","metadata":{}},{"cell_type":"code","source":"# Seperate dataframes for label 0 and label 1\ndf_0 = df[df['label'] == 0]\ndf_1 = df[df['label'] == 1]\n\ndf_0_ids = df_0['id'].head(3).tolist()\ndf_1_ids = df_1['id'].head(3).tolist()\n\ns = '.tif'\npath = '/kaggle/input/histopathologic-cancer-detection/train/'\n\ndef add_tif(id_list):\n    for i in range(len(id_list)):\n        id_list[i] = str(path + id_list[i] + s)\n    return id_list\n    \ndf_0_ids_tif = add_tif(df_0_ids)\ndf_1_ids_tif = add_tif(df_1_ids)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:19:01.261146Z","iopub.execute_input":"2025-11-05T15:19:01.261872Z","iopub.status.idle":"2025-11-05T15:19:01.282488Z","shell.execute_reply.started":"2025-11-05T15:19:01.261843Z","shell.execute_reply":"2025-11-05T15:19:01.281939Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, axs = plt.subplots(2, 3, figsize=(7, 5))\n\n# Flatten the `axs` array for easier iteration\naxs = axs.flatten()\n\n# Combine both image lists into one \nall_image_paths = df_0_ids_tif + df_1_ids_tif\ntitles = [f'Label 0 ' for i in range(3)] + [f'Label 1 ' for i in range(3)]\n\n# Display each image in its own subplot\nfor i, img_path in enumerate(all_image_paths):\n    img = Image.open(img_path)\n    axs[i].imshow(img)\n    axs[i].set_title(titles[i])\n    axs[i].axis('off')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:11:44.028979Z","iopub.execute_input":"2025-11-04T16:11:44.029388Z","iopub.status.idle":"2025-11-04T16:11:44.704626Z","shell.execute_reply.started":"2025-11-04T16:11:44.029360Z","shell.execute_reply":"2025-11-04T16:11:44.703376Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Sample 80K Images to Train CNN Model","metadata":{}},{"cell_type":"code","source":"#df_0.shape # (130908, 2)\n#df_1.shape # (89117, 2)\n#len(df_0_ids_tif) #3\ndef add_tif1(id_list):\n    for i in range(len(id_list)):\n        id_list[i] = str(id_list[i] + s)\n    return id_list\n    \ndf_0_ids_tif = add_tif(df_0_ids)\ndf_1_ids_tif = add_tif(df_1_ids)\n\ndf_0_ids_20k = df_0['id'].sample(n=40000, random_state=42).tolist()\ndf_1_ids_20k = df_1['id'].sample(n=40000, random_state=42).tolist()\n    \ndf_0_ids_20k_tif = add_tif1(df_0_ids_20k)\ndf_1_ids_20k_tif = add_tif1(df_1_ids_20k)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:19:10.770803Z","iopub.execute_input":"2025-11-05T15:19:10.771081Z","iopub.status.idle":"2025-11-05T15:19:10.810730Z","shell.execute_reply.started":"2025-11-05T15:19:10.771063Z","shell.execute_reply":"2025-11-05T15:19:10.809963Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Sample image of each label 0 and 1 are chosen and plotted for visual comparison. Non-cancer and cancer images scan in train dataset.","metadata":{}},{"cell_type":"markdown","source":"### Properties of Images","metadata":{}},{"cell_type":"markdown","source":"### Size of Image","metadata":{}},{"cell_type":"code","source":"import PIL\nPIL.Image.open(\"/kaggle/input/histopathologic-cancer-detection/test/f5be692779144a3ecdaed9b82f9487564edcbccb.tif\").size","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:14:21.623567Z","iopub.execute_input":"2025-11-04T16:14:21.624003Z","iopub.status.idle":"2025-11-04T16:14:21.643187Z","shell.execute_reply.started":"2025-11-04T16:14:21.623976Z","shell.execute_reply":"2025-11-04T16:14:21.641476Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Image as Array","metadata":{}},{"cell_type":"code","source":"# Compress Image to array for easy computation\nimg = PIL.Image.open('/kaggle/input/histopathologic-cancer-detection/test/f5be692779144a3ecdaed9b82f9487564edcbccb.tif').resize((96, 96))\n\nrgb_pixels = np.array(img)\nrgb_pixels.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:14:27.354698Z","iopub.execute_input":"2025-11-04T16:14:27.355184Z","iopub.status.idle":"2025-11-04T16:14:27.367857Z","shell.execute_reply.started":"2025-11-04T16:14:27.355155Z","shell.execute_reply":"2025-11-04T16:14:27.366583Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Image Mode","metadata":{}},{"cell_type":"code","source":"# Find the mode of given image : RGB or RGBA\nprint(img.mode)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:14:37.029884Z","iopub.execute_input":"2025-11-04T16:14:37.030377Z","iopub.status.idle":"2025-11-04T16:14:37.037264Z","shell.execute_reply.started":"2025-11-04T16:14:37.030345Z","shell.execute_reply":"2025-11-04T16:14:37.035628Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The size of image is 96 x 96 pixels and have 3 channels RGB which is the mode of the images. The image can be represented as Numpy array and its shape is (96, 96, 3)","metadata":{}},{"cell_type":"markdown","source":"### Image Conversion to Grayscale","metadata":{}},{"cell_type":"code","source":"grayscale_image = img.convert(\"L\")\ngrayscale_image","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:14:48.166903Z","iopub.execute_input":"2025-11-04T16:14:48.167341Z","iopub.status.idle":"2025-11-04T16:14:48.182256Z","shell.execute_reply.started":"2025-11-04T16:14:48.167314Z","shell.execute_reply":"2025-11-04T16:14:48.180817Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Image Resizing","metadata":{}},{"cell_type":"code","source":"resized_img = img.resize((200, 200))\nresized_img","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:15:02.089766Z","iopub.execute_input":"2025-11-04T16:15:02.090242Z","iopub.status.idle":"2025-11-04T16:15:02.112634Z","shell.execute_reply.started":"2025-11-04T16:15:02.090210Z","shell.execute_reply":"2025-11-04T16:15:02.110497Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Rotated Image","metadata":{}},{"cell_type":"code","source":"rotated_img = resized_img.rotate(90)\nrotated_img","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:15:09.561415Z","iopub.execute_input":"2025-11-04T16:15:09.561810Z","iopub.status.idle":"2025-11-04T16:15:09.581456Z","shell.execute_reply.started":"2025-11-04T16:15:09.561782Z","shell.execute_reply":"2025-11-04T16:15:09.579898Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Images can be greyscaled (black and white version), resized (creates a copy of the given image to the specified dimensions (width and height)) and rotated to angles (The rotation angle in degrees, measured counter-clockwise. Use a negative value to rotate clockwise). The PIL package refers to the Python Imaging Library, which is the standard library for image processing in Python","metadata":{}},{"cell_type":"markdown","source":"## Three-channel RGB Representation","metadata":{}},{"cell_type":"code","source":"plt.imshow(rgb_pixels)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:24:00.463491Z","iopub.execute_input":"2025-11-04T16:24:00.464035Z","iopub.status.idle":"2025-11-04T16:24:00.668918Z","shell.execute_reply.started":"2025-11-04T16:24:00.463916Z","shell.execute_reply":"2025-11-04T16:24:00.667269Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The imshow function in matplotlib takes a NumPy array as input allocates a color to every element in the array according to its value, and convert those numbers into a visual image. The rgb_pixels array must have a shape that represents image data, such as (M, N, 3) for an RGB image or (M, N, 4) for an RGBA image, where M and N are the height and width of the image.","metadata":{}},{"cell_type":"markdown","source":"## Plot of 1-Channel Image on Monochrome","metadata":{}},{"cell_type":"code","source":"fig, axs = plt.subplots(1, 3, figsize=(9, 5))\n\n# Flatten the `axs` array \naxs = axs.flatten()\ntitle = ['red channel', 'green channel', 'blue channel']\n\n# Display each image in its subplot\nfor i in range(0, 3):\n    axs[i].imshow(rgb_pixels[:, :, i], cmap='Greys')\n    axs[i].set_title(title[i])\n    axs[i].axis('off')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:24:15.208892Z","iopub.execute_input":"2025-11-04T16:24:15.209320Z","iopub.status.idle":"2025-11-04T16:24:15.606918Z","shell.execute_reply.started":"2025-11-04T16:24:15.209293Z","shell.execute_reply":"2025-11-04T16:24:15.605501Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The above plots shows the red, green, and blue color channels of an image as three separate grayscale images. Images will be a grayscale representation of each channel in the pixel of the original rgb_pixels array. This is a common technique used in image processing to analyze the intensity of each channel independently. ","metadata":{}},{"cell_type":"markdown","source":"## Image Crop - Left and Right","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(nrows=1, ncols=3, figsize=(15, 5)) \n\n# Plot the full image \naxes[0].imshow(rgb_pixels)\naxes[0].set_title(\"Full image\")\n\n# Plot the left half \naxes[1].imshow(rgb_pixels[0:96, 0:47])\naxes[1].set_title(\"Left half\")\n\n# Plot the right half\naxes[2].imshow(rgb_pixels[0:96, 47:96])\naxes[2].set_title(\"Right half\")\n\nfig.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:24:39.337902Z","iopub.execute_input":"2025-11-04T16:24:39.338311Z","iopub.status.idle":"2025-11-04T16:24:39.888152Z","shell.execute_reply.started":"2025-11-04T16:24:39.338286Z","shell.execute_reply":"2025-11-04T16:24:39.886853Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The above plot shows three subplots displaying first the complete image, second and third displays the left and right half of the full image demonstrating the image crop.","metadata":{}},{"cell_type":"markdown","source":"## Pixel Distribution of Sample Image","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\n\n# getting a sample image\nsample_image = mpimg.imread(\"/kaggle/input/histopathologic-cancer-detection/test/f5be692779144a3ecdaed9b82f9487564edcbccb.tif\")\n\ncolors = ('red', 'green', 'blue')\nfor i, color in enumerate(colors):\n    plt.hist(sample_image[:, :, i].ravel(), bins=256, color=color, alpha=0.5)\nplt.title('Pixel Distribution (RGB Histograms)')\nplt.xlabel('Pixel Intensity')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:25:54.381760Z","iopub.execute_input":"2025-11-04T16:25:54.382201Z","iopub.status.idle":"2025-11-04T16:25:55.414612Z","shell.execute_reply.started":"2025-11-04T16:25:54.382176Z","shell.execute_reply":"2025-11-04T16:25:55.413301Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In the Pillow (PIL) library, analyzing the pixel distribution is typically done by generating an image histogram. The primary methods for this are Image.histogram() and the ImageStat module. This method calculates the frequency of each pixel value within the image. It returns a list of pixel counts.","metadata":{}},{"cell_type":"markdown","source":"## Image Statistics","metadata":{}},{"cell_type":"code","source":"from PIL import Image, ImageStat\nstat = ImageStat.Stat(img)\n\nprint(f\"Mean pixel value: {stat.mean[0]:.2f}\")\nprint(f\"Median pixel value: {stat.median[0]}\")\nprint(f\"Min/Max pixel values: {stat.extrema[0]}\")\nprint(f\"Standard deviation: {stat.stddev[0]:.2f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:31:14.334821Z","iopub.execute_input":"2025-11-04T16:31:14.335200Z","iopub.status.idle":"2025-11-04T16:31:14.343151Z","shell.execute_reply.started":"2025-11-04T16:31:14.335177Z","shell.execute_reply":"2025-11-04T16:31:14.341991Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## BoxPlot of Pixel Distribution","metadata":{}},{"cell_type":"code","source":"if sample_image.ndim == 3 :\n    red_channel = sample_image[:, :, 0].flatten()\n    green_channel = sample_image[:, :, 1].flatten()\n    blue_channel = sample_image[:, :, 2].flatten()\n    \n    data = [red_channel, green_channel, blue_channel]\n    labels = ['Red', 'Green', 'Blue']\n    title = 'Pixel Distribution by Color Channel'\n\nplt.figure(figsize=(10, 6))\nplt.boxplot(data, vert=True, patch_artist=True, labels=labels)\nplt.title(title)\nplt.ylabel('Pixel Value')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T16:31:20.666265Z","iopub.execute_input":"2025-11-04T16:31:20.666599Z","iopub.status.idle":"2025-11-04T16:31:20.873589Z","shell.execute_reply.started":"2025-11-04T16:31:20.666577Z","shell.execute_reply":"2025-11-04T16:31:20.872112Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Preprocessing Images","metadata":{}},{"cell_type":"markdown","source":"## Transfer Label 0,1 images to /kaggle/working/ 0,1 subdirectories","metadata":{}},{"cell_type":"code","source":"# Define the names of the directories to be created\ndirectory1_name = \"0\"\ndirectory2_name = \"1\"\n\n# Construct the full paths for the new directories\npath_dir1 = os.path.join(\"/kaggle/working/\", directory1_name)\npath_dir2 = os.path.join(\"/kaggle/working/\", directory2_name)\n\n# Create the first directory \nif not os.path.exists(path_dir1):\n    os.mkdir(path_dir1)\n    print(f\"Directory '{directory1_name}' created at {path_dir1}\")\nelse:\n    print(f\"Directory '{directory1_name}' already exists at {path_dir1}\")\n\n# Create the second directory \nif not os.path.exists(path_dir2):\n    os.mkdir(path_dir2)\n    print(f\"Directory '{directory2_name}' created at {path_dir2}\")\nelse:\n    print(f\"Directory '{directory2_name}' already exists at {path_dir2}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:19:31.413342Z","iopub.execute_input":"2025-11-05T15:19:31.413929Z","iopub.status.idle":"2025-11-05T15:19:31.420057Z","shell.execute_reply.started":"2025-11-05T15:19:31.413905Z","shell.execute_reply":"2025-11-05T15:19:31.419420Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def transfer_images(source_dir: str, destination_dir: str, image_list: list[str]):\n    \"\"\"\n    Copies a list of images from a source directory to a destination directory.\n\n    Args:\n        source_dir (str): The path to the source directory.\n        destination_dir (str): The path to the destination directory.\n        image_list (list[str]): A list of filenames (strings) to be transferred.\n    \"\"\"\n    \n    # Use Path objects for clean and reliable path joining\n    source = Path(source_dir)\n    destination = Path(destination_dir)\n\n    # Ensure the destination directory exists; create it if necessary\n    destination.mkdir(parents=True, exist_ok=True)\n    \n    for image_name in image_list:\n        source_path = source / image_name\n        destination_path = destination / image_name\n\n        if source_path.is_file():\n            # shutil.copy2 copies file data and metadata (timestamps, etc.)\n            shutil.copy2(source_path, destination_path)\n\n    print(\"Files copied\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:19:39.233919Z","iopub.execute_input":"2025-11-05T15:19:39.234664Z","iopub.status.idle":"2025-11-05T15:19:39.239799Z","shell.execute_reply.started":"2025-11-05T15:19:39.234640Z","shell.execute_reply":"2025-11-05T15:19:39.238997Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Copies Train Label 1 Files to /kaggle/working/0 Directory","metadata":{}},{"cell_type":"code","source":"# 1. Define paths and list of files\nSOURCE_DIR = \"/kaggle/input/histopathologic-cancer-detection/train/\"\nDEST_DIR = \"/kaggle/working/0/\"\nIMAGES_TO_COPY = df_0_ids_20k_tif\n\ntransfer_images(SOURCE_DIR, DEST_DIR, IMAGES_TO_COPY)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:19:46.505574Z","iopub.execute_input":"2025-11-05T15:19:46.505859Z","iopub.status.idle":"2025-11-05T15:25:58.555976Z","shell.execute_reply.started":"2025-11-05T15:19:46.505840Z","shell.execute_reply":"2025-11-05T15:25:58.555302Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Copies Train Label 1 Files to /kaggle/working/1 Directory","metadata":{}},{"cell_type":"code","source":"DEST_DIR = \"/kaggle/working/1/\"\nIMAGES_TO_COPY = df_1_ids_20k_tif\n\ntransfer_images(SOURCE_DIR, DEST_DIR, IMAGES_TO_COPY)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:26:02.713381Z","iopub.execute_input":"2025-11-05T15:26:02.713657Z","iopub.status.idle":"2025-11-05T15:32:37.084524Z","shell.execute_reply.started":"2025-11-05T15:26:02.713638Z","shell.execute_reply":"2025-11-05T15:32:37.083717Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Function converts .tif to .bmp ","metadata":{}},{"cell_type":"code","source":"def convert_tif_to_bmp(tif_path, bmp_path):\n    \"\"\"\n    Converts a single TIF file to a BMP file losslessly.\n    \"\"\" \n    # Open the TIF image\n    with Image.open(tif_path) as im:\n        # Save the image as a BMP file. BMP is lossless by nature.\n        im.save(bmp_path, 'BMP')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:33:12.213864Z","iopub.execute_input":"2025-11-05T15:33:12.214155Z","iopub.status.idle":"2025-11-05T15:33:12.218316Z","shell.execute_reply.started":"2025-11-05T15:33:12.214133Z","shell.execute_reply":"2025-11-05T15:33:12.217530Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Converts .tif Files in /kaggle/working/0 and 1 files to .bmp","metadata":{}},{"cell_type":"code","source":"image_directory_working0 = '/kaggle/working/0/'\nimage_files_working0 = [f for f in os.listdir(image_directory_working0) if f.endswith(('.tif'))]\nsource_tif_dir = image_directory_working0\ndestination_bmp_dir = image_directory_working0 # Destination is same directory\n\nfor filename in image_files_working0:\n    source_path = os.path.join(source_tif_dir, filename)\n    \n    # 1. Change the file extension from .tif to .bmp for the destination filename\n    base, ext = os.path.splitext(filename)\n    destination_filename = base + '.bmp'\n\n    # 2. destination file path\n    destination_path = os.path.join(destination_bmp_dir, destination_filename)\n    \n    # Call  conversion function \n    convert_tif_to_bmp(source_path, destination_path)\n\n\n# --- Loop 1 ---\nimage_directory_working1 = '/kaggle/working/1/'\nimage_files_working1 = [f for f in os.listdir(image_directory_working1) if f.endswith(('.tif'))]\nsource_tif_dir_1 = image_directory_working1\ndestination_bmp_dir_1 = image_directory_working1\n\nfor filename in image_files_working1:\n    source_path = os.path.join(source_tif_dir_1, filename)\n    \n    # 1. Change the file extension from .tif to .bmp for the destination filename\n    base, ext = os.path.splitext(filename)\n    destination_filename = base + '.bmp'\n\n    # 2. destination file path\n    destination_path = os.path.join(destination_bmp_dir_1, destination_filename)\n    \n    # Call  conversion function \n    convert_tif_to_bmp(source_path, destination_path)\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:33:15.017417Z","iopub.execute_input":"2025-11-05T15:33:15.017700Z","iopub.status.idle":"2025-11-05T15:34:03.889669Z","shell.execute_reply.started":"2025-11-05T15:33:15.017681Z","shell.execute_reply":"2025-11-05T15:34:03.888826Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Preprocessing Images is a multi step process. \n\n* First step is to create 2 subdirectories in /kaggle/working directory with label names '0' and '1'.\n* Copy the sample train with label 0 images to /kaggle/working/0 directory.\n* Copy the sample train with label 1 images to /kaggle/working/1 directory.\n* Convert the tif format image files in /kaggle/working/0 directory to bmp lossless format to train the CNN models.\n* Convert the tif format image files in /kaggle/working/1 directory to bmp lossless format to train the CNN models.","metadata":{}},{"cell_type":"markdown","source":"There is another approach of converting images to Numpy array. I tried that method but it did not work for me.","metadata":{}},{"cell_type":"markdown","source":"Since it took longer and not enough GPU (got only limited use hours of GPU), I have taken only 40K from label 0 and 40K of label 1 from train images to train models. Also, the test images are not able to convert to bmp format, so I am using a fraction of train images itsel as validation or test set. ","metadata":{}},{"cell_type":"markdown","source":"## Train and Test Dataset (Preprocessed)","metadata":{}},{"cell_type":"code","source":"\n# This function loads images, and gets labels from the directory structure\ntrain_ds = tf.keras.utils.image_dataset_from_directory(\n    '/kaggle/working/',\n    labels='inferred',  \n    label_mode='binary', \n    image_size=(96, 96),  # The image size for this dataset\n    batch_size=32,\n    validation_split=0.3,\n    subset='training',\n    shuffle=True,seed = 42\n)\n\nval_ds = tf.keras.utils.image_dataset_from_directory(\n    '/kaggle/working/',\n    labels='inferred',\n    label_mode='binary',\n    image_size=(96, 96),\n    batch_size=32,\n    validation_split=0.3,\n    subset='validation', # This is the validation set\n    shuffle=True,seed = 42\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:34:12.288471Z","iopub.execute_input":"2025-11-05T15:34:12.289471Z","iopub.status.idle":"2025-11-05T15:34:22.276682Z","shell.execute_reply.started":"2025-11-05T15:34:12.289440Z","shell.execute_reply":"2025-11-05T15:34:22.275929Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Execution Plan","metadata":{}},{"cell_type":"markdown","source":"The Stochastic Gradient Descent (SGD) optimizer is a fundamental and widely used iterative optimization algorithm in machine learning and deep learning. Its goal is to find the set of model parameters (weights and biases) that minimizes a loss function. The plan is to use SGD optimizer with some pyperparameter tuning like adjusting learning rate with number of epochs. Accuracy and loss will be recorded. Activation functions sigmoid and softmax will be used for comparison.\n\nNext Adam optimizer will be used for its efficiency and robustness and to train the model with some pyperparameter tuning. Accuracy and loss will be recorded and compared.\n\nFinally RMSProp optimizer that adjusts the learning rate for each parameter individually based on the history of its gradients. Model is trained using this optimizer and performances are compared.","metadata":{}},{"cell_type":"markdown","source":"## Model Construction","metadata":{}},{"cell_type":"markdown","source":"### SGD Optimizer","metadata":{}},{"cell_type":"code","source":"input_shape = (96, 96, 3)\n\n# Create a Sequential model with an Input layer\nmodel = Sequential()\n# Image Normalization \ntf.keras.layers.Rescaling(1./255, input_shape=(96, 96, 3)), \n\n# Add the hidden layer\n# Conv2D: 32 filters, 3x3 kernel, ReLU activation\n# MaxPooling2D: 2x2 pooling to reduce dimensions\n\nmodel.add(Conv2D(32, (3, 3), activation='relu', input_shape=input_shape))                  \nmodel.add(BatchNormalization()) # Add Batch Normalization\nmodel.add(MaxPooling2D((2, 2)))\n\n# Conv2D: 64 filters, 3x3 kernel, ReLU activation\n# MaxPooling2D: 2x2 pooling to reduce dimensions\nmodel.add(Conv2D(64, (3, 3), activation='relu', input_shape=input_shape))\nmodel.add(BatchNormalization()) # Added Batch Normalization\nmodel.add(MaxPooling2D((2, 2)))\n\n# Flatten the output of the convolutional layers\n\nmodel.add(Flatten())\n\n# Add the output layer\n# Dense: 1 neuron with a sigmoid activation for binary classification\nmodel.add(Dense(1, activation='sigmoid', \n                kernel_regularizer=regularizers.l1_l2(l1=0.001, l2=0.0001)))\n\n# Compile the model\n# 'binary_crossentropy' for binary classification\nsgd_optimizer = keras.optimizers.SGD(learning_rate=0.0001) \nmodel.compile(optimizer=sgd_optimizer, \n              loss='binary_crossentropy',\n              metrics=['accuracy'])\n\n#model's architecture\nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:35:23.983694Z","iopub.execute_input":"2025-11-05T15:35:23.984000Z","iopub.status.idle":"2025-11-05T15:35:25.074335Z","shell.execute_reply.started":"2025-11-05T15:35:23.983980Z","shell.execute_reply":"2025-11-05T15:35:25.072928Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Model Architecture","metadata":{}},{"cell_type":"markdown","source":"The model processes images sequentially :\n\n**Input and Normalization**\n \n**Input Shape**: Expects images of size (96, 96) pixels with 3 color channels (RGB).\n\n**Rescaling**(1./255, ...):  It normalizes the raw pixel values (which range from 0 to 255) down to floating-point values between 0.0 and 1.0. \n\n**First Convolutional Block**\n \n**Conv2D**(32, (3, 3), activation='relu', ...): Applies 32 different filters (kernels) of size 3x3 pixels across the image. The ReLU activation function introduces non-linearity. \n\n**Batch Normalization**: Included for faster training, convergence and bit of regularization effects.\n\n**MaxPooling2D**((2, 2)):Reduce feature maps dimension (from 96x96 down to 48x48).\n\n**Second Convolutional Block**\n\n**Conv2D**(64, (3, 3), activation='relu'): Similar to the first block, but 64 filters. \n\n**Batch Normalization**: Included for faster training, convergence and bit of regularization effects.\n\n**MaxPooling2D**((2, 2)): Further reduces the spatial dimensions (from 48x48 down to 24x24).\n\n**Output Layer**\n\n**Dense**(1, activation='sigmoid', ...): The final classification layer.\n\n**Units** = 1: A single neuron is used, which is standard for binary classification.\n\n**activation** ='sigmoid': The sigmoid function squashes the output into a single probability value between 0 and 1.\n\n**Regularization**: Includes combined L1 (0.001) and L2 (0.01) regularization to control the final weights of the decision boundary.\n\n**Compilation** \n\n**optimizer**=SGD(learning_rate=0.0001): Uses Stochastic Gradient Descent with a relatively small \n\n**learning rate** (0.0001). The small rate suggests a cautious approach to training, perhaps to avoid instability common in deep CNNs when using basic SGD.\n\n**loss**='binary_crossentropy': The standard and correct loss function for a binary classification task where the output layer uses a sigmoid activation.\n\n**metrics**=['accuracy']: The model will track and report classification accuracy during training in addition to the loss.","metadata":{}},{"cell_type":"markdown","source":"**Model Summary**\n\n**Architecture**: A Sequential CNN with two convolutional/pooling blocks, a flatten layer, and a single neuron dense output layer for binary classification.\n\n**Normalization**: Pixels are scaled from 0-255 to 0.0-1.0.\n\n**Regularization**: L2 is applied to the first Conv2D layer, and L1/L2 combined regularization is applied to the final Dense layer.\n\n**Optimizer**: Stochastic Gradient Descent (SGD) with a learning rate of 0.0001.\n\n**Loss/Metrics**: Binary cross-entropy loss and accuracy metric.","metadata":{}},{"cell_type":"markdown","source":"## Train the model","metadata":{}},{"cell_type":"code","source":"history = model.fit(train_ds,epochs=150)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T22:58:53.267268Z","iopub.execute_input":"2025-11-04T22:58:53.267895Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_loss = history.history['loss']\ntrain_accuracy = history.history['accuracy']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T20:09:48.449846Z","iopub.execute_input":"2025-11-04T20:09:48.450441Z","iopub.status.idle":"2025-11-04T20:09:48.453987Z","shell.execute_reply.started":"2025-11-04T20:09:48.450418Z","shell.execute_reply":"2025-11-04T20:09:48.453218Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Test the model","metadata":{}},{"cell_type":"code","source":"history = model.fit(val_ds,epochs=150)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"val_loss = history.history['loss']\nval_accuracy = history.history['accuracy'] ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T20:19:34.089652Z","iopub.execute_input":"2025-11-04T20:19:34.090203Z","iopub.status.idle":"2025-11-04T20:19:34.093655Z","shell.execute_reply.started":"2025-11-04T20:19:34.090181Z","shell.execute_reply":"2025-11-04T20:19:34.092902Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Training (Vs) Validation Accuracy","metadata":{}},{"cell_type":"code","source":"#epochs_range = range(100) # First Tuning\nepochs_range = range(150) # Second Tuning\n\n# Plot for Accuracy - Train and Validation set\nplt.figure(figsize=(12, 4))\nplt.subplot(1, 2, 1)\nplt.plot(epochs_range, train_accuracy, label='Training Accuracy')\nplt.plot(epochs_range, val_accuracy, label='Validation Accuracy')\nplt.legend(loc='lower right')\nplt.title('Training and Validation Accuracy')\n\n# Plot for Loss - Train and Validation set\nplt.subplot(1, 2, 2)\nplt.plot(epochs_range, train_loss, label='Training Loss')\nplt.plot(epochs_range, val_loss, label='Validation Loss')\nplt.legend(loc='upper right')\nplt.title('Training and Validation Loss')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-04T20:19:49.833508Z","iopub.execute_input":"2025-11-04T20:19:49.834215Z","iopub.status.idle":"2025-11-04T20:19:50.120774Z","shell.execute_reply.started":"2025-11-04T20:19:49.834192Z","shell.execute_reply":"2025-11-04T20:19:50.120053Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Adam Optimizer","metadata":{}},{"cell_type":"code","source":"# adam optimizer\nadam_optimizer = keras.optimizers.Adam(learning_rate=0.0001)\n\n# Compile the model\n# 'binary_crossentropy'  for binary classification\nmodel.compile(optimizer=adam_optimizer, \n              loss='binary_crossentropy',\n              metrics=['accuracy'])\n\n# model's architecture\nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:35:36.641712Z","iopub.execute_input":"2025-11-05T15:35:36.642586Z","iopub.status.idle":"2025-11-05T15:35:36.663940Z","shell.execute_reply.started":"2025-11-05T15:35:36.642561Z","shell.execute_reply":"2025-11-05T15:35:36.663442Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Training the Model","metadata":{}},{"cell_type":"code","source":"history = model.fit(train_ds,epochs=50, batch_size = 32)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"history = model.fit(train_ds,epochs=75, batch_size = 64)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:51:17.329466Z","iopub.execute_input":"2025-11-05T15:51:17.330093Z","iopub.status.idle":"2025-11-05T16:04:14.984381Z","shell.execute_reply.started":"2025-11-05T15:51:17.330073Z","shell.execute_reply":"2025-11-05T16:04:14.983757Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Collect Accuracy and Loss Scores","metadata":{}},{"cell_type":"code","source":"train_loss = history.history['loss']\ntrain_accuracy = history.history['accuracy']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:49:17.328867Z","iopub.execute_input":"2025-11-05T15:49:17.329468Z","iopub.status.idle":"2025-11-05T15:49:17.332888Z","shell.execute_reply.started":"2025-11-05T15:49:17.329445Z","shell.execute_reply":"2025-11-05T15:49:17.332056Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Testing the Model","metadata":{}},{"cell_type":"code","source":"val_history = model.fit(val_ds,epochs=50, batch_size = 32)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:45:02.498414Z","iopub.execute_input":"2025-11-05T15:45:02.499077Z","iopub.status.idle":"2025-11-05T15:48:44.376560Z","shell.execute_reply.started":"2025-11-05T15:45:02.499052Z","shell.execute_reply":"2025-11-05T15:48:44.375886Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"val_history = model.fit(val_ds,epochs=75, batch_size = 64)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T16:05:10.684721Z","iopub.execute_input":"2025-11-05T16:05:10.685047Z","iopub.status.idle":"2025-11-05T16:10:44.859204Z","shell.execute_reply.started":"2025-11-05T16:05:10.685027Z","shell.execute_reply":"2025-11-05T16:10:44.858621Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"val_loss = val_history.history['loss']\nval_accuracy = val_history.history['accuracy'] ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:49:09.242044Z","iopub.execute_input":"2025-11-05T15:49:09.242359Z","iopub.status.idle":"2025-11-05T15:49:09.246340Z","shell.execute_reply.started":"2025-11-05T15:49:09.242337Z","shell.execute_reply":"2025-11-05T15:49:09.245509Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Train (Vs) Validation Accuracy (Adam Optimizer)","metadata":{}},{"cell_type":"code","source":"#epochs_range = range(50)  # for first tuning\nepochs_range = range(75)   # for second tuning\n\n# Plot for Accuracy - Train and Validation set\nplt.figure(figsize=(12, 4))\nplt.subplot(1, 2, 1)\nplt.plot(epochs_range, train_accuracy, label='Training Accuracy')\nplt.plot(epochs_range, val_accuracy, label='Validation Accuracy')\nplt.legend(loc='lower right')\nplt.title('Training and Validation Accuracy')\n\n# Plot for Loss - Train and Validation set\nplt.subplot(1, 2, 2)\nplt.plot(epochs_range, train_loss, label='Training Loss')\nplt.plot(epochs_range, val_loss, label='Validation Loss')\nplt.legend(loc='upper right')\nplt.title('Training and Validation Loss')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T15:50:24.201101Z","iopub.execute_input":"2025-11-05T15:50:24.201908Z","iopub.status.idle":"2025-11-05T15:50:24.511982Z","shell.execute_reply.started":"2025-11-05T15:50:24.201883Z","shell.execute_reply":"2025-11-05T15:50:24.511390Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## RMSProp Optimizer","metadata":{}},{"cell_type":"code","source":"# Compile the model\n# 'binary_crossentropy' for binary classification\nrmsprop_optimizer = keras.optimizers.RMSprop(learning_rate=0.0001)\nmodel.compile(optimizer=rmsprop_optimizer, #'adam',\n              loss='binary_crossentropy',\n              metrics=['accuracy'])\n\n#model's architecture\nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T16:11:03.078790Z","iopub.execute_input":"2025-11-05T16:11:03.079601Z","iopub.status.idle":"2025-11-05T16:11:03.104622Z","shell.execute_reply.started":"2025-11-05T16:11:03.079577Z","shell.execute_reply":"2025-11-05T16:11:03.103865Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Train the model","metadata":{}},{"cell_type":"code","source":"history = model.fit(train_ds,epochs=20, batch_size = 32)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T16:11:23.181535Z","iopub.execute_input":"2025-11-05T16:11:23.182135Z","iopub.status.idle":"2025-11-05T16:14:50.616669Z","shell.execute_reply.started":"2025-11-05T16:11:23.182114Z","shell.execute_reply":"2025-11-05T16:14:50.616064Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"history = model.fit(train_ds,epochs=40, batch_size = 64)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T16:17:27.276822Z","iopub.execute_input":"2025-11-05T16:17:27.277617Z","iopub.status.idle":"2025-11-05T16:24:15.878988Z","shell.execute_reply.started":"2025-11-05T16:17:27.277590Z","shell.execute_reply":"2025-11-05T16:24:15.878429Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# get the accuracy and loss scores at each epoch\ntrain_loss = history.history['loss']\ntrain_accuracy = history.history['accuracy']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T16:29:42.957871Z","iopub.execute_input":"2025-11-05T16:29:42.958136Z","iopub.status.idle":"2025-11-05T16:29:42.961952Z","shell.execute_reply.started":"2025-11-05T16:29:42.958120Z","shell.execute_reply":"2025-11-05T16:29:42.961094Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Test the model","metadata":{}},{"cell_type":"code","source":"val_history = model.fit(val_ds,epochs=20, batch_size = 32)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T16:15:18.129010Z","iopub.execute_input":"2025-11-05T16:15:18.129411Z","iopub.status.idle":"2025-11-05T16:16:45.684459Z","shell.execute_reply.started":"2025-11-05T16:15:18.129390Z","shell.execute_reply":"2025-11-05T16:16:45.683703Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"val_history = model.fit(val_ds,epochs=40, batch_size = 64)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T16:24:46.941721Z","iopub.execute_input":"2025-11-05T16:24:46.942001Z","iopub.status.idle":"2025-11-05T16:27:42.735238Z","shell.execute_reply.started":"2025-11-05T16:24:46.941981Z","shell.execute_reply":"2025-11-05T16:27:42.734655Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"val_loss = val_history.history['loss']\nval_accuracy = val_history.history['accuracy'] ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T16:29:28.277667Z","iopub.execute_input":"2025-11-05T16:29:28.278158Z","iopub.status.idle":"2025-11-05T16:29:28.282084Z","shell.execute_reply.started":"2025-11-05T16:29:28.278135Z","shell.execute_reply":"2025-11-05T16:29:28.281303Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## RMSProp Optimizer Accuracy and Loss Plots for Train and Validation Data","metadata":{}},{"cell_type":"code","source":"#epochs_range = range(20) # for first tuning \nepochs_range = range(40) # for second tuning\n\n# Plot for Accuracy - Train and Validation set\nplt.figure(figsize=(12, 4))\nplt.subplot(1, 2, 1)\nplt.plot(epochs_range, train_accuracy, label='Training Accuracy')\nplt.plot(epochs_range, val_accuracy, label='Validation Accuracy')\nplt.legend(loc='lower right')\nplt.title('Training and Validation Accuracy')\n\n# Plot for Loss - Train and Validation set\nplt.subplot(1, 2, 2)\nplt.plot(epochs_range, train_loss, label='Training Loss')\nplt.plot(epochs_range, val_loss, label='Validation Loss')\nplt.legend(loc='upper right')\nplt.title('Training and Validation Loss')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-05T16:30:18.573102Z","iopub.execute_input":"2025-11-05T16:30:18.573792Z","iopub.status.idle":"2025-11-05T16:30:18.869265Z","shell.execute_reply.started":"2025-11-05T16:30:18.573767Z","shell.execute_reply":"2025-11-05T16:30:18.868638Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Results and Analysis","metadata":{}},{"cell_type":"markdown","source":"## SGD Optimizer - Accuracy and Loss Score - at each tuning            ","metadata":{}},{"cell_type":"markdown","source":"Train data is 56K with equal proportions of label 0 and label 1, so the data is balanced. In the same way validation data is 24K, equally labelled and balanced.\n\nKeeping Learning rate 0.0001, L1 and L2 regularizations are 0.001 and 0.0001. Batch size as 32 and number of epochs as 100, accuracy = 0.9131 abd loss  = 0.3028 in train and accuracy = 0.9717 and loss = 0.2109 in validation data.\n\nBatch size as 64 and number of epochs as 150, accuracy = 0.9156 abd loss  = 0.3010 in train and accuracy = 0.9736 and loss = 0.2009 in validation data.\n\nOn increasing the batch size and number of epochs, there is an increase in accuracy and decrease in loss.\n","metadata":{}},{"cell_type":"markdown","source":"## Adam Optimizer - Accuracy and Loss scores","metadata":{}},{"cell_type":"markdown","source":"Keeping Learning rate 0.0001, L1 and L2 regularizations are 0.001 and 0.0001. Batch size as 32 and number of epochs as 50, accuracy = 0.9742 abd loss  = 0.2236 in train and accuracy = 0.9969 and loss = 0.1180 in validation data.\n\nBatch size as 64 and number of epochs as 75, accuracy = 0.9894 abd loss  = 0.1617 in train and accuracy = 0.9972 and loss = 0.1010 in validation data.\n\nOn increasing the batch size and number of epochs, there is an increase in accuracy and decrease in loss.\n","metadata":{}},{"cell_type":"markdown","source":"## RMSProp Optimizer - Accuracy and Loss Score","metadata":{}},{"cell_type":"markdown","source":"Keeping Learning rate 0.0001, L1 and L2 regularizations are 0.001 and 0.0001. Batch size as 32 and number of epochs as 20, accuracy = 0.9791 abd loss  = 0.1586 in train and accuracy = 0.9953 and loss = 0.0916 in validation data.\n\nBatch size as 64 and number of epochs as 75, accuracy = 0.9840 abd loss  = 0.1352 in train and accuracy = 0.9966 and loss = 0.0810 in validation data.\n\nOn increasing the batch size and number of epochs, there is an increase in accuracy and decrease in loss.\n","metadata":{}},{"cell_type":"markdown","source":"## Conclusion","metadata":{}},{"cell_type":"markdown","source":"**Data characteristics:**\n\nThe study uses a well-balanced dataset for both training (56K samples, equal proportions of label 0 and 1) and validation (24K samples, equally labelled), which ensures the performance metrics are not skewed.\n\n**Impact of Hyperparameter Tuning:**\n\nIncreased Batch Size and Epochs Lead to Improved Performance: In all three scenarios presented, increasing the batch size (from 32 to 64) and the number of epochs resulted in a  increase in both training and validation accuracy, alongside a corresponding decrease in loss. \n\n**Optimal Performance Configuration:** \n\nThe final experiment with a batch size of 64 and 75 epochs (train accuracy = 0.9840, validation accuracy = 0.9966) achieved the highest overall validation accuracy and lowest loss, suggesting this combination provided the most effective training criterion for the specific model architecture and the given dataset, regularization, and learning rate parameters.\n\n**Minimal Overfitting Observed:** \n\nThe consistent high validation accuracy (approaching 99.7%) and low loss values across the later experiments indicate the model is generalizing, due to the combined effects of learning rate, L1/L2 regularization, and balanced data.\n","metadata":{}},{"cell_type":"markdown","source":"## Thank you!","metadata":{}}]}