{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":14774,"databundleVersionId":875431,"sourceType":"competition"}],"dockerImageVersionId":30527,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Detect diabetic retinopathy to stop blindness before it's too late","metadata":{}},{"cell_type":"markdown","source":"The objective of this analysis is to detect and prevent diabetic retinopathy, among people living in rural areas.\n\n\n\nGiven with a large set of retina images taken using fundus photography with imaging conditions.\n\nThe images are rated for the severity of diabetic retinopathy on a scale of 0 to 4:\n* 0 - No DR\n* 1 - Mild\n* 2 - Moderate\n* 3 - Severe\n* 4 - Proliferative DR","metadata":{}},{"cell_type":"markdown","source":"## List of Pre-Trained Models Used\n\n1. **ResNet-18** (Residual Network 18), is a specific variant of the ResNet architecture and a popular CNN that allows the training of much deeper networks with improved performance.\n     * 18 - refers to the layers that includes convolutional layers, pooling layers, fully connected layers, and the residual blocks.         \n   \n\n2. **Tf_efficientnet_b4**, is a specific variant of the EfficientNet (CNN)model designed for efficient and scalable image classification.\n     * b4 - the compound scaling factor used to determine the depth, width, and resolution of the network.  \n   \n\n3.  **Tf_efficientnet_lite4**, belongs to EfficientNet Lite models which are designed to be lightweight and efficient while maintaining good performance. In this lite versoin a ReLU6 activation functions and removing squeeze-and-excitation blocks.\n     * lite4 - refers to the version-4 lite model.\n       \n  \n4.  **Efficientnet_b0** - is the smallest and baseline variant of the EfficientNet family of CNN models.\n     * b0 - the compound scaling factor used to determine the depth, width, and resolution of the network.\n      \n      \n5. **Inception-v4** is a Inception architecture variant, designed for image classification and other computer vision tasks.\n      * v4 - refers to the version, 4th version is an update of v3 with inclusion of factorized convolutions, reduction blocks, auxiliary classifiers, inception-ResNet-like Connections, etc. \n      \n    \n6. **Seresnext50_32x4d** is a specific CNN architecture which is a  part of the SE-ResNeXt family. SE-ResNeXt stands for \"Squeeze-and-Excitation ResNeXt,\" which is an extension of the ResNeXt architecture that incorporates a \"squeeze-and-excitation\" module to enhance the learning capabilities of the network.\n  \n    * se    - refers to the usage of Squeeze-and-Excitation (SE) blocks\n    * next  - refers to the \"next\" version of the model architecture\n    * 50    - refers to the depth or number of layers in the network\n    * 32x4d - refers to the cardinality (no.of parallel paths,32) and the width of each path(4).This configuration allows the model to capture a diverse set of features by processing the input data through multiple paths.\n    \n \n7. **Ensemble Models taken for consideration**\n     * Efficientnet_b0\n     * Tf_efficientnet_lite4\n     * Seresnext50_32x4d\n     * Tf_efficientnet_b4\n     \n     \n8. **Augmentation is performed on Tf_efficientnet_lite4**","metadata":{}},{"cell_type":"markdown","source":"### **Experiments with No-Preprocessing**\n\n1. **No Augmentation** is performed for this analysis \n2. Error Metrics **CrossEntropyLoss()**, multiclass (5-features)\n3. Learning Rate --> **0.001**\n4. Epoch  --> **(10-15)** saves the best model","metadata":{}},{"cell_type":"markdown","source":"#### **Experiment 1:** [***resnet18***]\n  \n   \n     \n     \n   * To begin with, this model (relatively simpler and shallow comparatively) is choosen for the image classification model.\n    * *Parameters*\n        * Model Tuning - Softmax layer added afer O/P layer\n        * Image Size - (512,512)\n        * Epoch - 10 (saves the best model)\n        * Optimizer - Adam \n        * Stratified Split\n    * *Cross Validation and Leaderboard Scores*\n        * CV Score - Train(**79.33%**),   Validatian(**76.86%**)\n        * LB Score - Private(**72.33%**), Public(**48.96%**)","metadata":{}},{"cell_type":"markdown","source":"#### **Experiment 2** [***tf_efficientnet_b4***]      \n   * EfficientNet models were proved to perfom well on image classification tasks and hecnce is tried for this problem set.\n    * *Parameters*\n      * Model Tuning - None\n      * Image Size - (256,256)\n      * Epoch - 10 (saves the best model)\n      * Optimizer - AdamW\n    * *Cross Validation and Leaderboard Scores*\n       * CV Score - Train(**91.71%**),   Validatian(**81.36%**)\n       * LB Score - Private(**83.97%**), Public(**66.48%**)","metadata":{}},{"cell_type":"markdown","source":"#### **Experiment 3** [***tf_efficientnet_lite4***]     \n   * This lighter verison of the tf_efficientnet is performed on this dataset, since it is proven to high accuracies .\n    * *Parameters*\n      * Model Tuning - None\n      * Image Size - (512,512)\n      * Epoch - 10 (saves the best model)\n      * Optimizer - Adam\n    * *Cross Validation and Leaderboard Scores*\n       * CV Score - Train(**92.99%**),   Validatian(**82.42%**)\n       * LB Score - Private(**84.61%**), Public(**63.22%**)","metadata":{}},{"cell_type":"markdown","source":"#### **Experiment 4** [***efficientnet_b0***]\n  * Since efficientNet models are known for their excellent trade-off between accuracy and computation, a simpler version is tried in this experiment.\n     * *Parameters*\n        * Model Tuning - None\n        * Image Size - (512,512)\n        * Epoch - 10 (saves the best model)\n        * Optimizer - Adam \n        * Stratified Split\n     * *Cross Validation and Leaderboard Scores*\n        * CV Score - Train(**94.93%**),   Validatian(**82.21%**)\n        * LB Score - Private(**84.17%**), Public(**64.20%**)","metadata":{}},{"cell_type":"markdown","source":"#### **Experiment 5** [***Inception_v4***]\n   * Inception models - designed by Google are characterized by the innovative use of \"inception\" modules, allowing the network to capture features at different scales while optimizing computation efficiency.\n     \n     * *Variant1*\n      * Model Tuning - None\n      * Image Size - (256,256)\n      * Epoch - 20 (saves the best model)\n      * Optimizer - AdamW\n      * Cross Validation and Leaderboard Scores\n        1. CV Score - Train(**91.13%**),   Validatian(**83.49%**)\n        2. LB Score - Private(**80.39%**), Public(**57.94%**)\n    * *Variant2*\n      * Model Tuning - None\n      * Image Size - (512,512)\n      * Epoch - 10 (saves the best model)\n      * Optimizer - Adam\n      * Stratified Split\n      * Cross Validation and Leaderboard Scores\n        1. CV Score - Train(**84.74%**),   Validatian(**79.15%**)\n        2. LB Score - Private(**79.54%**), Public(**56.83%**)","metadata":{}},{"cell_type":"markdown","source":"#### **Experiment 6** [***seresnext50_32x4d***]\n  * Further research and understanding of discussions, the ensembles of seresnext50 model performed well for this dataset and hence is studied for the current analysis.\n    * *Variant1*\n      * Model Tuning - None\n      * Image Size - (512,512)\n      * Epoch - 10 (saves the best model)\n      * Optimizer - AdamW\n      * Cross Validation and Leaderboard Scores\n        1. CV Score - Train(**91.32%**),   Validatian(**84.45%**)\n        2. LB Score - Private(**84.12%**), Public(**53.32%**)\n    * *Variant2*\n      * Model Tuning - None\n      * Image Size - (256,256)\n      * Epoch - 20 (saves the best model)\n      * Optimizer - RMSProp\n      * Cross Validation and Leaderboard Scores\n        1. CV Score - Train(**85.46%**),   Validatian(**79.81%**)\n        2. LB Score - Private(**76.90%**), Public(**59.74%**)\n    * *Variant3*\n      * Model Tuning - None\n      * Image Size - (512,512)\n      * Epoch - 10 (saves the best model)\n      * Optimizer - AdamW\n      * Stratified Split\n      * Cross Validation and Leaderboard Scores\n        1. CV Score - Train(**93.22%**),   Validatian(**82.42%**)\n        2. LB Score - Private(**83.32%**), Public(**61.33%**)","metadata":{}},{"cell_type":"markdown","source":"#### **Experiment 7** [***Ensembling***]\n\n   * ***Ensemble1 (Efficientnet_b0 + Tf_efficientnet_lite4 + Seresnext50_32x4d)***\n      * Model Tuning - None\n      * Image Size - (512,512)\n      * Epoch - 10 (saves the best model)\n      * Optimizer - Adam\n      * Leaderboard Scores\n        * LB Score - Private(**87.80%**), Public(**66.58%**)\n\n\n  * ***Ensemble2 (Efficientnet_b0 + Tf_efficientnet_lite4)***\n      * Model Tuning - None\n      * Image Size - (512,512)\n      * Epoch - 10 (saves the best model)\n      * Optimizer - Adam\n      * Leaderboard Scores\n        * LB Score - Private(**86.00%**), Public(**65.59%**)\n        \n        \n  * ***Ensemble3 (Tf_efficientnet_lite4 + Tf_efficientnet_b4)***\n      * Model Tuning - None\n      * Image Size - (512,512)\n      * Epoch - 10 (saves the best model)\n      * Optimizer - Adam\n      * Leaderboard Scores\n        * LB Score - Private(**82.71%**), Public(**58.69%**)\n        \n        \n   * ***Ensemble4 (Seresnext50_32x4d + Efficientnet_b0)***\n      * Model Tuning - None\n      * Image Size - (512,512)\n      * Epoch - 10 (saves the best model)\n      * Optimizer - Adam\n      * Leaderboard Scores\n        * LB Score - Private(**86.77%**), Public(**65.99%**)","metadata":{}},{"cell_type":"markdown","source":"### **Experiments with Augmentation**\n\n1. **Augmentations - [RandomContrast, RandomBrightness, Blur]** is performed for the analysis \n2. Error Metrics **CrossEntropyLoss()**, multiclass (5-features)\n3. Learning Rate --> **0.001**\n4. Epoch  --> **(10-15)** Saves the best model\n\n\n#### **Experiment 8** [***tf_efficientnet_lite4***]     \n    \n   * *Parameters*\n      * Model Tuning - None\n      * Image Size - (512,512)\n      * Epoch - 10 (saves the best model)\n      * Optimizer - Adam\n    * *Cross Validation and Leaderboard Scores*\n       * CV Score - Train(**90.07%**),   Validatian(**82.31%**)\n       * LB Score - Private(**84.67%**), Public(**64.35%**)","metadata":{}},{"cell_type":"markdown","source":"#### **1. Load the required libraries**","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.image as mpimg\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n#from sklearn.metrics import cohen_kappa_score\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import confusion_matrix\n\nimport cv2 # Computer vision module for image processing\nimport os  # os module with methods for interacting with the operating system\n\nimport timm\nimport torch\nfrom torch import nn\nfrom torch.utils.data import Dataset,DataLoader,random_split\nfrom torchvision import datasets, transforms\n\nfrom PIL import Image\nimport torchvision\nfrom tqdm.notebook import tqdm\n\nseed = 42\ntorch.manual_seed(seed)\nprint(torch.__version__)","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:17:16.910027Z","iopub.execute_input":"2025-10-18T15:17:16.910393Z","iopub.status.idle":"2025-10-18T15:17:21.671173Z","shell.execute_reply.started":"2025-10-18T15:17:16.910363Z","shell.execute_reply":"2025-10-18T15:17:21.670260Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### **2. Setup Device**","metadata":{}},{"cell_type":"code","source":"device = \"cuda\" if torch.cuda.is_available() else \"cpu\"\nprint('Active Device:',device)","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:17:37.708446Z","iopub.execute_input":"2025-10-18T15:17:37.709114Z","iopub.status.idle":"2025-10-18T15:17:37.738276Z","shell.execute_reply.started":"2025-10-18T15:17:37.709082Z","shell.execute_reply":"2025-10-18T15:17:37.737259Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### **3.a. Save the folder paths for training and testing images**","metadata":{}},{"cell_type":"code","source":"train_dir = \"/kaggle/input/aptos2019-blindness-detection/train_images\"\ntest_dir = \"/kaggle/input/aptos2019-blindness-detection/test_images\"\nprint(\"Training Folder path:  \",train_dir)\nprint(\"Testing Folder path:   \",test_dir)","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:17:43.698079Z","iopub.execute_input":"2025-10-18T15:17:43.698678Z","iopub.status.idle":"2025-10-18T15:17:43.703406Z","shell.execute_reply.started":"2025-10-18T15:17:43.698644Z","shell.execute_reply":"2025-10-18T15:17:43.702517Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### **3.b. Open the lables of train and test csv files**","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/aptos2019-blindness-detection/train.csv')\ntest  = pd.read_csv('/kaggle/input/aptos2019-blindness-detection/test.csv')\nprint(train.head())\nprint(test.head())","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:18:00.341912Z","iopub.execute_input":"2025-10-18T15:18:00.342257Z","iopub.status.idle":"2025-10-18T15:18:00.379110Z","shell.execute_reply.started":"2025-10-18T15:18:00.342230Z","shell.execute_reply":"2025-10-18T15:18:00.378238Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.diagnosis.value_counts()","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:18:03.150166Z","iopub.execute_input":"2025-10-18T15:18:03.150496Z","iopub.status.idle":"2025-10-18T15:18:03.163162Z","shell.execute_reply.started":"2025-10-18T15:18:03.150469Z","shell.execute_reply":"2025-10-18T15:18:03.162038Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sns.countplot(x=train['diagnosis'])","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:18:54.994509Z","iopub.execute_input":"2025-10-18T15:18:54.994846Z","iopub.status.idle":"2025-10-18T15:18:55.240797Z","shell.execute_reply.started":"2025-10-18T15:18:54.994811Z","shell.execute_reply":"2025-10-18T15:18:55.239930Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### 4.**Train and Validation Split**","metadata":{}},{"cell_type":"code","source":"# Perform stratified train-test split\ntrain, val = train_test_split(train, test_size=0.25, stratify=train.diagnosis, random_state=seed)","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:19:26.413656Z","iopub.execute_input":"2025-10-18T15:19:26.414041Z","iopub.status.idle":"2025-10-18T15:19:26.423578Z","shell.execute_reply.started":"2025-10-18T15:19:26.414009Z","shell.execute_reply":"2025-10-18T15:19:26.422811Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.diagnosis.value_counts()","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:19:36.194368Z","iopub.execute_input":"2025-10-18T15:19:36.194969Z","iopub.status.idle":"2025-10-18T15:19:36.201772Z","shell.execute_reply.started":"2025-10-18T15:19:36.194937Z","shell.execute_reply":"2025-10-18T15:19:36.200912Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"val.diagnosis.value_counts()","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:19:38.403089Z","iopub.execute_input":"2025-10-18T15:19:38.403421Z","iopub.status.idle":"2025-10-18T15:19:38.410627Z","shell.execute_reply.started":"2025-10-18T15:19:38.403395Z","shell.execute_reply":"2025-10-18T15:19:38.409739Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### **5.a. Checking the size of the images**","metadata":{}},{"cell_type":"code","source":"#dataset_path = '/kaggle/input/aptos2019-blindness-detection/test_images'\n#image_files = os.listdir(dataset_path)\n\n#for image_file in image_files:\n#    image_path = os.path.join(dataset_path, image_file)\n#    image      = cv2.imread(image_path)\n#    if image is not None:\n#        height, width, _ = image.shape\n#        print(f\"Image: {image_file}, Width: {width}, Height: {height}\")","metadata":{"execution":{"iopub.execute_input":"2023-08-23T11:29:56.736141Z","iopub.status.busy":"2023-08-23T11:29:56.735400Z","iopub.status.idle":"2023-08-23T11:29:56.741655Z","shell.execute_reply":"2023-08-23T11:29:56.740497Z","shell.execute_reply.started":"2023-08-23T11:29:56.736096Z"},"scrolled":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The images are of different sizes (640,480) is the majority, with the variations (2588,1958), (819,614), (2416,1736), (1050, 1050) etc","metadata":{}},{"cell_type":"markdown","source":"#### **5.b. Visualizing the actual images**","metadata":{}},{"cell_type":"code","source":"t_dir = \"/kaggle/input/aptos2019-blindness-detection/train_images/\"\nplt.figure(figsize=[15,15])\ni = 1\nfor img_name in train['id_code'][:10]:\n    img = mpimg.imread(t_dir + img_name + '.png')\n    plt.subplot(6,5,i)\n    plt.imshow(img)\n    i += 1\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:19:43.113327Z","iopub.execute_input":"2025-10-18T15:19:43.114070Z","iopub.status.idle":"2025-10-18T15:19:50.183629Z","shell.execute_reply.started":"2025-10-18T15:19:43.114037Z","shell.execute_reply":"2025-10-18T15:19:50.182796Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### **6.a. Resize the image and convert it to tensors for both train and test sets**","metadata":{}},{"cell_type":"code","source":"#train_data = datasets.ImageFolder(root=train_dir, transform=data_transform)\n#test_data  = datasets.ImageFolder(root=test_dir, transform=data_transform)\n# ---> Can't apply this because we dont have folders arranged in label sequence \n# and hence need to customize it\n\nclass Customized_Data(Dataset):\n    def __init__(self, image_folder, csv_file, transform=None):\n        super().__init__()\n        self.image_folder = image_folder\n        self.label_csv = csv_file\n        self.transform = transform\n    \n    def __getitem__(self, idx):\n        img_name = self.label_csv.id_code.values[idx] + '.png'\n        img_path = os.path.join(self.image_folder, img_name)\n        image = cv2.imread(img_path)\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        \n        \n        cropped_image = self.transform(image=image)[\"image\"]\n        label = self.label_csv.diagnosis.values[idx]\n        \n        return cropped_image, label\n        \n\n    def __len__(self):\n        return len(self.label_csv)","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:24:39.819490Z","iopub.execute_input":"2025-10-18T15:24:39.819850Z","iopub.status.idle":"2025-10-18T15:24:39.826580Z","shell.execute_reply.started":"2025-10-18T15:24:39.819822Z","shell.execute_reply":"2025-10-18T15:24:39.825727Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### **6.b. Applying Basic Augmentations**","metadata":{}},{"cell_type":"code","source":"import albumentations as A\nfrom albumentations.pytorch import ToTensorV2\n\n# Define augmentation transformations\ndata_transforms = A.Compose([\n    A.Resize(height=512, width=512),\n    #A.RandomCrop(width=512, height=512),\n    A.RandomBrightnessContrast(p=0.2),\n    #A.HueSaturationValue(hue_shift_limit=10, sat_shift_limit=20, val_shift_limit=20, p=1.0),\n    A.Blur(p=1.0),\n    #A.Rotate(limit=180, p=1.0),\n    #A.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=0, p=1.0),\n    #A.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=0, p=1.0),\n    #A.HorizontalFlip(p=1.0),\n    ToTensorV2(),\n])","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:25:45.745617Z","iopub.execute_input":"2025-10-18T15:25:45.746001Z","iopub.status.idle":"2025-10-18T15:25:45.750946Z","shell.execute_reply.started":"2025-10-18T15:25:45.745974Z","shell.execute_reply":"2025-10-18T15:25:45.749945Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#data_transforms = transforms.Compose([transforms.Resize(size=(512,512)),transforms.ToTensor()])\n\ncustom_train = Customized_Data(image_folder=train_dir, csv_file=train, transform=data_transforms)\ncustom_val   = Customized_Data(image_folder=train_dir, csv_file=val,   transform=data_transforms)","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:25:54.788727Z","iopub.execute_input":"2025-10-18T15:25:54.789086Z","iopub.status.idle":"2025-10-18T15:25:54.793857Z","shell.execute_reply.started":"2025-10-18T15:25:54.789059Z","shell.execute_reply":"2025-10-18T15:25:54.792897Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### **6.c. Visualize the augmented and resized images**","metadata":{}},{"cell_type":"code","source":"def show_images(dataset):\n    loader = torch.utils.data.DataLoader(dataset, batch_size=8, shuffle=True)\n    batch = next(iter(loader))\n    images, labels = batch\n    grid = torchvision.utils.make_grid(images, nrow=4)\n    plt.figure(figsize=(10,10))\n    plt.imshow(np.transpose(grid, (1,2,0)))\n    print('labels : ', labels)\n    \nshow_images(custom_train)","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:26:00.618365Z","iopub.execute_input":"2025-10-18T15:26:00.618721Z","iopub.status.idle":"2025-10-18T15:26:02.384760Z","shell.execute_reply.started":"2025-10-18T15:26:00.618691Z","shell.execute_reply":"2025-10-18T15:26:02.383934Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### 7.**Load Train, Val and Test Datasets into DataLoaders**","metadata":{}},{"cell_type":"code","source":"train_dataloader = DataLoader(dataset=custom_train,batch_size=16, shuffle=True) \nval_dataloader   = DataLoader(dataset=custom_val,  batch_size=16, shuffle=True)","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:26:36.789558Z","iopub.execute_input":"2025-10-18T15:26:36.789910Z","iopub.status.idle":"2025-10-18T15:26:36.794834Z","shell.execute_reply.started":"2025-10-18T15:26:36.789866Z","shell.execute_reply":"2025-10-18T15:26:36.793817Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"image, label = next(iter(train_dataloader))\nprint(f\"Image shape: {image.shape} -> [batch_size, color_channels, height, width]\")\nprint(f\"Label shape: {label.shape}\")","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:26:41.968493Z","iopub.execute_input":"2025-10-18T15:26:41.969070Z","iopub.status.idle":"2025-10-18T15:26:43.565062Z","shell.execute_reply.started":"2025-10-18T15:26:41.969040Z","shell.execute_reply":"2025-10-18T15:26:43.564130Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"image, label = next(iter(val_dataloader))\nprint(f\"Image shape: {image.shape} -> [batch_size, color_channels, height, width]\")\nprint(f\"Label shape: {label.shape}\")","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:26:54.108784Z","iopub.execute_input":"2025-10-18T15:26:54.109633Z","iopub.status.idle":"2025-10-18T15:26:55.674610Z","shell.execute_reply.started":"2025-10-18T15:26:54.109603Z","shell.execute_reply":"2025-10-18T15:26:55.673668Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### **8. Building Model - Pretrained**","metadata":{}},{"cell_type":"code","source":"model1 = timm.list_models('*tf_efficientnet_lite4*')\nprint(model1)","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:26:59.348066Z","iopub.execute_input":"2025-10-18T15:26:59.348362Z","iopub.status.idle":"2025-10-18T15:26:59.354194Z","shell.execute_reply.started":"2025-10-18T15:26:59.348339Z","shell.execute_reply":"2025-10-18T15:26:59.352724Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Creat a model with pretrained (weights as true from Imagenet) and change the number of classes to 5 (0-4)","metadata":{}},{"cell_type":"code","source":"efficientnet = timm.create_model('tf_efficientnet_lite4', pretrained = True, num_classes=5)\n\n# Initializing Error and Optimizer\n\n# Cross Entropy Loss \nerror = nn.CrossEntropyLoss()\n\n# AdamW Optimizer\nlearning_rate = 0.001\noptimizer = torch.optim.Adam(efficientnet.parameters(),lr=learning_rate)","metadata":{"_kg_hide-output":true,"scrolled":true,"trusted":true,"execution":{"iopub.status.busy":"2025-10-18T15:27:04.833244Z","iopub.execute_input":"2025-10-18T15:27:04.833580Z","iopub.status.idle":"2025-10-18T15:27:08.183766Z","shell.execute_reply.started":"2025-10-18T15:27:04.833552Z","shell.execute_reply":"2025-10-18T15:27:08.183090Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Move the model to device(GPU-cuda)","metadata":{}},{"cell_type":"code","source":"model = efficientnet.to(device)\nprint(model.to(device))\ndevice","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:27:23.628918Z","iopub.execute_input":"2025-10-18T15:27:23.629267Z","iopub.status.idle":"2025-10-18T15:27:23.811613Z","shell.execute_reply.started":"2025-10-18T15:27:23.629241Z","shell.execute_reply":"2025-10-18T15:27:23.810739Z"},"scrolled":true,"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Display the Summary**","metadata":{}},{"cell_type":"code","source":"from torchinfo import summary\nsummary(model)","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:29:13.704285Z","iopub.execute_input":"2025-10-18T15:29:13.705010Z","iopub.status.idle":"2025-10-18T15:29:13.742124Z","shell.execute_reply.started":"2025-10-18T15:29:13.704978Z","shell.execute_reply":"2025-10-18T15:29:13.741290Z"},"scrolled":true,"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if torch.cuda.is_available():\n    device = torch.device(\"cuda\") \nelse:\n    device = torch.device(\"cpu\")\nprint(device)","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:30:08.708660Z","iopub.execute_input":"2025-10-18T15:30:08.709336Z","iopub.status.idle":"2025-10-18T15:30:08.714195Z","shell.execute_reply.started":"2025-10-18T15:30:08.709303Z","shell.execute_reply":"2025-10-18T15:30:08.713278Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"##### **8. Training and Validation Loops**","metadata":{}},{"cell_type":"code","source":"def Cmatrix(rater_a, rater_b, min_rating=None, max_rating=None):\n    #assert(len(rater_a) == len(rater_b))\n    if min_rating is None:\n        min_rating = min(rater_a + rater_b)\n    if max_rating is None:\n        max_rating = max(rater_a + rater_b)\n    num_ratings = int(max_rating - min_rating + 1)\n    conf_mat = [[0 for i in range(num_ratings)]\n                for j in range(num_ratings)]\n    for a, b in zip(rater_a, rater_b):\n        conf_mat[a - min_rating][b - min_rating] += 1\n    return conf_mat\n\ndef histogram(ratings, min_rating=None, max_rating=None):\n    if min_rating is None:\n        min_rating = min(ratings)\n    if max_rating is None:\n        max_rating = max(ratings)\n    num_ratings = int(max_rating - min_rating + 1)\n    hist_ratings = [0 for x in range(num_ratings)]\n    for r in ratings:\n        hist_ratings[r - min_rating] += 1\n    return hist_ratings\n\ndef quadratic_weighted_kappa(y, y_pred):\n    rater_a = y\n    rater_b = y_pred\n    min_rating=None\n    max_rating=None\n    rater_a = np.array(rater_a, dtype=int)\n    rater_b = np.array(rater_b, dtype=int)\n    #assert(len(rater_a) == len(rater_b))\n    if min_rating is None:\n        min_rating = min(min(rater_a), min(rater_b))\n    if max_rating is None:\n        max_rating = max(max(rater_a), max(rater_b))\n    conf_mat = Cmatrix(rater_a, rater_b,\n                                min_rating, max_rating)\n    num_ratings = len(conf_mat)\n    num_scored_items = float(len(rater_a))\n\n    hist_rater_a = histogram(rater_a, min_rating, max_rating)\n    hist_rater_b = histogram(rater_b, min_rating, max_rating)\n\n    numerator = 0.0\n    denominator = 0.0\n\n    for i in range(num_ratings):\n        for j in range(num_ratings):\n            expected_count = (hist_rater_a[i] * hist_rater_b[j]\n                              / num_scored_items)\n            d = pow(i - j, 2.0) / pow(num_ratings - 1, 2.0)\n            numerator += d * conf_mat[i][j] / num_scored_items\n            denominator += d * expected_count / num_scored_items\n\n    return (1.0 - numerator / denominator)","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:30:11.419649Z","iopub.execute_input":"2025-10-18T15:30:11.420004Z","iopub.status.idle":"2025-10-18T15:30:11.429630Z","shell.execute_reply.started":"2025-10-18T15:30:11.419973Z","shell.execute_reply":"2025-10-18T15:30:11.428778Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"epochs = 10\noutput1 = []\nlabel1  = []\n\n# Training Loop\nfor epoch in tqdm(range(epochs)):\n    \n    epoch_loss = 0\n    epoch_accuracy = 0\n    best_loss = float('inf')\n    for image, label in train_dataloader:\n        model.train()\n        \n        image  = image.to(device)\n        label  = label.to(device)\n        \n        output = model(image.float())\n        loss = error(output, label)\n        \n        optimizer.zero_grad()\n        loss.backward()\n        optimizer.step()\n        \n        acc = ((output.argmax(dim=1) == label).float().mean())\n        epoch_accuracy += acc/len(train_dataloader)\n        epoch_loss += loss/len(train_dataloader)\n        preds = output.argmax(dim=1)\n        output1.extend(preds.view(-1).detach().to(\"cpu\").numpy())\n        label1.extend(label.view(-1).detach().to(\"cpu\").numpy())\n        #print(output1)\n        #print(label1)\n        #print(quadratic_weighted_kappa(label1,output1))\n    \n    qwk = quadratic_weighted_kappa(label1,output1)\n    print('Epoch : {}, train accuracy : {}, train loss : {}, kappa_value : {}'.format(epoch+1,epoch_accuracy,epoch_loss,qwk))","metadata":{"execution":{"iopub.status.busy":"2025-10-18T15:34:16.909410Z","iopub.execute_input":"2025-10-18T15:34:16.910086Z","execution_failed":"2025-10-18T15:40:36.715Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Validation loop\ntotal_correct = 0\ntotal_samples = 0\nwith torch.no_grad():\n    for image, label in val_dataloader:\n        model.eval()\n        \n        image  = image.to(device)\n        label  = label.to(device)\n        \n        outputs = model(image.float())\n        val_loss = error(outputs, label)\n        \n        predicted = torch.argmax(outputs, 1)\n        total_samples += label.size(0)\n        total_correct += (predicted == label).sum().item()\n        \n\n    # Check for overfitting\n    if val_loss < best_loss:\n        best_val_loss = val_loss\n        torch.save(model.state_dict(), 'best_model_effnetb0.pth')\n    \nvalidation_accuracy = total_correct / total_samples\n\nprint(\"Training Accuracy  : {:.2f}%\" .format(epoch_accuracy * 100))\nprint(\"Validation Accuracy: {:.2f}%\" .format(validation_accuracy * 100))","metadata":{"execution":{"iopub.execute_input":"2023-08-23T12:36:06.880951Z","iopub.status.busy":"2023-08-23T12:36:06.880012Z","iopub.status.idle":"2023-08-23T12:38:00.936902Z","shell.execute_reply":"2023-08-23T12:38:00.931780Z","shell.execute_reply.started":"2023-08-23T12:36:06.880902Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### **9. Inferences**","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nmodels = ['ResNet-18', 'Tf_efficientnet_b4', 'Tf_efficientnet_lite4', 'Efficientnet_b0','Inception-v4-1', \n          'Inception-v4-2', 'Seresnext50_32x4d-1', 'Seresnext50_32x4d-2', 'Seresnext50_32x4d-3','Ensemble1','Ensemble2','Ensemble3','Ensemble4','Tf_efficientnet_lite4-Aug']\n\ntrain_acc     = [0.7933, 0.9171, 0.9362, 0.9493, 0.9113, 0.8474, 0.9132, 0.8546, 0.9322,0.0,0.0,0.0,0.0,0.9007]\nval_acc       = [0.7686, 0.8136, 0.8275, 0.8221, 0.8349, 0.7915, 0.8445, 0.7981, 0.8242,0.8302,0.0,0.0,0.0,0.8231]\nLB_priv_score = [0.7233, 0.8397, 0.8461, 0.8417, 0.8039, 0.7954, 0.8412, 0.7690, 0.8332,0.8780,0.8600,0.8271,0.8677,0.8467]\nLB_pub_score  = [0.4892, 0.6648, 0.6322, 0.6420, 0.5794, 0.5683, 0.5332, 0.5974, 0.6133,0.6658,0.6559,0.5869,0.6599,0.6435]\n\nx = range(len(models))\n\nfig, ax = plt.subplots(figsize=(10, 12))\nbar_width=0.2\nax.barh(x, train_acc, height=0.2, label='Train Acc', color='b')\nax.barh([pos + bar_width for pos in x], val_acc, height=0.2, label='Val Acc', color='g')\nax.barh([pos + 2 * bar_width for pos in x], LB_priv_score, height=0.2, label='Private score', color='r')\nax.barh([pos + 3 * bar_width for pos in x], LB_pub_score, height=0.2, label='Public score', color='y')\n\n# Display values for each bar (100% scale)\ndef add_values(rects, ax):\n    for rect in rects:\n        width = rect.get_width()\n        ax.annotate(f'{width:.2%}',  # Display as percentage with no decimal places\n                    xy=(width, rect.get_y() + rect.get_height() / 2),\n                    xytext=(5, 0),  # 5 points horizontal offset\n                    textcoords='offset points',\n                    va='center')\n\nadd_values(ax.containers[0], ax)\nadd_values(ax.containers[1], ax)\nadd_values(ax.containers[2], ax)\nadd_values(ax.containers[3], ax)\n\n# Set labels and title\nax.set_ylabel('Models')\nax.set_xlabel('Scores')\nax.set_title('Model Performance Comparison')\nax.set_yticks([pos + 1.5 * bar_width for pos in x])\nax.set_yticklabels(models)  # Keep y-axis labels vertical\nax.invert_yaxis()  # Invert y-axis to have the highest score at the top\nax.legend(loc='upper right', bbox_to_anchor=(1, 1))  # Move legend to the top right corner\n\n# Customize x-axis ticks\nax.set_xticks([0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1])\nax.set_xticklabels(['0', '10', '20' ,'30', '40', '50', '60', '70', '80', '90', '100','110'])\n\n# Add a little space between each model's bars\nax.set_xlim(0, 1.2)\n\n# Show the plot\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.execute_input":"2023-08-24T05:59:50.414283Z","iopub.status.busy":"2023-08-24T05:59:50.413868Z","iopub.status.idle":"2023-08-24T05:59:51.303389Z","shell.execute_reply":"2023-08-24T05:59:51.302170Z","shell.execute_reply.started":"2023-08-24T05:59:50.414249Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\n# Comparison of Cross Validation Scores and Leaderboard Score\nComparison_Table = {'Model Name '     : ['ResNet-18', 'Tf_efficientnet_b4', 'Tf_efficientnet_lite4', 'Efficientnet_b0','Inception-v4-1', 'Inception-v4-2', 'Seresnext50_32x4d-1', 'Seresnext50_32x4d-2', 'Seresnext50_32x4d-3','Ensemble1','Ensemble2','Ensemble3','Ensemble4','Tf_efficientnet_lite4-Aug'],\n                    'Train Accuracy'  : [0.7933, 0.9171, 0.9299, 0.9493, 0.9113, 0.8474, 0.9132, 0.8546, 0.9322,'--','--','--','--',0.9007],\n                    'Validation Accuracy'    : [0.7686, 0.8136, 0.8242, 0.8221, 0.8349, 0.7915, 0.8445, 0.7981, 0.8242,'--','--','--','--',0.8231],\n                    'LB Private Score'       : [0.7233, 0.8397, 0.8461, 0.8417, 0.8039, 0.7954, 0.8412, 0.7690, 0.8332,0.8780,0.8600,0.8271,0.8677,0.8467],\n                    'LB Public Score'        : [0.4892, 0.6648, 0.6322, 0.6420, 0.5794, 0.5683, 0.5332, 0.5974, 0.6133,0.6658,0.6559,0.5869,0.6599,0.6435],}\n \n# Create a DataFrame from the dictionary\ndf = pd.DataFrame(Comparison_Table)\n\n# Print the DataFrame\ndf = df.sort_values(by=['LB Private Score'],ascending=False)\ndf = df.reset_index(drop=True)\ndf.index += 1\ndf","metadata":{"execution":{"iopub.execute_input":"2023-08-24T05:58:41.719676Z","iopub.status.busy":"2023-08-24T05:58:41.718918Z","iopub.status.idle":"2023-08-24T05:58:41.744786Z","shell.execute_reply":"2023-08-24T05:58:41.743556Z","shell.execute_reply.started":"2023-08-24T05:58:41.719638Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### CONCLUSION","metadata":{}},{"cell_type":"markdown","source":"From the above list of experiments,\n\n1. **Tf_efficientnet_lite4** model with basic augmentations \n    [**RandomBrightnessContrast(p=0.2)**, and **A.Blur(p=1.0)**] performed well\n    * *With Parameters* - Image Size(512,512) with Adam-Optimizer and cross-entropy loss.\n    * *Cross Validation and Leaderboard Scores*\n       * CV Score - Train(**90.07%**),   Validatian(**82.31%**)\n       * LB Score - Private(**84.67%**), Public(**63.35%**)\n\n\n 2. **Ensemble of (Efficientnet_b0 + Tf_efficientnet_lite4 + Seresnext50_32x4d)**\n    * *With Parameters* - Image Size(512,512) with Adam-Optimizer and cross-entropy loss, No-augmentation.\n    * *Leaderboard Scores*\n       * LB Score - Private(**87.80%**), Public(**66.58%**)","metadata":{}}]}