{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":7428962,"sourceType":"datasetVersion","datasetId":3625455},{"sourceId":6372738,"sourceType":"datasetVersion","datasetId":3671867}],"dockerImageVersionId":25160,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is Instance Segmentation?\n\n<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n\n<p style=\"padding: 10px; color:black; text-align: justify;\">Image classification can be defined as the prediction of a class that the whole image belongs to, fabric classification, for instance. On the other hand, object detection and localisation techniques aim for a more refined classification approach. In these cases, the goal is to discern and locate individual instances within the image, assigning each instance to its respective class <b>(Hafiz and Bhat, 2020)</b>. Taking a step further in classification refinement, the semantic segmentation technique adopts an incremental approach by attributing each pixel in an image to one of the predefined output classes.</p>\n</div>\n\n<center>\n<img src=\"https://i.ibb.co/3y1QksP/Screenshot-2024-01-09-at-10-20-23-AM.png\" alt=\"Instance Segmentation\" border=\"0\" width=900>\n</center>","metadata":{}},{"cell_type":"markdown","source":"# Experimental Set-up for Training\n## 1. Pre-processing the DeepFashion2 Dataset\n\n<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n<p style=\"padding: 10px; color:black; text-align: justify;\">An exploratory analysis of DeepFashion2 indicates that the image dimensions fluctuate vastly. In the dataset, the largest images have dimensions exceeding <b>3000 pixels</b> in height and <b>2000 pixels</b> in width. Conversely, the smallest images have dimensions of around <b>200 pixels</b> for both height and width.</p> \n<p style=\"padding: 10px; color:black; text-align: justify;\">To handle the variation in image sizes, one alternative is to specify a target height and width, along with a resizing method in the Mask R-CNN configuration which will then perform resizing as part of training. However, this approach can result in longer training times and higher computational costs. As an alternative, we resize the entire dataset to a dimension of <b>256 × 256</b> along with their corresponding annotations, as part of data pre-processing. </p>\n<p style=\"padding: 10px; color:black; text-align: justify;\">The Kaggle notebook where I pre-processed and resized DeepFashion2 can be found here - <a href=\"https://www.kaggle.com/code/thusharanair/resizing-deepfashion2-256-x-256-and-basic-eda\">Resizing DeepFashion2 (256 x 256) and Basic EDA</a></p>\n</div>","metadata":{}},{"cell_type":"markdown","source":"## 2. Download Libraries and Pretrained Weights","metadata":{}},{"cell_type":"code","source":"import os\n\n! git clone https://www.github.com/matterport/Mask_RCNN.git\n! pip install -q silence-tensorflow\nos.chdir('Mask_RCNN')\n\n!rm -rf .git # to prevent an error when the kernel is committed\n!rm -rf images assets # to prevent displaying images at the bottom of a kernel","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.status.busy":"2024-01-18T16:01:40.691382Z","iopub.execute_input":"2024-01-18T16:01:40.691921Z","iopub.status.idle":"2024-01-18T16:01:58.964676Z","shell.execute_reply.started":"2024-01-18T16:01:40.691846Z","shell.execute_reply":"2024-01-18T16:01:58.963147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import json\nimport datetime\nimport numpy as np\nimport pandas as pd\nimport skimage\nimport skimage.draw\nfrom skimage.draw import polygon\nimport time\nimport cv2\nimport matplotlib.pyplot as plt\nfrom mrcnn.config import Config\nfrom mrcnn import utils\nimport mrcnn.model as modellib\nfrom mrcnn import visualize\nfrom mrcnn.model import log\nfrom os import listdir\nimport tensorflow as tf\nimport random\n# ignore warnings to make outputs clearer\nimport warnings\nimport skimage\nimport imageio\nimport glob\nimport imgaug\nimport multiprocessing\nimport seaborn as sns\nfrom collections import Counter\nimport gc\nimport sys\nimport json\nimport glob\nimport random\nfrom pathlib import Path\n\nimport cv2\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nfrom mrcnn.model import log\n\nimport itertools\nfrom tqdm import tqdm\n\nfrom imgaug import augmenters as iaa\nfrom sklearn.model_selection import StratifiedKFold, KFold\n\n\nfrom PIL import Image, ImageDraw\nfrom tqdm import tqdm\nfrom sklearn.model_selection import KFold, train_test_split\nfrom PIL import Image, ImageEnhance\nfrom mrcnn import visualize\nimport sys\nfrom IPython.display import FileLink\nfrom imgaug import augmenters as iaa\nfrom sklearn.metrics import precision_score, recall_score, f1_score, confusion_matrix\n\nwarnings.filterwarnings('ignore')\ntqdm.pandas()\n\nDATA_DIR = Path('/kaggle/input/')\nROOT_DIR = Path('/kaggle/working/')\n\nprint(f'Python Version: {sys.version}')\nprint(f'Tensorflow Version: {tf.__version__}')\nprint(f'Tensorflow Keras Version: {tf.keras.__version__}')","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:01:58.966767Z","iopub.execute_input":"2024-01-18T16:01:58.967207Z","iopub.status.idle":"2024-01-18T16:01:58.980841Z","shell.execute_reply.started":"2024-01-18T16:01:58.967126Z","shell.execute_reply":"2024-01-18T16:01:58.979980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sys.path.append(ROOT_DIR/'Mask_RCNN')\nfrom mrcnn.config import Config\nfrom mrcnn import utils\nimport mrcnn.model as modellib\nfrom mrcnn import visualize\nfrom mrcnn.model import log","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","execution":{"iopub.status.busy":"2024-01-18T16:01:58.982143Z","iopub.execute_input":"2024-01-18T16:01:58.982673Z","iopub.status.idle":"2024-01-18T16:01:59.005727Z","shell.execute_reply.started":"2024-01-18T16:01:58.982608Z","shell.execute_reply":"2024-01-18T16:01:59.004587Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!wget --quiet https://github.com/matterport/Mask_RCNN/releases/download/v2.0/mask_rcnn_coco.h5\n!ls -lh mask_rcnn_coco.h5\n\nCOCO_WEIGHTS_PATH = 'mask_rcnn_coco.h5'","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:01:59.007781Z","iopub.execute_input":"2024-01-18T16:01:59.008152Z","iopub.status.idle":"2024-01-18T16:02:04.907046Z","shell.execute_reply.started":"2024-01-18T16:01:59.008101Z","shell.execute_reply":"2024-01-18T16:02:04.906062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Dataframe pre-processing","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/deepfashion2-256x256/DeepFashion2 Resized/input/train.csv')\nvalidation_df = pd.read_csv('/kaggle/input/deepfashion2-256x256/DeepFashion2 Resized/input/validation.csv')","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:04:19.274728Z","iopub.execute_input":"2024-01-18T16:04:19.275098Z","iopub.status.idle":"2024-01-18T16:04:28.193352Z","shell.execute_reply.started":"2024-01-18T16:04:19.275047Z","shell.execute_reply":"2024-01-18T16:04:28.192222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Minor pre-processing\ntrain_df['path'] = train_df['path'].str.replace('working', 'input/deepfashion2-256x256/DeepFashion2 Resized')\nvalidation_df['path'] = validation_df['path'].str.replace('working', 'input/deepfashion2-256x256/DeepFashion2 Resized')","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:04:28.194976Z","iopub.execute_input":"2024-01-18T16:04:28.195255Z","iopub.status.idle":"2024-01-18T16:04:28.599979Z","shell.execute_reply.started":"2024-01-18T16:04:28.195206Z","shell.execute_reply":"2024-01-18T16:04:28.599153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"validation_df['segmentation'] = validation_df['segmentation'].progress_apply(lambda x: json.loads(x.replace(\"'\", \"\\\"\")))\ntrain_df['segmentation'] = train_df['segmentation'].progress_apply(lambda x: json.loads(x.replace(\"'\", \"\\\"\")))\n\nvalidation_df['b_box'] = validation_df['b_box'].progress_apply(lambda x: json.loads(x.replace(\"'\", \"\\\"\")))\ntrain_df['b_box'] = train_df['b_box'].progress_apply(lambda x: json.loads(x.replace(\"'\", \"\\\"\")))","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:04:28.601654Z","iopub.execute_input":"2024-01-18T16:04:28.602253Z","iopub.status.idle":"2024-01-18T16:04:51.264316Z","shell.execute_reply.started":"2024-01-18T16:04:28.602185Z","shell.execute_reply":"2024-01-18T16:04:51.263482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_df.tail())\nprint(\"\\n\\n\")\ndisplay(validation_df.tail())","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:05:01.613201Z","iopub.execute_input":"2024-01-18T16:05:01.613555Z","iopub.status.idle":"2024-01-18T16:05:01.697617Z","shell.execute_reply.started":"2024-01-18T16:05:01.613501Z","shell.execute_reply":"2024-01-18T16:05:01.696580Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Sampling the dataset\n\n<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n\n<p style=\"padding: 10px; color:black; text-align: justify;\"> DeepFashion2 is a large dataset, containing over <b>491k images</b> with a total of 801k items <b>(Ge et al., 2019)</b>. Due to constraints on time and computational resources, working with the full dataset can be impractical. One solution is to create a subset of the training and validation sets while keeping the original class balance. However, care must be taken to ensure that when an image is sampled, the information for all garment instances within the image is retained.  </p>\n</div>","metadata":{}},{"cell_type":"code","source":"def balanced_sampling(df, n_samples_per_category):\n    # Store which images (by path) have been sampled\n    sampled_images = set()\n    \n    # Output dataframe to store sampled rows\n    sampled_df = pd.DataFrame()\n\n    # Loop over categories\n    for category in df['category_name'].unique():\n        # Find unique images containing this category\n        category_images = set(df[df['category_name'] == category]['path'].unique())\n        \n        # Exclude already sampled images to ensure new samples\n        available_images = category_images - sampled_images\n        if len(available_images) < n_samples_per_category:\n            selected_images = available_images\n        else:\n            selected_images = set(np.random.choice(list(available_images), size=n_samples_per_category, replace=False))\n        \n        # Add selected images to the sampled_images set\n        sampled_images = sampled_images.union(selected_images)\n\n        # Gather all rows related to the selected images and append to sampled_df\n        sampled_df = pd.concat([sampled_df, df[df['path'].isin(selected_images)]], axis=0)\n\n    return sampled_df.sample(frac=1).reset_index(drop=True)  # shuffle and return","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:05:05.730167Z","iopub.execute_input":"2024-01-18T16:05:05.730515Z","iopub.status.idle":"2024-01-18T16:05:05.737819Z","shell.execute_reply.started":"2024-01-18T16:05:05.730467Z","shell.execute_reply":"2024-01-18T16:05:05.736839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = balanced_sampling(train_df, 1300) #1320\nvalidation_df = balanced_sampling(validation_df, 1500) #400","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:05:06.313918Z","iopub.execute_input":"2024-01-18T16:05:06.314562Z","iopub.status.idle":"2024-01-18T16:05:09.156570Z","shell.execute_reply.started":"2024-01-18T16:05:06.314256Z","shell.execute_reply":"2024-01-18T16:05:09.155656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grouped_train = train_df.groupby('path')\ngrouped_validation = validation_df.groupby('path')\n\nagg_funcs = {\n    'b_box': list,\n    'category_name': list,\n    'segmentation': list,\n    'category_id': list,\n    'img_height': 'first',\n    'img_width': 'first'\n}\n\ntrain_df = grouped_train.agg(agg_funcs).reset_index()\nvalidation_df = grouped_validation.agg(agg_funcs).reset_index()","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:05:12.019147Z","iopub.execute_input":"2024-01-18T16:05:12.019546Z","iopub.status.idle":"2024-01-18T16:05:24.387474Z","shell.execute_reply.started":"2024-01-18T16:05:12.019468Z","shell.execute_reply":"2024-01-18T16:05:24.386420Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_df.tail())\nprint(\"\\n\\n\")\ndisplay(validation_df.tail())\nprint(\"\\n\\n\")","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:05:24.388830Z","iopub.execute_input":"2024-01-18T16:05:24.389133Z","iopub.status.idle":"2024-01-18T16:05:24.472000Z","shell.execute_reply.started":"2024-01-18T16:05:24.389076Z","shell.execute_reply":"2024-01-18T16:05:24.470888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n\n<p style=\"padding: 10px; color:black; text-align: justify;\"> We analyse the distributions of different <b>garment category combinations</b> in the dataset, such as the number of images containing both short sleeve tops and skirts. The dataset was sampled based on these category combinations rather than individual categories. This approach ensures that if an image is included in the sample, the information for all garment instances within the image is also retained. </p>\n</div>","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(15, 45), sharey=True)\n\ncombinations_df_train = pd.DataFrame()\ncombinations_df_validation = pd.DataFrame()\n\n# Count the occurrences of each category combination\ncombinations_df_train['category_name'] = train_df['category_name']\ncombinations_df_train['category_combination'] = combinations_df_train['category_name'].apply(lambda x: ', '.join(sorted(x)))\ncounter = Counter(combinations_df_train['category_combination'])\ncombinations_df_train = combinations_df_train.join(pd.DataFrame(list(counter.items()), columns=['Category Combination', 'Count']))\nsns.barplot(data=combinations_df_train, y='Category Combination', x='Count', ax=axes[0]).set_title('Category pair distribution in Training Set')\n\ncombinations_df_validation['category_name'] = validation_df['category_name']\ncombinations_df_validation['category_combination'] = combinations_df_validation['category_name'].apply(lambda x: ', '.join(sorted(x)))\ncounter = Counter(combinations_df_validation['category_combination'])\ncombinations_df_validation = combinations_df_validation.join(pd.DataFrame(list(counter.items()), columns=['Category Combination', 'Count']))\nsns.barplot(data=combinations_df_validation, y='Category Combination', x='Count', ax=axes[1]).set_title('Category pair distribution in Validation Set')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:05:39.919888Z","iopub.execute_input":"2024-01-18T16:05:39.920235Z","iopub.status.idle":"2024-01-18T16:05:47.568217Z","shell.execute_reply.started":"2024-01-18T16:05:39.920181Z","shell.execute_reply":"2024-01-18T16:05:47.567201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"validation_df, test_df = train_test_split(validation_df, test_size=0.50, random_state=42) ","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:05:47.569510Z","iopub.execute_input":"2024-01-18T16:05:47.569765Z","iopub.status.idle":"2024-01-18T16:05:47.660990Z","shell.execute_reply.started":"2024-01-18T16:05:47.569721Z","shell.execute_reply":"2024-01-18T16:05:47.660005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 3, figsize=(18, 6), sharey=True)\n\ncategories = [category for sublist in train_df['category_name'] for category in sublist]\ncounter = Counter(categories)\ncategories = pd.DataFrame(list(counter.items()), columns=['Category', 'Count'])\nsns.barplot(data=categories, x='Category', y='Count', ax=axes[0], order=categories.sort_values('Count',ascending = False).Category).set_title('Category distribution in Training Set')\naxes[0].tick_params('x', labelrotation=90)\n\ncategories = [category for sublist in validation_df['category_name'] for category in sublist]\ncounter = Counter(categories)\ncategories = pd.DataFrame(list(counter.items()), columns=['Category', 'Count'])\nsns.barplot(data=categories, x='Category', y='Count', ax=axes[1], order=categories.sort_values('Count',ascending = False).Category).set_title('Category distribution in Validation Set')\naxes[1].tick_params('x', labelrotation=90)\n\ncategories = [category for sublist in test_df['category_name'] for category in sublist]\ncounter = Counter(categories)\ncategories = pd.DataFrame(list(counter.items()), columns=['Category', 'Count'])\nsns.barplot(data=categories, x='Category', y='Count', ax=axes[2], order=categories.sort_values('Count',ascending = False).Category).set_title('Category distribution in Test Set')\naxes[2].tick_params('x', labelrotation=90)\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:05:47.669058Z","iopub.execute_input":"2024-01-18T16:05:47.669408Z","iopub.status.idle":"2024-01-18T16:05:48.620566Z","shell.execute_reply.started":"2024-01-18T16:05:47.669342Z","shell.execute_reply":"2024-01-18T16:05:48.619737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n<p style=\"padding: 10px; color:black; text-align: justify;\"> To address the class imbalance observed in the original dataset, we strategically refined our sampling method to ensure an approximately balanced representation of each garment category. With this approach, we generated a training dataset of <b>16k</b> distinct images and a validation set of <b>8k</b> images. An inspection of the categories within these sets reveals a more even distribution of garment classes. </p>\n</div>","metadata":{}},{"cell_type":"code","source":"if len(validation_df)%2 != 0:\n    random_index = validation_df.sample(n=1).index\n    validation_df = validation_df.drop(random_index)\n    \n    \nif len(train_df)%2 != 0:\n    random_index = train_df.sample(n=1).index\n    train_df = train_df.drop(random_index)\n    \nif len(test_df)%2 != 0:\n    random_index = test_df.sample(n=1).index\n    test_df = test_df.drop(random_index)\n    \n    \n    \nprint(len(validation_df))\nprint(len(train_df))\nprint(len(test_df))","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:05:58.761688Z","iopub.execute_input":"2024-01-18T16:05:58.762039Z","iopub.status.idle":"2024-01-18T16:05:58.790023Z","shell.execute_reply.started":"2024-01-18T16:05:58.761985Z","shell.execute_reply":"2024-01-18T16:05:58.789227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_instances_train = train_df['category_id'].apply(len).tolist()\nnum_instances_validation = validation_df['category_id'].apply(len).tolist()\nnum_instances_test = test_df['category_id'].apply(len).tolist()\n\nfig, axs = plt.subplots(1, 3, figsize=(20, 5), sharey=True, tight_layout=True)\n\nsns.countplot(x=num_instances_train, ax=axs[0]).set_title('Training Dataset')\naxs[0].set_xlabel('Number of Instances')\n\nsns.countplot(x=num_instances_validation, ax=axs[1]).set_title('Validation Dataset')\naxs[1].set_xlabel('Number of Instances')\n\nsns.countplot(x=num_instances_test, ax=axs[2]).set_title('Test Dataset')\naxs[2].set_xlabel('Number of Instances')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:05:59.598996Z","iopub.execute_input":"2024-01-18T16:05:59.599540Z","iopub.status.idle":"2024-01-18T16:06:00.808725Z","shell.execute_reply.started":"2024-01-18T16:05:59.599470Z","shell.execute_reply":"2024-01-18T16:06:00.807610Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5. Create Dataset & Config\n\n<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n<p style=\"padding: 10px; color:black; text-align: justify;\"> \nWe use an open-source variant <b>(Wijkhuizen, 2021)</b> of Matterplot’s Mask R-CNN implementation <b>(Abdulla, 2017)</b> as well as the original implementation for the experiment. This variant allows the use of any model from the EfficientNetV2 family as the backbone <b>(Wijkhuizen, 2021)</b>, in contrast to the original implementation, which uses a ResNet backbone <b>(Abdulla, 2017)</b>. Several default settings of Mask R-CNN were modified to better fit our dataset and the available computational resources.\n</p>   \n</div>\n\n<center><a href=\"https://ibb.co/xG7qfSD\"><img src=\"https://i.ibb.co/tcJpxM4/Screenshot-2024-01-09-at-10-17-27-AM.png\" alt=\"Model Architecture\" border=\"0\" width=750></a></center>","metadata":{}},{"cell_type":"code","source":"# Target Image Dimensions which are divisable by 64 as required by the MASK-RCNN model\nHEIGHT_TARGET = 256\nWIDTH_TARGET = 256\nSHAPE_TARGET = (HEIGHT_TARGET, WIDTH_TARGET)","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:06:03.399248Z","iopub.execute_input":"2024-01-18T16:06:03.399824Z","iopub.status.idle":"2024-01-18T16:06:03.404332Z","shell.execute_reply.started":"2024-01-18T16:06:03.399748Z","shell.execute_reply":"2024-01-18T16:06:03.403067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n<p style=\"padding: 10px; color:black; text-align: justify;\"> \nA <i>Dataset</i> class is configured which transforms the tabular data created in the earlier phase into variables accessible by the network functions. A custom configuration sets up the hyperparameters and other options for training. For instance, some hyperparameters control the number of instances identified in an image during training and another set of hyperparameters define the size and shape of the instance masks detected by the model.\n</p>\n</div>","metadata":{}},{"cell_type":"code","source":"class ClothDataset(utils.Dataset):\n\n    def __init__(self, df):\n        super().__init__()\n        self.df = df\n\n    def load_dataset(self):\n        categories = [\"short sleeve top\", \"long sleeve top\", \"short sleeve outwear\", \"long sleeve outwear\", \"vest\", \"sling\", \n                      \"shorts\", \"trousers\", \"skirt\", \"short sleeve dress\", \"long sleeve dress\", \"vest dress\", \"sling dress\"]\n        for index, category in enumerate(categories):\n            self.add_class(\"fashion\", index+1, category.lower())\n            \n        for vertical_flip in [True, False]:\n            for horizontal_flip in [True, False]:\n                for index, row in self.df.iterrows():\n                    self.add_image('fashion', \n                                   image_id=index, \n                                   path=row['path'], \n                                   bounding_box=row['b_box'], \n                                   segmentation=row['segmentation'],\n                                   category_name=row['category_name'], \n                                   category_id=row['category_id'],\n                                   img_height=row['img_height'],\n                                   img_width=row['img_width'],\n                                   vertical_flip=vertical_flip, \n                                   horizontal_flip=horizontal_flip)\n            \n    def extract_boxes(self, image_id):\n        image_info = self.image_info[image_id]\n        boxes = np.array(image_info['bounding_box'])  # Convert to numpy array\n        category_names = image_info['category_name']\n        return boxes, category_names, image_info['img_width'], image_info['img_height']\n\n    def load_mask(self, image_id):\n        info = self.image_info[image_id]\n    \n        img_width = info['img_width']\n        img_height = info['img_height']\n\n        mask_list = info['segmentation']\n\n        # Initialize an empty mask with all zeros\n        masks = np.zeros([img_height, img_width, len(mask_list)], dtype='uint8')\n\n        # For each mask in the mask_list\n        for i, masks_per_instance in enumerate(mask_list):\n            # Masks for each instance could be multiple in case an instance has disjoint parts; combine them into a single mask here.\n            full_mask = np.zeros([img_height, img_width], dtype='uint8')\n            for seg_mask in masks_per_instance:\n                rr, cc = polygon(seg_mask[1::2], seg_mask[0::2], (img_height, img_width))\n                full_mask[rr, cc] = 1\n            masks[:, :, i] = full_mask\n\n        # Convert the list of category IDs into an array\n        class_ids = np.array(info['category_id'], dtype='int32')\n        return masks, class_ids\n\n    def image_reference(self, image_id):\n        info = self.image_info[image_id]\n        return info['path']","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:06:05.835916Z","iopub.execute_input":"2024-01-18T16:06:05.836746Z","iopub.status.idle":"2024-01-18T16:06:05.854770Z","shell.execute_reply.started":"2024-01-18T16:06:05.836668Z","shell.execute_reply":"2024-01-18T16:06:05.853763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset = ClothDataset(train_df)\ndataset.load_dataset()\ndataset.prepare()\n\nplt.figure(figsize=(25, 80))\nfor i in range(5):\n    ax = plt.subplot(1, 5, i+1)\n    image_id = random.choice(dataset.image_ids)\n    image = dataset.load_image(image_id)\n    mask, class_ids = dataset.load_mask(image_id)\n    bbox = utils.extract_bboxes(mask)\n    print(dataset.image_reference(image_id))\n    visualize.display_instances(image, bbox, mask, class_ids, dataset.class_names, ax=ax)\n\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:06:07.467302Z","iopub.execute_input":"2024-01-18T16:06:07.467636Z","iopub.status.idle":"2024-01-18T16:06:21.418539Z","shell.execute_reply.started":"2024-01-18T16:06:07.467579Z","shell.execute_reply":"2024-01-18T16:06:21.417676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATASET_LENGTH = len(train_df)\nMAX_BATCH = 2\nBATCH_SIZE = sorted([int(DATASET_LENGTH/n) for n in range(1,DATASET_LENGTH+1) if DATASET_LENGTH % n ==0 and DATASET_LENGTH/n<=MAX_BATCH],reverse=True)[0]  \nSTEPS = int(DATASET_LENGTH/BATCH_SIZE)\nSTEPS","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:06:35.072175Z","iopub.execute_input":"2024-01-18T16:06:35.072525Z","iopub.status.idle":"2024-01-18T16:06:35.081639Z","shell.execute_reply.started":"2024-01-18T16:06:35.072472Z","shell.execute_reply":"2024-01-18T16:06:35.080675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"VAL_DATASET_LENGTH = len(validation_df)\nMAX_BATCH = 2\nVAL_BATCH_SIZE = sorted([int(VAL_DATASET_LENGTH/n) for n in range(1,VAL_DATASET_LENGTH+1) if VAL_DATASET_LENGTH % n ==0 and VAL_DATASET_LENGTH/n<=MAX_BATCH],reverse=True)[0]  \nVAL_STEPS = int(VAL_DATASET_LENGTH/VAL_BATCH_SIZE)\nVAL_STEPS","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:06:35.463833Z","iopub.execute_input":"2024-01-18T16:06:35.464152Z","iopub.status.idle":"2024-01-18T16:06:35.472107Z","shell.execute_reply.started":"2024-01-18T16:06:35.464101Z","shell.execute_reply":"2024-01-18T16:06:35.471436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class ClothConfig(Config):\n    # name of the configuration\n    NAME = \"fashion_config\"\n    \n    GPU_COUNT = 1\n    IMAGES_PER_GPU = 2\n    STEPS_PER_EPOCH = STEPS\n    VALIDATION_STEPS = VAL_STEPS\n    \n    # Number of Classes\n    NUM_CLASSES = 1 + 13\n\n    # Image Dimensions\n    IMAGE_MIN_DIM = HEIGHT_TARGET\n    IMAGE_MAX_DIM = WIDTH_TARGET\n    IMAGE_SHAPE = [HEIGHT_TARGET, WIDTH_TARGET, 3]\n    IMAGE_RESIZE_MODE = 'square'\n    BACKBONE = 'resnet50'\n    \n    TRAIN_BN = False\n    \n    # Learning Rate\n    LEARNING_RATE = 0.0001\n    WEIGHT_DECAY = 0.0\n    \n    # Dataloader Queue Size (was set to 100 but resulted in OOM error)\n    MAX_QUEUE_SIZE = 3\n    \n    # Debug mode will disable model checkpoints\n    DEBUG = False\n    \n    # Do not use multithreading as this slows down the dataloader!\n    WORKERS = 0\n    \n    # Losses\n    LOSS_WEIGHTS = {\n        'rpn_class_loss': 1.0,    # is the class of the bbox correct? / RPN anchor classifier loss (Forground/Background)\n        'rpn_bbox_loss': 1.0,     # is the size of the bbox correct? / RPN bounding box loss graph (bbox of generic object)\n        'mrcnn_class_loss': 1.0,  # loss for the classifier head of Mask R-CNN (Background / specific class)\n        'mrcnn_bbox_loss': 1.0,   # is the size of the bounding box correct or not? / loss for Mask R-CNN bounding box refinement\n        'mrcnn_mask_loss': 1.0,   # is the class correct? is the pixel correctly assign to the class? / mask binary cross-entropy loss for the masks head\n    }\n    \n    # Training Structure\n    RPN_ANCHOR_SCALES = (16, 32, 64, 128, 256)\n    \n    \n    # Regions of Interest\n    PRE_NMS_LIMIT = 2000\n    \n    # Non Max Supression\n    POST_NMS_ROIS_TRAINING = 600\n    POST_NMS_ROIS_INFERENCE = 600\n    \n    # Instances\n    MAX_GT_INSTANCES = 2\n    TRAIN_ROIS_PER_IMAGE = 100\n    DETECTION_MAX_INSTANCES = 3\n    \n    # Thresholds\n    RPN_NMS_THRESHOLD = 0.70        # IoU Threshold for RPN proposals and GT\n    DETECTION_MIN_CONFIDENCE = 0.50 # Non-Background Confidence Threshold\n    DETECTION_NMS_THRESHOLD = 0.30  # IoU Threshold for ROI and GT\n    ROI_POSITIVE_RATIO = 0.33\n    \n    \nconfig = ClothConfig()\nconfig.display()","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:06:37.007649Z","iopub.execute_input":"2024-01-18T16:06:37.008217Z","iopub.status.idle":"2024-01-18T16:06:37.022313Z","shell.execute_reply.started":"2024-01-18T16:06:37.008142Z","shell.execute_reply":"2024-01-18T16:06:37.021007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training the model on COCO weights\n\n<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n<p style=\"padding: 10px; color:black; text-align: justify;\"> \nThe <b>ResNet50</b> backbone model is initialised with <b>COCO</b> weights. During training and validation, we begin by training only the top layers of the model, using a learning rate of <b>0.0002</b>. Subsequently, all layers are trained with a reduced learning rate of <b>0.0001</b>, but with a momentum of 0.9. This step is repeated with an even smaller learning rate of <b>0.00002</b>. To introduce variety and enhance the generalising capabilities of the model, in-place data augmentation is employed. The input images are flipped horizontally with a <b>50%</b> probability. This augmentation technique leads to about half of the images in the batch being flipped horizontally, while the others are left as is.\n</p>\n</div>","metadata":{}},{"cell_type":"code","source":"# training set\ntraining_data = ClothDataset(train_df)\ntraining_data.load_dataset()\ntraining_data.prepare()\n\n# validation set\nvalidation_data = ClothDataset(validation_df)\nvalidation_data.load_dataset()\nvalidation_data.prepare()\n\n# load fashion config\nconfig = ClothConfig()","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:06:39.714809Z","iopub.execute_input":"2024-01-18T16:06:39.715152Z","iopub.status.idle":"2024-01-18T16:06:56.289561Z","shell.execute_reply.started":"2024-01-18T16:06:39.715088Z","shell.execute_reply":"2024-01-18T16:06:56.288623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = modellib.MaskRCNN(mode='training', config=config, model_dir=ROOT_DIR)\n\nmodel.load_weights(COCO_WEIGHTS_PATH, by_name=True, exclude=[\n    'mrcnn_class_logits', 'mrcnn_bbox_fc', 'mrcnn_bbox', 'mrcnn_mask'])","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:06:56.290964Z","iopub.execute_input":"2024-01-18T16:06:56.291452Z","iopub.status.idle":"2024-01-18T16:07:05.121923Z","shell.execute_reply.started":"2024-01-18T16:06:56.291259Z","shell.execute_reply":"2024-01-18T16:07:05.120746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"augmentation = iaa.Sequential([\n    iaa.Fliplr(0.5) # only horizontal flip here\n])","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:07:05.123517Z","iopub.execute_input":"2024-01-18T16:07:05.123865Z","iopub.status.idle":"2024-01-18T16:07:05.130380Z","shell.execute_reply.started":"2024-01-18T16:07:05.123800Z","shell.execute_reply":"2024-01-18T16:07:05.129295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n<p style=\"padding: 10px; color:black; text-align: justify;\"> \nFor garment instance segmentation, the model is trained utilising a single NVIDIA TESLA P100 Graphics Processing Unit (GPU). Due to resource limitations, the training is capped at <b>12 epochs</b>, selecting the epoch with the smallest validation loss for the network weights to optimise the model performance.\n</p>\n</div>","metadata":{}},{"cell_type":"markdown","source":"## 1. Training only the heads","metadata":{}},{"cell_type":"code","source":"start_train = time.time()\nmodel.train(train_dataset=training_data, \n            val_dataset=validation_data, \n            learning_rate=config.LEARNING_RATE*2, \n            layers='heads',\n            epochs=3,\n            augmentation=None)\nend_train = time.time()\nminutes = round((end_train - start_train) / 60, 2)\nprint(f'Training heads took {minutes} minutes')\n\nhistory = model.keras_model.history.history","metadata":{"execution":{"iopub.status.busy":"2024-01-09T04:11:12.158885Z","iopub.execute_input":"2024-01-09T04:11:12.159185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Training all layers Pt.1","metadata":{}},{"cell_type":"code","source":"start_train = time.time()\nmodel.train(train_dataset=training_data, \n            val_dataset=validation_data, \n            learning_rate=config.LEARNING_RATE, \n            layers='all',\n            epochs=9,\n            augmentation=augmentation)\nend_train = time.time()\nminutes = round((end_train - start_train) / 60, 2)\nprint(f'Training all took {minutes} minutes')\n\nnew_history = model.keras_model.history.history\nfor k in new_history: \n    history[k] = history[k] + new_history[k]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Train all layers Pt.2","metadata":{}},{"cell_type":"code","source":"start_train = time.time()\nmodel.train(train_dataset=training_data, \n            val_dataset=validation_data, \n            learning_rate=config.LEARNING_RATE/5, \n            layers='all',\n            epochs=12,\n            augmentation=augmentation)\nend_train = time.time()\nminutes = round((end_train - start_train) / 60, 2)\nprint(f'Training all took {minutes} minutes')\n\nnew_history = model.keras_model.history.history\nfor k in new_history: \n    history[k] = history[k] + new_history[k]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"np.save('/kaggle/working/history_resnet50.npy', history)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Visualising the trainign and validation metrics","metadata":{}},{"cell_type":"code","source":"history = np.load(\"/kaggle/input/resnet50-training/history_resnet50.npy\", allow_pickle=True).item()","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:07:11.969675Z","iopub.execute_input":"2024-01-18T16:07:11.970299Z","iopub.status.idle":"2024-01-18T16:07:11.981433Z","shell.execute_reply.started":"2024-01-18T16:07:11.970044Z","shell.execute_reply":"2024-01-18T16:07:11.980188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function to plot metrics\ndef plot_history_metric(ax, metric, f_best=np.argmax):\n    values = history[metric]\n    N_EPOCHS = len(values)\n    if N_EPOCHS <= 20:\n        x = np.arange(1, N_EPOCHS + 1)\n    else:\n        x = [1, 5] + [10 + 5 * idx for idx in range((N_EPOCHS - 10) // 5 + 1)]\n    x_ticks = np.arange(1, N_EPOCHS+1)\n    \n    ax.plot(x_ticks, values, label='train')\n    argmin = f_best(values)\n    ax.scatter(argmin + 1, values[argmin], color='red', s=75, marker='o', label='train_best')\n    ax.set_ylabel(metric, fontsize=15, labelpad=10)\n    ax.set_xlabel('epoch', fontsize=15, labelpad=10)\n    ax.tick_params(axis='x', labelsize=12)\n    ax.tick_params(axis='y', labelsize=12)\n    ax.set_xticks(x) # set tick step to 1 and let x axis start at 1\n    ax.legend(prop={'size': 15})","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:07:13.839461Z","iopub.execute_input":"2024-01-18T16:07:13.839773Z","iopub.status.idle":"2024-01-18T16:07:13.847607Z","shell.execute_reply.started":"2024-01-18T16:07:13.839725Z","shell.execute_reply":"2024-01-18T16:07:13.846498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_epoch = np.argmin(history[\"val_loss\"]) + 1\nprint(\"Best epoch: \", best_epoch)\nprint(\"Valid loss: \", history[\"val_loss\"][best_epoch-1])","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:07:14.274919Z","iopub.execute_input":"2024-01-18T16:07:14.275434Z","iopub.status.idle":"2024-01-18T16:07:14.283033Z","shell.execute_reply.started":"2024-01-18T16:07:14.275297Z","shell.execute_reply":"2024-01-18T16:07:14.281657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axs = plt.subplots(2, 3, figsize=(22,18))\nplot_history_metric(axs[0, 0], 'loss', f_best=np.argmin)\nplot_history_metric(axs[0, 1], 'rpn_class_loss', f_best=np.argmin)\nplot_history_metric(axs[0, 2], 'rpn_bbox_loss', f_best=np.argmin)\nplot_history_metric(axs[1, 0], 'mrcnn_class_loss', f_best=np.argmin)\nplot_history_metric(axs[1, 1], 'mrcnn_bbox_loss', f_best=np.argmin)\nplot_history_metric(axs[1, 2], 'mrcnn_mask_loss', f_best=np.argmin)\n\nplot_history_metric(axs[0, 0], 'val_loss', f_best=np.argmin)\nplot_history_metric(axs[0, 1], 'val_rpn_class_loss', f_best=np.argmin)\nplot_history_metric(axs[0, 2], 'val_rpn_bbox_loss', f_best=np.argmin)\nplot_history_metric(axs[1, 0], 'val_mrcnn_class_loss', f_best=np.argmin)\nplot_history_metric(axs[1, 1], 'val_mrcnn_bbox_loss', f_best=np.argmin)\nplot_history_metric(axs[1, 2], 'val_mrcnn_mask_loss', f_best=np.argmin)\n\naxs[0, 0].legend(['loss', 'val_loss'])\naxs[0, 0].grid()\n\naxs[0, 1].legend(['rpn_class_loss', 'val_rpn_class_loss '])\naxs[0, 1].grid()\n\naxs[0, 2].legend(['rpn_bbox_loss', 'val_rpn_bbox_loss'])\naxs[0, 2].grid()\n\naxs[1, 0].legend(['mrcnn_class_loss', 'val_mrcnn_class_loss'])\naxs[1, 0].grid()\n\naxs[1, 1].legend(['mrcnn_bbox_loss', 'val_mrcnn_bbox_loss'])\naxs[1, 1].grid()\n\naxs[1, 2].legend(['mrcnn_mask_loss', 'val_mrcnn_mask_loss'])\naxs[1, 2].grid()\n\n\nfig.tight_layout()\nfig.show() ","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:07:15.064233Z","iopub.execute_input":"2024-01-18T16:07:15.064571Z","iopub.status.idle":"2024-01-18T16:07:18.054714Z","shell.execute_reply.started":"2024-01-18T16:07:15.064511Z","shell.execute_reply":"2024-01-18T16:07:18.053921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction & Evaluation\n\n<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n<p style=\"padding: 10px; color:black; text-align: justify;\"> \nFor the testing phase, the model weights that yield the lowest validation loss are selected. The performance of the models is then assessed using a testing dataset comprised of approximately <b>8k</b> samples.\n</p>\n</div>","metadata":{}},{"cell_type":"code","source":"glob_list = glob.glob(f'/kaggle/working/fashion*/mask_rcnn_fashion_config_{best_epoch:04d}.h5')\nmodel_path = glob_list[0] if glob_list else ''","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:07:23.500026Z","iopub.execute_input":"2024-01-18T16:07:23.500622Z","iopub.status.idle":"2024-01-18T16:07:23.506127Z","shell.execute_reply.started":"2024-01-18T16:07:23.500568Z","shell.execute_reply":"2024-01-18T16:07:23.505068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_path = \"/kaggle/input/resnet50-training/mask_rcnn_fashion_config_0018.h5\"","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:07:23.883348Z","iopub.execute_input":"2024-01-18T16:07:23.883693Z","iopub.status.idle":"2024-01-18T16:07:23.887480Z","shell.execute_reply.started":"2024-01-18T16:07:23.883632Z","shell.execute_reply":"2024-01-18T16:07:23.886812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATASET_LENGTH = len(test_df)\nMAX_BATCH = 1\nBATCH_SIZE = sorted([int(DATASET_LENGTH/n) for n in range(1,DATASET_LENGTH+1) if DATASET_LENGTH % n ==0 and DATASET_LENGTH/n<=MAX_BATCH],reverse=True)[0]  \nSTEPS = int(DATASET_LENGTH/BATCH_SIZE)\nSTEPS","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:07:24.263940Z","iopub.execute_input":"2024-01-18T16:07:24.264420Z","iopub.status.idle":"2024-01-18T16:07:24.271352Z","shell.execute_reply.started":"2024-01-18T16:07:24.264358Z","shell.execute_reply":"2024-01-18T16:07:24.270714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class InferenceConfig(ClothConfig):\n    IMAGES_PER_GPU = 1\n    DETECTION_MIN_CONFIDENCE = 0.50\n    USE_MINI_MASK = False\n    STEPS_PER_EPOCH = STEPS\n    DETECTION_MAX_INSTANCES = 1\n\ninference_config = InferenceConfig()\ninference_config.display()","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:07:25.610260Z","iopub.execute_input":"2024-01-18T16:07:25.610846Z","iopub.status.idle":"2024-01-18T16:07:25.621066Z","shell.execute_reply.started":"2024-01-18T16:07:25.610790Z","shell.execute_reply":"2024-01-18T16:07:25.619925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = modellib.MaskRCNN(mode='inference', \n                          config=inference_config,\n                          model_dir=ROOT_DIR)\n\nassert model_path != '', \"Provide path to trained weights\"\nprint(\"Loading weights from \", model_path)\nmodel.load_weights(model_path, by_name=True)","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:07:41.697634Z","iopub.execute_input":"2024-01-18T16:07:41.697951Z","iopub.status.idle":"2024-01-18T16:07:50.247911Z","shell.execute_reply.started":"2024-01-18T16:07:41.697909Z","shell.execute_reply":"2024-01-18T16:07:50.247153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# validation set\ntesting_data = ClothDataset(test_df)\ntesting_data.load_dataset()\ntesting_data.prepare()","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:07:50.248976Z","iopub.execute_input":"2024-01-18T16:07:50.249353Z","iopub.status.idle":"2024-01-18T16:07:55.795702Z","shell.execute_reply.started":"2024-01-18T16:07:50.249311Z","shell.execute_reply":"2024-01-18T16:07:55.794448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1. Visualising some predictions","metadata":{}},{"cell_type":"code","source":"label_names = [\"short sleeve top\", \"long sleeve top\", \"short sleeve outwear\", \"long sleeve outwear\", \"vest\", \"sling\", \n              \"shorts\", \"trousers\", \"skirt\", \"short sleeve dress\", \"long sleeve dress\", \"vest dress\", \"sling dress\"]\nfig, axs = plt.subplots(figsize=(22, 8))\nfor i in range(10):\n    ax = plt.subplot(2, 5, i+1)\n    image_path = test_df.sample()['path'].values[0]\n    \n    img = cv2.imread(image_path)\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n    \n    result = model.detect([img])\n    r = result[0]\n    \n    if r['masks'].size > 0:\n        masks = np.zeros((img.shape[0], img.shape[1], r['masks'].shape[-1]), dtype=np.uint8)\n        for m in range(r['masks'].shape[-1]):\n            masks[:, :, m] = cv2.resize(r['masks'][:, :, m].astype('uint8'), \n                                        (img.shape[1], img.shape[0]), interpolation=cv2.INTER_NEAREST)\n        \n        y_scale = img.shape[0]/256\n        x_scale = img.shape[1]/256\n        rois = (r['rois'] * [y_scale, x_scale, y_scale, x_scale]).astype(int)\n    else:\n        masks, rois = r['masks'], r['rois']\n        \n    visualize.display_instances(img, rois, masks, r['class_ids'], \n                                ['bg']+label_names, r['scores'],\n                                title=image_id, ax=ax)\n\n    \nplt.tight_layout()    \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:08:31.233129Z","iopub.execute_input":"2024-01-18T16:08:31.233510Z","iopub.status.idle":"2024-01-18T16:08:38.119987Z","shell.execute_reply.started":"2024-01-18T16:08:31.233442Z","shell.execute_reply":"2024-01-18T16:08:38.119002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Predicted vs Actual","metadata":{}},{"cell_type":"code","source":"def plot_actual_vs_predicted(dataset, model, config, n_images):\n    image_ids = np.random.choice(dataset.image_ids, n_images)\n    fig, axs = plt.subplots(figsize=(30, 15))\n    for i, image_id in enumerate(image_ids):\n        image, image_meta, gt_class_id, gt_bbox, gt_mask =\\\n            modellib.load_image_gt(dataset, config, image_id)\n        info = dataset.image_info[image_id]\n        print(\"image ID: {}.{} ({}) {}\".format(info[\"source\"], info[\"id\"], image_id, \n                                               dataset.image_reference(image_id)))\n        print(\"Original image shape: \", modellib.parse_image_meta(image_meta[np.newaxis,...])[\"original_image_shape\"][0])\n\n        # Run object detection\n        molded_images = np.expand_dims(modellib.mold_image(image, inference_config), 0)\n        results = model.detect([image], verbose=0)\n#         results = model.detect_molded(np.expand_dims(image, 0), np.expand_dims(image_meta, 0), verbose=1)\n\n        # Display results\n        r = results[0]\n        log(\"gt_class_id\", gt_class_id)\n        log(\"gt_bbox\", gt_bbox)\n        log(\"gt_mask\", gt_mask)\n\n        # Compute AP over range 0.5 to 0.95 and print it\n        utils.compute_ap_range(gt_bbox, gt_class_id, gt_mask,\n                               r['rois'], r['class_ids'], r['scores'], r['masks'],\n                               verbose=1)\n        \n        ax = plt.subplot(2, 5, i+1)\n        visualize.display_differences(\n            image,\n            gt_bbox, gt_class_id, gt_mask,\n            r['rois'], r['class_ids'], r['scores'], r['masks'],\n            dataset.class_names,\n            show_box=False, show_mask=False, ax = ax,\n            iou_threshold=0.5, score_threshold=0.5)","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:09:38.445053Z","iopub.execute_input":"2024-01-18T16:09:38.445731Z","iopub.status.idle":"2024-01-18T16:09:38.456709Z","shell.execute_reply.started":"2024-01-18T16:09:38.445648Z","shell.execute_reply":"2024-01-18T16:09:38.455814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_actual_vs_predicted(testing_data, model, inference_config, 10)","metadata":{"execution":{"iopub.status.busy":"2024-01-18T16:09:39.414443Z","iopub.execute_input":"2024-01-18T16:09:39.415111Z","iopub.status.idle":"2024-01-18T16:09:47.198593Z","shell.execute_reply.started":"2024-01-18T16:09:39.415035Z","shell.execute_reply":"2024-01-18T16:09:47.197534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Evaluation mAP\n\n<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n<p style=\"padding: 10px; color:black; text-align: justify;\"> \nTypically, the evaluation of object detection and instance segmentation tasks is carried out using the <b>Average Precision</b> (AP) metric at various <b>Intersection over Union</b> (IoU) thresholds <b>(He et al., 2018)</b>. IoU is a measure that quantifies the overlap between predicted and ground truth bounding boxes or masks by calculating the ratio of their intersection to their union. An IoU value of 1 signifies perfectly aligned masks. \n</p>\n    \n<p style=\"padding: 10px; color:black; text-align: justify;\"> \nTo evaluate and compare the performance of the models, the Mean Average Precision (mAP) is computed across all testing samples at four distinct IoU thresholds, 0.5, 0.65, 0.75 and 0.85, which correspond to 50%, 65%, 75% and 85% overlap, respectively.\n</p>\n</div>","metadata":{}},{"cell_type":"code","source":"all_scores = []\nall_class_ids = []\nactual_labels = []\npredicted_labels = []\ndef calculate_mAP(data_loader, threshold=0.5):\n    APs = []\n\n    for image_id in tqdm(data_loader.image_ids):\n        # Load image and ground truth data\n        image, image_meta, gt_class_id, gt_bbox, gt_mask =\\\n            modellib.load_image_gt(data_loader, inference_config,\n                                   image_id, use_mini_mask=False)\n        molded_images = np.expand_dims(modellib.mold_image(image, inference_config), 0)\n\n\n        results = model.detect([image], verbose=0)\n        r = results[0]\n\n        # Calculate mAP\n        AP, precisions, recalls, overlaps = utils.compute_ap(gt_bbox, gt_class_id, gt_mask, \n                                                             r[\"rois\"], r[\"class_ids\"], r[\"scores\"], r['masks'],\n                                                             iou_threshold=threshold)\n        \n        APs.append(AP)\n        all_scores.extend(r[\"scores\"])\n        all_class_ids.extend(r[\"class_ids\"])\n        actual_labels.extend(gt_class_id)\n        predicted_labels.extend(r[\"class_ids\"])\n        \n\n    print(\"mAP@IoU{:.2f}: {:.2f}\".format(threshold, np.mean(APs)))","metadata":{"execution":{"iopub.status.busy":"2024-01-09T07:58:38.806057Z","iopub.execute_input":"2024-01-09T07:58:38.806417Z","iopub.status.idle":"2024-01-09T07:58:38.816464Z","shell.execute_reply.started":"2024-01-09T07:58:38.806349Z","shell.execute_reply":"2024-01-09T07:58:38.815605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"calculate_mAP(testing_data, 0.65)\ncalculate_mAP(testing_data, 0.75)\ncalculate_mAP(testing_data, 0.85)\ncalculate_mAP(testing_data, 0.95)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"calculate_mAP(testing_data)","metadata":{"execution":{"iopub.status.busy":"2024-01-09T09:12:03.869527Z","iopub.status.idle":"2024-01-09T09:12:03.870400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Histogram of confidence intervals\nfig, axs = plt.subplots(figsize=(15, 8))\nplt.hist(all_scores, bins=50, facecolor='blue', alpha=0.7)\nplt.title('Histogram of Confidence Scores Across All Garment Categories')\nplt.xlabel('Score')\nplt.ylabel('Number of Predictions')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-09T09:12:03.871523Z","iopub.status.idle":"2024-01-09T09:12:03.872348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"boxprops = dict(linestyle='-', linewidth=4, color='b')\nmedianprops = dict(linestyle='-', linewidth=4, color='r')\n\ndf = pd.DataFrame({\n    'Scores': all_scores,\n    'Class': all_class_ids\n})\n\ndf.boxplot(column='Scores', by='Class', showmeans=True,\n           whiskerprops=dict(linestyle='-', linewidth=1.5), \n           capprops=dict(linestyle='-', linewidth=1.5),\n           boxprops=boxprops, medianprops=medianprops, figsize=(20, 12))\nplt.title('Box Plot of Confidence Scores by Class Across All Images')\nplt.suptitle('')  # Suppress the default title\nplt.xticks(range(1, 14), labels=[\"short sleeve top\", \"long sleeve top\", \"short sleeve outwear\", \"long sleeve outwear\", \"vest\", \"sling\", \n                 \"shorts\", \"trousers\", \"skirt\", \"short sleeve dress\", \"long sleeve dress\", \"vest dress\", \"sling dress\"], rotation=45)\nplt.xlabel('Class ID')\nplt.ylabel('Score')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-09T09:12:03.873434Z","iopub.status.idle":"2024-01-09T09:12:03.874271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Confusion matrix for images with only one garment\n\n<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n<p style=\"padding: 10px; color:black; text-align: justify;\"> \nTo evaluate the performance of the models across various garment categories, the validation dataset was filtered to include only images annotated with a <b>single garment</b>.\n</p>\n    \n<p style=\"padding: 10px; color:black; text-align: justify;\"> \nThe ResNet50 model exhibits good precision for certain garment categories, such as <b>trousers</b>, <b>shorts</b>, and <b>short sleeve tops</b> with average confidences of <b>0.9</b>, <b>0.79</b> and <b>0.72</b>, respectively. However, there is a possibility of the model performing poorly with other garment categories such as skirt where the highest confidence score of 0.84 is when the predicted value is long sleeve top.\n</p>\n</div>","metadata":{}},{"cell_type":"code","source":"all_scores = []\nall_class_ids = []\nactual_labels = []\npredicted_labels = []\ndef calculate_mAP_conf(data_loader, threshold=0.5):\n    APs = []\n\n    for image_id in tqdm(data_loader.image_ids):\n        # Load image and ground truth data\n        image, image_meta, gt_class_id, gt_bbox, gt_mask =\\\n            modellib.load_image_gt(data_loader, inference_config,\n                                   image_id)\n        molded_images = np.expand_dims(modellib.mold_image(image, inference_config), 0)\n\n\n        results = model.detect([image], verbose=0)\n        r = results[0]\n\n        # Calculate mAP\n        AP, precisions, recalls, overlaps = utils.compute_ap(gt_bbox, gt_class_id, gt_mask, \n                                                             r[\"rois\"], r[\"class_ids\"], r[\"scores\"], r['masks'],\n                                                             iou_threshold=threshold)\n        \n        APs.append(AP)\n        if len(r[\"class_ids\"]) == 1 and r[\"class_ids\"][0] != 0:\n            all_scores.append(r[\"scores\"][0])\n            all_class_ids.append(r[\"class_ids\"][0])\n            actual_labels.extend(gt_class_id)\n            predicted_labels.append(r[\"class_ids\"][0])\n\n    print(\"mAP@IoU{:.2f}: {:.2f}\".format(threshold, np.mean(APs)))","metadata":{"execution":{"iopub.status.busy":"2024-01-09T09:12:03.875410Z","iopub.status.idle":"2024-01-09T09:12:03.876206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def filter_single_garment_rows(df):\n    \"\"\"Filter rows of a DataFrame that have only one garment.\"\"\"\n    return df[df['category_name'].apply(lambda x: len(x) == 1)]","metadata":{"execution":{"iopub.status.busy":"2024-01-09T09:12:03.877328Z","iopub.status.idle":"2024-01-09T09:12:03.878186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filtered_df = filter_single_garment_rows(test_df)\nfiltered_data = ClothDataset(filtered_df)\nfiltered_data.load_dataset()\nfiltered_data.prepare()\n\ncalculate_mAP_conf(filtered_data)","metadata":{"execution":{"iopub.status.busy":"2024-01-09T09:12:03.879237Z","iopub.status.idle":"2024-01-09T09:12:03.880034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_confidence_matrix(actual_labels, predicted_labels, all_scores, class_names):\n    n_classes = len(class_names)\n    \n    # Initialize confusion matrix with zeros\n    confidence_matrix = np.zeros((n_classes, n_classes))\n    count_matrix = np.zeros((n_classes, n_classes))\n    \n    # Populate the matrix\n    for i in range(len(predicted_labels)):\n        confidence_matrix[actual_labels[i], predicted_labels[i]] += all_scores[i]\n        count_matrix[actual_labels[i], predicted_labels[i]] += 1\n    \n    # Normalize the matrix\n    confidence_matrix = np.divide(confidence_matrix, count_matrix, out=np.zeros_like(confidence_matrix), where=count_matrix!=0)\n    \n    mask = np.ones_like(confidence_matrix, dtype=bool)\n    mask[confidence_matrix.argmax(axis=0), np.arange(n_classes)] = False\n    \n    # Plot the matrix\n    plt.figure(figsize=(20, 12))\n    ax = sns.heatmap(confidence_matrix, annot=True, cmap=\"Reds\", cbar=False, mask=mask,\n                     xticklabels=class_names, yticklabels=class_names)\n    sns.heatmap(confidence_matrix, annot=True, cmap=\"Blues\", cbar=False, mask=~mask, vmin=0.52,\n                xticklabels=class_names, yticklabels=class_names, ax=ax)\n    plt.xlabel('Actual labels')\n    plt.ylabel('Predicted labels')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2024-01-09T09:12:03.881099Z","iopub.status.idle":"2024-01-09T09:12:03.881708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_confidence_matrix(actual_labels, predicted_labels, all_scores, [\"background\", \"short sleeve top\", \"long sleeve top\", \"short sleeve outwear\", \"long sleeve outwear\", \"vest\", \"sling\", \n                 \"shorts\", \"trousers\", \"skirt\", \"short sleeve dress\", \"long sleeve dress\", \"vest dress\", \"sling dress\"])","metadata":{"execution":{"iopub.status.busy":"2024-01-09T09:12:03.882379Z","iopub.status.idle":"2024-01-09T09:12:03.882767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5. Images with one category only\n\n<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n<p style=\"padding: 10px; color:black; text-align: justify;\"> \nWe filter the validation data to include images with <b>only one annotated item</b> and record the AP at 50% and 75% thresholds. This process is repeated across all garment categories to see how the model performs for each individual category.\n</p>\n    \n<p style=\"padding: 10px; color:black; text-align: justify;\"> \nAlthough confidence scores cannot be directly interpreted as the prediction efficiency of the model, this suggests that the model tends to predict with higher confidence for categories that are more frequently represented in the dataset.\n</p>\n    \n</div>","metadata":{}},{"cell_type":"code","source":"def filter_by_category_name(df, category):\n    \"\"\"Filter rows of a DataFrame where 'category_name' column contains only the specified category.\"\"\"\n    return df[df['category_name'].apply(lambda x: x == [category])]","metadata":{"execution":{"iopub.status.busy":"2024-01-09T09:12:03.883400Z","iopub.status.idle":"2024-01-09T09:12:03.883969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for category in [\"short sleeve top\", \"long sleeve top\", \"short sleeve outwear\", \"long sleeve outwear\", \"vest\", \"sling\", \n                 \"shorts\", \"trousers\", \"skirt\", \"short sleeve dress\", \"long sleeve dress\", \"vest dress\", \"sling dress\"]:\n    filtered_df = filter_by_category_name(test_df, category)\n    filtered_data = ClothDataset(filtered_df)\n    filtered_data.load_dataset()\n    filtered_data.prepare()\n    print(\"\\n\\nEvaluation for images with only \" + category)\n    calculate_mAP(filtered_data)\n    calculate_mAP(filtered_data, 0.75)","metadata":{"execution":{"iopub.status.busy":"2024-01-09T09:12:03.884623Z","iopub.status.idle":"2024-01-09T09:12:03.885156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6. Images with atleast one category\n\n<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n<p style=\"padding: 10px; color:black; text-align: justify;\"> \nWe filter the validation data to include images with <b>atleast one annotated item</b> and record the AP at 50% and 75% thresholds. This process is repeated across all garment categories to see how the model performs for each individual category.\n</p>\n    \n</div>","metadata":{}},{"cell_type":"code","source":"def filter_by_category_name(df, category):\n    \"\"\"Filter rows of a DataFrame where 'category_name' column contains the specified category.\"\"\"\n    return df[df['category_name'].apply(lambda x: category in x)]","metadata":{"execution":{"iopub.status.busy":"2024-01-09T09:12:03.885894Z","iopub.status.idle":"2024-01-09T09:12:03.886417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for category in [\"short sleeve top\", \"long sleeve top\", \"short sleeve outwear\", \"long sleeve outwear\", \"vest\", \"sling\", \n                 \"shorts\", \"trousers\", \"skirt\", \"short sleeve dress\", \"long sleeve dress\", \"vest dress\", \"sling dress\"]:\n    filtered_df = filter_by_category_name(test_df, category)\n    filtered_data = ClothDataset(filtered_df)\n    filtered_data.load_dataset()\n    filtered_data.prepare()\n    print(\"\\n\\nEvaluation for images with atleast one\" + category)\n    calculate_mAP(filtered_data)\n    calculate_mAP(filtered_data, 0.75)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#E7E4DD;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n<p style=\"padding: 10px; color:black; text-align: justify;\">\n    A similar version of this notebook but with the EfficientNetV2-B0 backbone can be found here - <a href=\"https://www.kaggle.com/code/thusharanair/fashion-segmentation-maskrcnn-efficientnetv2-b0\">Fashion Instance Segmentation using MaskRCNN(EfficientNetV2-B0)</a>\n</p>\n</div>\n\n<center>\n<img src=\"https://media.giphy.com/media/v1.Y2lkPTc5MGI3NjExa2lmeHhzaGU0b3ZlNnBsNzc0ZTdtY25obWtidm14MnFiZWdyNmxtZiZlcD12MV9pbnRlcm5hbF9naWZfYnlfaWQmY3Q9Zw/6tHy8UAbv3zgs/giphy.gif\">\n</center>","metadata":{}}]}