{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"colab":{"provenance":[],"collapsed_sections":["EvxPB-mqmFrW","wATcP3FVmFre","75TVsB85mFre","U0tL3mZemFrj"]},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":1188410,"sourceType":"datasetVersion","datasetId":675948}],"dockerImageVersionId":30648,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"*This notebook is particularly designed for Youtube Academic Projects Series by Spartificial Innovations Pvt Ltd. However, anyone is allowed to use it, who is interested in understanding applications of Machine Learning through real world projects.*       \n*Please contact team@spartificial.com or visit https://spartificial.com/ to know more.*         \n**This project is inspired from the work of K SCOTT MADER and team**","metadata":{"id":"F3ltk9LJcqT8"}},{"cell_type":"markdown","source":"# Detecting Ships using Satellite Images with Deep Learning\n\n---\n\n**Project ID: SDSI0922**      \n\n**Project Name: Ship Detection from Satellite Imagery**\n\n---\n\n<center> <img src = \"http://www.schneeberg.land/bilder/schiff.gif\" width = 35%>\n    ","metadata":{"id":"ZfuQuobXmFqt"}},{"cell_type":"markdown","source":"## Workflow of this notebook\n**1)** [Introducing Dataset](#h1)             \n**2)** [Quick guide on Object Detection, Image Classification and Segmentation](#h2)      \n**3)** [Importing necessary libraries and modules for this notebook](#h2.5)       \n**4)** [Exploring the Dataset](#h3)             \n**5)** [Grasping the idea of Run Length Encoding and Decoding](#h4)         \n**6)** [Preparing data for our model](#h5)           \n**7)** [Brief introduction to UNET model](#h6)         \n**8)** [Build and train UNET model](#h7)            \n**9)** [Tasks for you](#h8)          \n\n## Dataset & Aim<a class=\"anchor\"  id=\"h1\"></a>\n    \n<img src = \"https://eoimages.gsfc.nasa.gov/images/imagerecords/2000/2938/SinkingShip_10_15_02_lrg.jpg\" width = 35%>  \n       \n- We will be dealing with [this dataset](https://www.kaggle.com/competitions/airbus-ship-detection).\n- Search \"Airbus Ship Detection Dataset\" and add it to this kernel.\n- It consists of train and test image folders along with sample_submission and train_ship_segmentation csv files.\n- Soon we will jump into more details of these files.\n- We will build and train UNET model from scratch for image segmentaion.\n\n## Image Classification, Object Detection and Semantic Segmentation <a class=\"anchor\"  id=\"h2\"></a>\n\n- If you are new to computer vision then these terms may confuse you.\n- These are some of the most common applications of machine learning in computer vision.\n- Below is a quick guide before we start exploring our data.\n\n<img src = \"https://i1.wp.com/bdtechtalks.com/wp-content/uploads/2021/05/image-classification-vs-object-detection-vs-semantic-segmentation.jpg\" width = 65%>\n\n* **Image Classificaton:-** Determines whether a certain type of object is present in an image or not.    \n* **Object Detection:-** Object detection takes image classification one step further and provides the bounding box where detected objects are located.        \n* **Semantic Segmentation:-** Specifies the object class of each pixel in an input image. This is what we are targeting in this notebook!  \n* **Instance Segmentation:-** Separates individual instances of each type of object. \n\nThe images below and above shall help you visualize these ideas.\n\n<img src = \"https://deeplobe.ai/wp-content/uploads/2021/05/SENTIMENT-ANALYSIS-1024x683.png\" width = 65%>\n\n*You can read more about it over <u>[here](https://venturebeat.com/ai/new-deep-learning-model-brings-image-segmentation-to-edge-devices/)</u>.*\n\n## Importing libraries and modules needed for this notebook <a class=\"anchor\"  id=\"h2.5\"></a>","metadata":{"id":"UCL2tmN8mFqz"}},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')\n\nimport os\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom skimage.io import imread\nfrom skimage.segmentation import mark_boundaries\nfrom skimage.util import montage\nfrom skimage.morphology import label\n\nimport gc\ngc.enable()","metadata":{"id":"duMa5oo7mFrI","execution":{"iopub.status.busy":"2024-02-28T09:50:38.316522Z","iopub.execute_input":"2024-02-28T09:50:38.316894Z","iopub.status.idle":"2024-02-28T09:50:41.004394Z","shell.execute_reply.started":"2024-02-28T09:50:38.316864Z","shell.execute_reply":"2024-02-28T09:50:41.003507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Documents of the above used libraries and modules in case you aren't familiar and want to know more about it:-**\n- [os](https://docs.python.org/3/library/os.html)\n- [numpy](https://numpy.org/doc/1.23/user/absolute_beginners.html)\n- [pandas](https://pandas.pydata.org/docs/user_guide/index.html#user-guide)\n- [seaborn](https://seaborn.pydata.org/tutorial/introduction.html)\n- [matplotlib.pyplot](https://matplotlib.org/stable/tutorials/introductory/pyplot.html)\n- [skimage.io.imread](https://scikit-image.org/docs/stable/api/skimage.io.html#skimage.io.imread)\n- [skimage.segmentation.mark_boundaries](https://scikit-image.org/docs/stable/api/skimage.segmentation.html#skimage.segmentation.mark_boundaries)\n- [skimage.util.montage](https://scikit-image.org/docs/stable/api/skimage.util.html#skimage.util.montage)\n- [skimage.morphology.label](https://scikit-image.org/docs/stable/api/skimage.morphology.html#skimage.morphology.label)\n- [gc.enable()](https://docs.python.org/3/library/gc.html)\n","metadata":{"id":"h4nXx7X8mFrN"}},{"cell_type":"markdown","source":"## Exploring the data <a class=\"anchor\"  id=\"h3\"></a>","metadata":{"id":"crJdN2a9mFrO"}},{"cell_type":"code","source":"# Train and test directories\ntrain_image_dir = '/kaggle/input/airbus-224x224x3/train_v2'\ntest_image_dir = \"/kaggle/input/airbus-224x224x3/test_v2\"","metadata":{"id":"q1XfqoRvmFrP","execution":{"iopub.status.busy":"2024-02-28T09:50:41.006309Z","iopub.execute_input":"2024-02-28T09:50:41.006838Z","iopub.status.idle":"2024-02-28T09:50:41.011049Z","shell.execute_reply.started":"2024-02-28T09:50:41.006805Z","shell.execute_reply":"2024-02-28T09:50:41.010202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Getting into train directory\ntrain_images = os.listdir(train_image_dir)\ntrain_images.sort()\nprint(f\"Total of {len(train_images)} images in train directory.\\nHere is how first five train_images looks like:- {train_images[:5]}\")","metadata":{"id":"_v7uprIWmFrP","outputId":"2a11ea28-c293-408e-acf2-7b4cc2fe319a","execution":{"iopub.status.busy":"2024-02-28T09:50:41.012182Z","iopub.execute_input":"2024-02-28T09:50:41.013244Z","iopub.status.idle":"2024-02-28T09:50:43.660313Z","shell.execute_reply.started":"2024-02-28T09:50:41.013204Z","shell.execute_reply":"2024-02-28T09:50:43.65928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Using for loop to generate different images to understand how data looks like\nplt.figure(figsize=(15,15))\nplt.suptitle('TRAIN IMAGES\\n', weight = 'bold', fontsize = 15, color = 'r')\nfor i in range(16):\n    plt.subplot(4, 4, i+1)\n    plt.imshow(imread(train_image_dir + \"/\" + train_images[i]))\n    plt.title(f\"{train_images[i]}\", weight = 'bold')\n    plt.axis('off')\nplt.tight_layout()","metadata":{"id":"2IfuQJbDmFrQ","outputId":"4594231b-4728-4b56-fb79-57f85b7541a5","execution":{"iopub.status.busy":"2024-02-28T09:50:43.663071Z","iopub.execute_input":"2024-02-28T09:50:43.663423Z","iopub.status.idle":"2024-02-28T09:50:46.606543Z","shell.execute_reply.started":"2024-02-28T09:50:43.663391Z","shell.execute_reply":"2024-02-28T09:50:46.605529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{"id":"6IUkSmz-mFrR"}},{"cell_type":"code","source":"# Train ships segmented masks\nmasks = pd.read_csv(\"/kaggle/input/airbus-224x224x3/train_ship_segmentations_v2.csv\")\nmasks.head(10)","metadata":{"id":"zoIy1ed2mFrR","outputId":"58dfddd8-5ac8-417a-e644-e4424d88a4ae","execution":{"iopub.status.busy":"2024-02-28T09:50:46.608017Z","iopub.execute_input":"2024-02-28T09:50:46.608451Z","iopub.status.idle":"2024-02-28T09:50:47.751833Z","shell.execute_reply.started":"2024-02-28T09:50:46.608406Z","shell.execute_reply":"2024-02-28T09:50:47.75098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Here, we can see some image ids are repeated. That is because we are given masks for each ship in one image.\n- This simply means that number of ships in train image = total of repeated image ids in this masks data frame.\n- NaN simply means that there are no ships in that image.\n- We can combine this masks into one image for same image ids.\n- Before that we need to make ourselves comfortable with this encoded format.\n\n## Run Length Encoding (RLE) and Decoding <a class=\"anchor\"  id=\"h4\"></a>\n- Run length encoding is a lossless compression.\n- Lossless compression allows the original data to be perfectly reconstructed from the compressed data. \n- We simply run through the data and count how many times each data point is repeated without any breaks.\n- Consider a small example below to grasp this idea:-\n\n<img src = \"https://img.api.video/1628663040-run-length.png?auto=format&dpr=1&fm=jpg&w=1370\" width = 35%> \n\n- We tend to use RLE for data that contains long runs of the same value.\n- If your data is complex without long runs then RLE can result in negative compression.\n- There are many variations to it - run accross rows, run accross columns, run until pixel changes, etc.\n- We also need to decompress the RLE data using run length decoding in order to use it. \n\n**Image example to understand it in a better way:-**    \n- Let black pixels be 1 and white pixels be 0 in 10x10 image shown below.     \n\n<img src=\"https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcScKG8WJQhjdfUJU2-BBnj7UAgaoF3tK4CZYLRvSjcqjA6H-2BwIwFa3AXRb_Bt2WR5qi4&usqp=CAU\" width = 35%>         ","metadata":{"id":"As43_L3smFrS"}},{"cell_type":"code","source":"row_rle = ['10 1', \n           '4 1 2 0 4 1', \n           '3 1 4 0 3 1', \n           '2 1 6 0 2 1',\n           '1 1 2 0 1 1 2 0 1 1 2 0 1 1', \n           '1 1 8 0 1 1', \n           '3 1 1 0 2 1 1 0 3 1', \n           '2 1 1 0 1 1 2 0 1 1 1 0 2 1', \n           '1 1 1 0 1 1 1 0 2 1 1 0 1 1 1 0 1 1', \n           '10 1',\n           'Total']\n\npixels = [len(row.split(\" \")) for row in row_rle if row != 'Total']\nsum_pixels = np.array(pixels).sum()\npixels.append(sum_pixels)\n\ndata = {\n    'Row - RLE' : row_rle,\n    'Pixels' : pixels\n}\n\nrle_df = pd.DataFrame(data)\nrle_df.index+=1\nrle_df\n","metadata":{"id":"96TekMv2mFrS","outputId":"ea74af11-f7e3-4383-f9fd-8a790e118f79","execution":{"iopub.status.busy":"2024-02-28T09:50:47.752874Z","iopub.execute_input":"2024-02-28T09:50:47.753143Z","iopub.status.idle":"2024-02-28T09:50:47.766046Z","shell.execute_reply.started":"2024-02-28T09:50:47.753119Z","shell.execute_reply":"2024-02-28T09:50:47.764991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Clearly we have compressed 100 data points into 84 data points.\n- Some rows like 5, 8 and 14 gave more data points after encoding than it was in original.\n- Reason being short runs as there were multiple breaks between black and white pixels.\n- We can now apply this idea onto our data!","metadata":{"id":"85s6TxoEmFrS"}},{"cell_type":"code","source":"# Let us now see how it works for Image id:- 0005d01c8.jpg we have in the mask data frame\n\n# Original image from training set\nimg_arr = imread(train_image_dir + '/' + '0005d01c8.jpg')\nplt.figure(figsize=(15,8))\nplt.imshow(img_arr)\nplt.show()","metadata":{"id":"ZqlG5qDbmFrS","outputId":"36b8e417-8250-4876-e940-61981f73d79b","execution":{"iopub.status.busy":"2024-02-28T09:50:47.767456Z","iopub.execute_input":"2024-02-28T09:50:47.767808Z","iopub.status.idle":"2024-02-28T09:50:48.148305Z","shell.execute_reply.started":"2024-02-28T09:50:47.767777Z","shell.execute_reply":"2024-02-28T09:50:48.14735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_arr.shape","metadata":{"id":"IJhFOljtmFrT","outputId":"3d10f1b8-5828-4635-e7ee-c6c581ddb464","execution":{"iopub.status.busy":"2024-02-28T09:50:48.149581Z","iopub.execute_input":"2024-02-28T09:50:48.149908Z","iopub.status.idle":"2024-02-28T09:50:48.156266Z","shell.execute_reply.started":"2024-02-28T09:50:48.14988Z","shell.execute_reply":"2024-02-28T09:50:48.15532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filter out all 0005d01c8.jpg image ids and respective encoded data \n# 2 ships means 2 same image ids will be there!\nrle_0 = masks.query('ImageId==\"0005d01c8.jpg\"')['EncodedPixels']\nrle_0","metadata":{"id":"dz6on9OLmFrT","outputId":"98fcaf86-a88e-43e4-9f66-8ed365bb0f25","execution":{"iopub.status.busy":"2024-02-28T09:50:48.157378Z","iopub.execute_input":"2024-02-28T09:50:48.157668Z","iopub.status.idle":"2024-02-28T09:50:48.190802Z","shell.execute_reply.started":"2024-02-28T09:50:48.157645Z","shell.execute_reply":"2024-02-28T09:50:48.189962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make a list of each mask shown above also visualise whats happening!\nmask_lst, ct = [], 1\nfor mask in rle_0:\n    print(f\"Mask {ct} -\\n{mask}\\n\\n\")\n    mask_lst.append(mask)\n    ct+=1","metadata":{"id":"0ccDO-J0mFrT","outputId":"a262b896-b911-4abd-b8ab-f68d70fc603d","execution":{"iopub.status.busy":"2024-02-28T09:50:48.194517Z","iopub.execute_input":"2024-02-28T09:50:48.1948Z","iopub.status.idle":"2024-02-28T09:50:48.2001Z","shell.execute_reply.started":"2024-02-28T09:50:48.194776Z","shell.execute_reply":"2024-02-28T09:50:48.199059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split and Display how the first mask in the list looks like\nsplit = mask_lst[0].split()\nprint(split)","metadata":{"id":"JcXAQabMmFrU","outputId":"8f589ee8-5bd6-4878-cc8b-659f04dcf323","execution":{"iopub.status.busy":"2024-02-28T09:50:48.201122Z","iopub.execute_input":"2024-02-28T09:50:48.201366Z","iopub.status.idle":"2024-02-28T09:50:48.21074Z","shell.execute_reply.started":"2024-02-28T09:50:48.201344Z","shell.execute_reply":"2024-02-28T09:50:48.209715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- This data shows start_pixels and lenghts where we can think ship to exist in the original image.\n- For example, 56777 3 shows that pixels 56777, 56778, 56779 contributes to the ship.\n- Our target is to create an image with these pixels labeled as 1 and remaining as 0.\n- This is how we can produce a mask for respective image.","metadata":{"id":"0lTvIaOMmFrU"}},{"cell_type":"code","source":"# Grab all the starting pixels and lenghts and convert it into integers using numpy \nstarts, lengths = [np.array(x, dtype = int) for x in (split[::2], split[1::2])]\nstarts, lengths","metadata":{"id":"JYxNhVTemFrU","outputId":"c8faed91-d89e-4384-eb49-3d80f3f7df95","execution":{"iopub.status.busy":"2024-02-28T09:50:48.211973Z","iopub.execute_input":"2024-02-28T09:50:48.212258Z","iopub.status.idle":"2024-02-28T09:50:48.224589Z","shell.execute_reply.started":"2024-02-28T09:50:48.212233Z","shell.execute_reply":"2024-02-28T09:50:48.223716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get the ending pixels. \n'''Examples:- \n56010 1 ---> Starts at 56010 and ends at 56010\n56777 3 ---> Starts at 56777 and ends at 56779\n57544 6 ---> Starts at 57544 and ends at 57549'''\nends = starts + lengths - 1\npd.DataFrame({\n    'Starts' : starts,\n    'Lengths' : lengths,\n    'Ends' : ends\n}).head(10)","metadata":{"id":"m_jYpfHymFrU","outputId":"7eebca8d-5238-4a94-90bb-4866838f67b3","execution":{"iopub.status.busy":"2024-02-28T09:50:48.226013Z","iopub.execute_input":"2024-02-28T09:50:48.226309Z","iopub.status.idle":"2024-02-28T09:50:48.239185Z","shell.execute_reply.started":"2024-02-28T09:50:48.226285Z","shell.execute_reply":"2024-02-28T09:50:48.238144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create 1s in place of these pixels and rest should be 0\nimg = np.zeros(768*768, dtype = np.uint8)\nfor start, end in zip(starts, ends):\n    img[start:end+1] = 1","metadata":{"id":"5gXIUsCOmFrV","execution":{"iopub.status.busy":"2024-02-28T10:34:04.886562Z","iopub.execute_input":"2024-02-28T10:34:04.887258Z","iopub.status.idle":"2024-02-28T10:34:04.893445Z","shell.execute_reply.started":"2024-02-28T10:34:04.887213Z","shell.execute_reply":"2024-02-28T10:34:04.892331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check how output looks\nimg[56776:56781] # Should output 0, 1 , 1, 1 ,0 as we know 56777, 56778, 56779 ---> 1 and 5676, 56780 ---> 0","metadata":{"id":"wyzWtoRnmFrV","outputId":"08a86a66-609a-44d4-f689-b613a1fa92bb","execution":{"iopub.status.busy":"2024-02-28T10:34:10.22492Z","iopub.execute_input":"2024-02-28T10:34:10.225287Z","iopub.status.idle":"2024-02-28T10:34:10.231818Z","shell.execute_reply.started":"2024-02-28T10:34:10.225259Z","shell.execute_reply":"2024-02-28T10:34:10.230811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Copy-Paste this idea for another ship in the image\nsplit_1 = mask_lst[1].split()                                                                # Split the mask into start_pixels and lengths\nstarts, lengths = [np.array(x, dtype = int) for x in (split_1[0:][::2], split_1[1:][::2])]   # Generate arrays from only starts and lengths\nends = starts + lengths - 1                                                                  # Start pixel to end pixel will be start - 1 + length\nimg1 = np.zeros(768*768, dtype = np.uint8)                                                   # 1D array containing all zeros\nfor start, end in zip(starts, ends):                                                         # For each start to end pair\n    img1[start:end+1] = 1                                                                    # Convert the values from 0 to 1","metadata":{"id":"EzpbR0zNmFrV","execution":{"iopub.status.busy":"2024-02-28T10:34:23.764062Z","iopub.execute_input":"2024-02-28T10:34:23.764428Z","iopub.status.idle":"2024-02-28T10:34:23.772787Z","shell.execute_reply.started":"2024-02-28T10:34:23.7644Z","shell.execute_reply":"2024-02-28T10:34:23.771535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reshaping both the ship masks and combining it to form the final mask!\nimg = img.reshape(768, 768)\nimg1 = img1.reshape(768, 768)\nfinal = img+img1\nprint(final, '\\n\\n', final.shape, '\\n\\n', final.ndim)","metadata":{"id":"GaFH_dzQmFrV","outputId":"3bf72b85-38a8-4392-aace-437c46e4a9c3","execution":{"iopub.status.busy":"2024-02-28T10:34:32.811128Z","iopub.execute_input":"2024-02-28T10:34:32.811739Z","iopub.status.idle":"2024-02-28T10:34:32.818194Z","shell.execute_reply.started":"2024-02-28T10:34:32.811709Z","shell.execute_reply":"2024-02-28T10:34:32.817239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Expand dimension of this array to have only 1 channel in the mask and visualise original and final mask\nfinal = np.expand_dims(final, -1) # -1 means the last available dimenstion, in this case it is 2. Hence, on axis = 2 we will get 1.\noriginal = imread(train_image_dir+'/'+train_images[15])\nplt.figure(figsize=(15, 8))\nplt.subplot(1, 2, 1)\nplt.title(f\"Original - Train Image, {original.shape}\")\nplt.imshow(original)\nplt.subplot(1, 2, 2)\nplt.title(f\"Mask generated from the RLE data for each ship, {final.shape}\")\nplt.imshow(final, cmap = \"gray\")\nplt.tight_layout()\nplt.show()","metadata":{"id":"GQf7ZZ3XmFrW","outputId":"5d90b3bd-4511-4769-9b82-67c79c8fd62a","execution":{"iopub.status.busy":"2024-02-28T09:50:48.288051Z","iopub.execute_input":"2024-02-28T09:50:48.288365Z","iopub.status.idle":"2024-02-28T09:50:48.893195Z","shell.execute_reply.started":"2024-02-28T09:50:48.288338Z","shell.execute_reply":"2024-02-28T09:50:48.892214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Oh, something is off!\n- Our mask needs to be transposed. ","metadata":{"id":"6kZifw3AmFrW"}},{"cell_type":"code","source":"# Copy Paste the code from the prev cell with one change - Transpose!\nimg = img.reshape(768, 768).T     # Transpose the first ship mask\nimg1 = img1.reshape(768, 768).T   # Transpose the second ship mask\nfinal = img+img1                  # Generate the final mask with two ships \nfinal = np.expand_dims(final, -1) \nplt.figure(figsize=(15, 8))\nplt.subplot(1, 2, 1)\nplt.title(f\"Original - Train Image, {original.shape}\")\nplt.imshow(original)\nplt.subplot(1, 2, 2)\nplt.title(f\"Mask generated from the RLE data for each ship, {final.shape}\")\nplt.imshow(final, cmap = \"Blues_r\")\nplt.tight_layout()\nplt.show()","metadata":{"id":"jZrbe3d6mFrW","outputId":"530150c2-e503-401c-8fbf-3653cbd2c77e","execution":{"iopub.status.busy":"2024-02-28T09:50:48.89459Z","iopub.execute_input":"2024-02-28T09:50:48.894963Z","iopub.status.idle":"2024-02-28T09:50:49.502737Z","shell.execute_reply.started":"2024-02-28T09:50:48.894933Z","shell.execute_reply":"2024-02-28T09:50:49.501813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- So this is how the EncodedPixels data for one image id looks like!\n- We can build a function that can quickly generate such masks for all the EncodedPixels wrt to its ImageId.","metadata":{"id":"ccK3v3elmFrW"}},{"cell_type":"code","source":"# Define functions to do these tasks for all the training images\ndef rle_decode(mask_rle, shape=(768,768)):\n    '''\n    Input arguments -\n    mask_rle: Mask of one ship in the train image\n    shape: Output shape of the image array\n    '''\n    s = mask_rle.split()                                                               # Split the mask of each ship that is in RLE format\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]     # Get the start pixels and lengths for which image has ship\n    ends = starts + lengths - 1                                                        # Get the end pixels where we need to stop\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)                                  # A 1D vec full of zeros of size = 768*768\n    for lo, hi in zip(starts, ends):                                                   # For each start to end pixels where ship exists\n        img[lo:hi+1] = 1                                                               # Fill those values with 1 in the main 1D vector\n    '''\n    Returns -\n    Transposed array of the mask: Contains 1s and 0s. 1 for ship and 0 for background\n    '''\n    return img.reshape(shape).T                                                       \n\ndef masks_as_image(in_mask_list):\n    '''\n    Input - \n    in_mask_list: List of the masks of each ship in one whole training image\n    '''\n    all_masks = np.zeros((768, 768), dtype = np.int16)                                 # Creating 0s for the background\n    for mask in in_mask_list:                                                          # For each ship rle data in the list of mask rle \n        if isinstance(mask, str):                                                      # If the datatype is string\n            all_masks += rle_decode(mask)                                              # Use rle_decode to create one mask for whole image\n    '''\n    Returns - \n    Full mask of the training image whose RLE data has been passed as an input\n    '''\n    return np.expand_dims(all_masks, -1)","metadata":{"id":"dPjmkwHjmFrW","execution":{"iopub.status.busy":"2024-02-28T09:50:49.50415Z","iopub.execute_input":"2024-02-28T09:50:49.504438Z","iopub.status.idle":"2024-02-28T09:50:49.513864Z","shell.execute_reply.started":"2024-02-28T09:50:49.504413Z","shell.execute_reply":"2024-02-28T09:50:49.512755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for num in [3, 4, 5, 6]:\n    rle_0 = masks.query(f'ImageId==\"{train_images[num-1]}\"')['EncodedPixels']\n    img_0 = masks_as_image(rle_0)\n    original = imread(train_image_dir+\"/\"+train_images[num-1])\n    plt.figure(figsize=(15, 8))\n    plt.subplot(1, 2, 1)\n    plt.title(f\"Original - Train Image {original.shape}\")\n    plt.imshow(original)\n    plt.subplot(1, 2, 2)\n    plt.title(f\"Mask generated from the RLE data for each ship {final.shape}\")\n    plt.imshow(img_0, cmap = \"Blues_r\")\n    plt.tight_layout()\n    plt.show()","metadata":{"id":"N3JLeU2ps5zG","outputId":"3a33c31e-e02a-4e87-dd62-de750de7707e","execution":{"iopub.status.busy":"2024-02-28T09:50:49.514973Z","iopub.execute_input":"2024-02-28T09:50:49.515239Z","iopub.status.idle":"2024-02-28T09:50:51.968487Z","shell.execute_reply.started":"2024-02-28T09:50:49.515216Z","shell.execute_reply":"2024-02-28T09:50:51.9676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- We have succesfully constructed some functions that will take in the rle data and convert it into mask!\n- We can now begin with spliting the data into train and validation.\n\n## Preparing Train and Validation Data <a class=\"anchor\"  id=\"h5\"></a>","metadata":{"id":"EvxPB-mqmFrW"}},{"cell_type":"code","source":"'''Note that NaN values in the EncodedPixels are of float type and everything else is a string type'''   \n\n# Add a new feature to the masks data frame named as ship. If Encoded pixel in any row is a string, there is a ship else there isn't. \nmasks['ships'] = masks['EncodedPixels'].map(lambda c_row: 1 if isinstance(c_row, str) else 0)\nmasks.head(9)","metadata":{"id":"h0BarNjlmFrX","outputId":"f5207375-1076-4703-81c7-b77c4f55f013","execution":{"iopub.status.busy":"2024-02-28T09:50:51.969729Z","iopub.execute_input":"2024-02-28T09:50:51.970106Z","iopub.status.idle":"2024-02-28T09:50:52.14353Z","shell.execute_reply.started":"2024-02-28T09:50:51.970074Z","shell.execute_reply":"2024-02-28T09:50:52.142595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Making a new data frame with unique image ids where we are summing up the ship counts\nunique_img_ids = masks.groupby('ImageId').agg({'ships': 'sum'}).reset_index() \nunique_img_ids.index+=1 # Incrimenting all the index by 1\nunique_img_ids.head()","metadata":{"id":"iW_hlCevmFrX","outputId":"5c2f0605-a7f1-4018-c2db-e231138a817c","execution":{"iopub.status.busy":"2024-02-28T09:50:52.144728Z","iopub.execute_input":"2024-02-28T09:50:52.145024Z","iopub.status.idle":"2024-02-28T09:50:52.378923Z","shell.execute_reply.started":"2024-02-28T09:50:52.145Z","shell.execute_reply":"2024-02-28T09:50:52.377924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Adding two new features to unique_img_ids data frame. If ship exists in image, val is 1 else 0. And it's vec form\nunique_img_ids['has_ship'] = unique_img_ids['ships'].map(lambda x: 1.0 if x>0 else 0.0)\nunique_img_ids.head()","metadata":{"id":"aPaLVit8mFrX","outputId":"76a04e17-fa8c-4a27-fbdf-42f63a6fd876","execution":{"iopub.status.busy":"2024-02-28T09:50:52.380187Z","iopub.execute_input":"2024-02-28T09:50:52.380536Z","iopub.status.idle":"2024-02-28T09:50:52.459691Z","shell.execute_reply.started":"2024-02-28T09:50:52.380507Z","shell.execute_reply":"2024-02-28T09:50:52.458704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check the size of the files. Will take some time to run as there are loads of files!!!\nunique_img_ids['file_size_kb'] = unique_img_ids['ImageId'].map(lambda c_img_id: os.stat(os.path.join(train_image_dir, c_img_id)).st_size/1024)\n'''os.stat is used to get status of the specified path. Here, st_size represents size of the file in bytes. Converting it into kB!'''","metadata":{"id":"ETbI2Byss5zG","outputId":"ed08ef0c-db14-46ed-8fab-30af19a35c3b","execution":{"iopub.status.busy":"2024-02-28T09:50:52.460752Z","iopub.execute_input":"2024-02-28T09:50:52.461002Z","iopub.status.idle":"2024-02-28T10:04:13.297872Z","shell.execute_reply.started":"2024-02-28T09:50:52.46098Z","shell.execute_reply":"2024-02-28T10:04:13.296898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We can get rid of any images whose size is less than 35 Kb. As some of the files are corrupted! \nunique_img_ids[unique_img_ids.file_size_kb<5].head()","metadata":{"id":"dh7Z40aIs5zH","outputId":"c86e02f4-f0af-4561-d8de-e9d827cbd48b","execution":{"iopub.status.busy":"2024-02-28T10:04:13.298895Z","iopub.execute_input":"2024-02-28T10:04:13.299138Z","iopub.status.idle":"2024-02-28T10:04:13.315132Z","shell.execute_reply.started":"2024-02-28T10:04:13.299117Z","shell.execute_reply":"2024-02-28T10:04:13.314403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rle_0 = masks.query(f'ImageId==\"0318fc519.jpg\"')['EncodedPixels']\nimg_0 = masks_as_image(rle_0)\noriginal = imread(train_image_dir+\"/\"+'0318fc519.jpg')\nplt.figure(figsize=(15, 8))\nplt.subplot(1, 2, 1)\nplt.title(f\"Original - Train Image {original.shape}\")\nplt.imshow(original)\nplt.subplot(1, 2, 2)\nplt.title(f\"Mask generated from the RLE data for each ship {final.shape}\")\nplt.imshow(img_0, cmap = \"Blues_r\")\nplt.tight_layout()\nplt.show()","metadata":{"id":"nYF_0MYfmFrY","outputId":"cf664010-00ee-432a-dfdc-2f7b79705d98","execution":{"iopub.status.busy":"2024-02-28T10:04:13.316243Z","iopub.execute_input":"2024-02-28T10:04:13.316505Z","iopub.status.idle":"2024-02-28T10:04:13.926512Z","shell.execute_reply.started":"2024-02-28T10:04:13.316482Z","shell.execute_reply":"2024-02-28T10:04:13.922937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Keep the files whose size > 35 kB\nunique_img_ids = unique_img_ids[unique_img_ids.file_size_kb > 5]\nunique_img_ids.head()","metadata":{"id":"Vcom6VhkmFrY","outputId":"022411aa-a97c-4e5d-a31d-02ea013dd24b","execution":{"iopub.status.busy":"2024-02-28T10:04:13.934152Z","iopub.execute_input":"2024-02-28T10:04:13.93454Z","iopub.status.idle":"2024-02-28T10:04:13.958702Z","shell.execute_reply.started":"2024-02-28T10:04:13.934515Z","shell.execute_reply":"2024-02-28T10:04:13.957836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Also, retrive the old masks data frame\nmasks.drop(['ships'], axis=1, inplace=True)\nmasks.index+=1 \nmasks.head()","metadata":{"id":"DfJCvsCHmFrY","outputId":"43d5f845-b7a9-4c85-861d-4eb5dd059880","execution":{"iopub.status.busy":"2024-02-28T10:04:13.95972Z","iopub.execute_input":"2024-02-28T10:04:13.960069Z","iopub.status.idle":"2024-02-28T10:04:13.975685Z","shell.execute_reply.started":"2024-02-28T10:04:13.960044Z","shell.execute_reply":"2024-02-28T10:04:13.974824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Now, its the time to use the train_test_split.\n- Stratify to split the dataset into train and test sets in a way that preserves the same proportions of examples in each class as observed in the original dataset.","metadata":{"id":"zCJ9sHZgmFrY"}},{"cell_type":"code","source":"# Train - Test split\nfrom sklearn.model_selection import train_test_split                   \ntrain_ids, valid_ids = train_test_split(unique_img_ids, test_size = 0.3, stratify = unique_img_ids['ships'])","metadata":{"id":"bbQHbyzxmFrY","execution":{"iopub.status.busy":"2024-02-28T10:04:13.976734Z","iopub.execute_input":"2024-02-28T10:04:13.97703Z","iopub.status.idle":"2024-02-28T10:04:14.21835Z","shell.execute_reply.started":"2024-02-28T10:04:13.977007Z","shell.execute_reply":"2024-02-28T10:04:14.217234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create train data frame\ntrain_df = pd.merge(masks, train_ids)\n\n# Create test data frame\nvalid_df = pd.merge(masks, valid_ids)","metadata":{"id":"mObH089jmFrZ","execution":{"iopub.status.busy":"2024-02-28T10:04:14.219592Z","iopub.execute_input":"2024-02-28T10:04:14.220097Z","iopub.status.idle":"2024-02-28T10:04:14.485678Z","shell.execute_reply.started":"2024-02-28T10:04:14.220071Z","shell.execute_reply":"2024-02-28T10:04:14.484888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:39:09.878333Z","iopub.execute_input":"2024-02-28T10:39:09.878724Z","iopub.status.idle":"2024-02-28T10:39:09.893616Z","shell.execute_reply.started":"2024-02-28T10:39:09.878695Z","shell.execute_reply":"2024-02-28T10:39:09.892709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"There are ~\")\nprint(train_df.shape[0], 'training masks,')\nprint(valid_df.shape[0], 'validation masks.')","metadata":{"id":"5Po1UEcGmFrZ","outputId":"0bbb1cb4-603c-43d2-97a6-a2647072d42a","execution":{"iopub.status.busy":"2024-02-28T10:04:14.486973Z","iopub.execute_input":"2024-02-28T10:04:14.487951Z","iopub.status.idle":"2024-02-28T10:04:14.493343Z","shell.execute_reply.started":"2024-02-28T10:04:14.487915Z","shell.execute_reply":"2024-02-28T10:04:14.492468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.ships.tail()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:43:41.053271Z","iopub.execute_input":"2024-02-28T10:43:41.054015Z","iopub.status.idle":"2024-02-28T10:43:41.060439Z","shell.execute_reply.started":"2024-02-28T10:43:41.053984Z","shell.execute_reply":"2024-02-28T10:43:41.059615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualise the ship counts\nplt.figure(figsize=(10, 6))\nsns.countplot(train_df.ships)\nplt.show()","metadata":{"id":"hGt5Qns4mFrZ","outputId":"19bbeba3-2baf-450e-d908-3770ea8b439f","execution":{"iopub.status.busy":"2024-02-28T10:50:17.513403Z","iopub.execute_input":"2024-02-28T10:50:17.513818Z","iopub.status.idle":"2024-02-28T10:50:17.654351Z","shell.execute_reply.started":"2024-02-28T10:50:17.513782Z","shell.execute_reply":"2024-02-28T10:50:17.653453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"colors = ['red', 'green', 'blue', 'orange', 'purple']\nplt.bar(ship_counts.index, ship_counts.values, color=colors[:len(ship_counts)])\n\nplt.xlabel(\"Ships in an image\")\nplt.ylabel(\"Count\")\nplt.title(\"Ship Counts\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-02-28T10:48:18.968399Z","iopub.execute_input":"2024-02-28T10:48:18.968782Z","iopub.status.idle":"2024-02-28T10:48:19.189512Z","shell.execute_reply.started":"2024-02-28T10:48:18.968741Z","shell.execute_reply":"2024-02-28T10:48:19.188633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Oh!! Very huge imbalance in the data...\n- We need to have a better balanced data!!!","metadata":{"id":"gISXjCmZmFrd"}},{"cell_type":"markdown","source":"## Random Undersampling to generate a better balanced data to work with","metadata":{"id":"wATcP3FVmFre"}},{"cell_type":"code","source":"# Clipping the max value of grouped_ship_count to be 7, minimum to be 0\ntrain_df['grouped_ship_count'] = train_df.ships.map(lambda x: (x+1)//2).clip(0,7)","metadata":{"id":"GW1DYwrfmFre","execution":{"iopub.status.busy":"2024-02-28T10:04:14.762317Z","iopub.execute_input":"2024-02-28T10:04:14.762575Z","iopub.status.idle":"2024-02-28T10:04:14.882249Z","shell.execute_reply.started":"2024-02-28T10:04:14.762553Z","shell.execute_reply":"2024-02-28T10:04:14.881448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check\ntrain_df.grouped_ship_count.value_counts()","metadata":{"id":"QHic4kjdmFre","outputId":"900177b5-aff2-4aba-ab08-fc1124e93356","execution":{"iopub.status.busy":"2024-02-28T10:04:14.88452Z","iopub.execute_input":"2024-02-28T10:04:14.885208Z","iopub.status.idle":"2024-02-28T10:04:14.89472Z","shell.execute_reply.started":"2024-02-28T10:04:14.885171Z","shell.execute_reply":"2024-02-28T10:04:14.893821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Difference between head(n) and sample(n) in pandas\n- df.head(n) returns only top n data from the df\n- df.sample(n) returns random n data from the df","metadata":{"execution":{"iopub.status.busy":"2022-10-07T10:50:30.915451Z","iopub.execute_input":"2022-10-07T10:50:30.916382Z","iopub.status.idle":"2022-10-07T10:50:30.941993Z","shell.execute_reply.started":"2022-10-07T10:50:30.91634Z","shell.execute_reply":"2022-10-07T10:50:30.940812Z"},"id":"75TVsB85mFre"}},{"cell_type":"code","source":"# Top 10 data\ntrain_df.head(10)","metadata":{"id":"pnr2H5_WmFre","outputId":"c657d012-f519-44f4-ae97-e46928698014","execution":{"iopub.status.busy":"2024-02-28T10:04:14.895958Z","iopub.execute_input":"2024-02-28T10:04:14.896246Z","iopub.status.idle":"2024-02-28T10:04:14.912039Z","shell.execute_reply.started":"2024-02-28T10:04:14.896223Z","shell.execute_reply":"2024-02-28T10:04:14.911129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Random 10 data\ntrain_df.sample(10)","metadata":{"id":"lHfspu_7mFre","outputId":"8da9cdfb-1932-472d-c844-78a5e427dbba","execution":{"iopub.status.busy":"2024-02-28T10:04:14.913169Z","iopub.execute_input":"2024-02-28T10:04:14.913442Z","iopub.status.idle":"2024-02-28T10:04:14.932169Z","shell.execute_reply.started":"2024-02-28T10:04:14.91342Z","shell.execute_reply":"2024-02-28T10:04:14.931324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Random Under-Sampling ships\ndef sample_ships(in_df, base_rep_val=1500):\n    '''\n    Input Args:\n    in_df - dataframe we want to apply this function\n    base_val - random sample of this value to be taken from the data frame\n    '''\n    if in_df['ships'].values[0]==0:                                                 \n        return in_df.sample(base_rep_val//3)  # Random 1500//3 = 500 samples taken whose ship count is 0 in an image \n    else:                                 \n        return in_df.sample(base_rep_val)    # Random 1500 samples taken whose ship count is not 0 in an image","metadata":{"id":"yvk146BImFre","execution":{"iopub.status.busy":"2024-02-28T10:04:14.933347Z","iopub.execute_input":"2024-02-28T10:04:14.933629Z","iopub.status.idle":"2024-02-28T10:04:14.941265Z","shell.execute_reply.started":"2024-02-28T10:04:14.933606Z","shell.execute_reply":"2024-02-28T10:04:14.940422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating groups of ship counts and applying the sample_ships functions to randomly undersample the ships\nbalanced_train_df = train_df.groupby('grouped_ship_count').apply(sample_ships)\nbalanced_train_df.grouped_ship_count.value_counts() # In each group we have total of 1500 ships except 0 as we have decreased it even more to 500","metadata":{"id":"Sh2SgP5lmFrf","outputId":"52adb1b2-7c42-4e5a-9275-dc61f8aee600","execution":{"iopub.status.busy":"2024-02-28T10:04:14.942187Z","iopub.execute_input":"2024-02-28T10:04:14.942433Z","iopub.status.idle":"2024-02-28T10:04:14.983871Z","shell.execute_reply.started":"2024-02-28T10:04:14.942411Z","shell.execute_reply":"2024-02-28T10:04:14.982975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Explaining what we just did if still not clear\nfor i in range(8):\n    df_val_counts = balanced_train_df[balanced_train_df.grouped_ship_count==i].ships.value_counts()\n    print(f\"Data frame for grouped ship count = {i}:-\\n{df_val_counts}\\nSum of Values:- {df_val_counts.values.sum()}\\n\\n\")\n","metadata":{"id":"6SERf32pmFrf","outputId":"612c6483-a49f-4cf1-d6c2-cd3a5486bfb6","execution":{"iopub.status.busy":"2024-02-28T10:04:14.985192Z","iopub.execute_input":"2024-02-28T10:04:14.985543Z","iopub.status.idle":"2024-02-28T10:04:15.002945Z","shell.execute_reply.started":"2024-02-28T10:04:14.985511Z","shell.execute_reply":"2024-02-28T10:04:15.002088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15, 5))\nplt.suptitle(\"Train Data\", fontsize=18, color='r', weight='bold')\n\nplt.subplot(1, 2, 1)\nship_counts_before = train_df['ships'].value_counts()\nplt.bar(ship_counts_before.index, ship_counts_before.values, color='lightblue')  # Change color here\nplt.title(\"Ship Counts - Before Balancing\", color='m', fontsize=15)\nplt.ylabel(\"Count\", color='tab:pink', fontsize=13)\nplt.xlabel(\"# Ships in an image\", color='tab:pink', fontsize=13)\n\nplt.subplot(1, 2, 2)\nship_counts_after = balanced_train_df['ships'].value_counts()\nplt.bar(ship_counts_after.index, ship_counts_after.values, color='lightgreen')  # Change color here\nplt.title(\"Ship Counts - After Balancing\", color='m', fontsize=15)\nplt.xlabel(\"# Ships in an image\", color='tab:pink', fontsize=13)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-02-28T11:08:10.861654Z","iopub.execute_input":"2024-02-28T11:08:10.862468Z","iopub.status.idle":"2024-02-28T11:08:11.333215Z","shell.execute_reply.started":"2024-02-28T11:08:10.862437Z","shell.execute_reply":"2024-02-28T11:08:11.33231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Parameters\nBATCH_SIZE = 4                 # Train batch size\nEDGE_CROP = 16                 # While building the model\nNB_EPOCHS = 5                  # Training epochs\nGAUSSIAN_NOISE = 0.1           # To be used in a layer in the model\nUPSAMPLE_MODE = 'SIMPLE'       # SIMPLE ==> UpSampling2D, else Conv2DTranspose\nNET_SCALING = None             # Downsampling inside the network                        \nIMG_SCALING = (1, 1)           # Downsampling in preprocessing\nVALID_IMG_COUNT = 400          # Valid batch size\nMAX_TRAIN_STEPS = 200          # Maximum number of steps_per_epoch in training","metadata":{"id":"YnVyTCqUmFrf","execution":{"iopub.status.busy":"2024-02-28T10:04:15.011928Z","iopub.execute_input":"2024-02-28T10:04:15.012188Z","iopub.status.idle":"2024-02-28T10:04:15.022034Z","shell.execute_reply.started":"2024-02-28T10:04:15.012165Z","shell.execute_reply":"2024-02-28T10:04:15.021064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Image and Mask Generator\ndef make_image_gen(in_df, batch_size = BATCH_SIZE):\n    '''\n    Inputs -\n    in_df - data frame on which the function will be applied\n    batch_size - number of training examples in one iteration\n    '''\n    all_batches = list(in_df.groupby('ImageId'))                             # Group ImageIds and create list of that dataframe\n    out_rgb = []                                                             # Image list\n    out_mask = []                                                            # Mask list\n    while True:                                                              # Loop for every data\n        np.random.shuffle(all_batches)                                       # Shuffling the data\n        for c_img_id, c_masks in all_batches:                                # For img_id and msk_rle in all_batches\n            rgb_path = os.path.join(train_image_dir, c_img_id)               # Get the img path\n            c_img = imread(rgb_path)                                         # img array\n            c_mask = masks_as_image(c_masks['EncodedPixels'].values)         # Create mask of rle data for each ship in an img\n            out_rgb += [c_img]                                               # Append the current img in the out_rgb / img list\n            out_mask += [c_mask]                                             # Append the current mask in the out_mask / mask list\n            if len(out_rgb)>=batch_size:                                     # If length of list is more or equal to batch size then\n                yield np.stack(out_rgb)/255.0, np.stack(out_mask)            # Yeild the scaled img array (b/w 0 and 1) and mask array (0 for bg and 1 for ship)\n                out_rgb, out_mask=[], []                                     # Empty the lists to create another batch","metadata":{"id":"_l69s0W_mFrf","execution":{"iopub.status.busy":"2024-02-28T10:04:15.023238Z","iopub.execute_input":"2024-02-28T10:04:15.023564Z","iopub.status.idle":"2024-02-28T10:04:15.033549Z","shell.execute_reply.started":"2024-02-28T10:04:15.023534Z","shell.execute_reply":"2024-02-28T10:04:15.032676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Generate train data \ntrain_gen = make_image_gen(balanced_train_df)\n\n# Image and Mask\ntrain_x, train_y = next(train_gen)\n\n# Print the summary\nprint(f\"train_x ~\\nShape: {train_x.shape}\\nMin value: {train_x.min()}\\nMax value: {train_x.max()}\")\nprint(f\"\\ntrain_y ~\\nShape: {train_y.shape}\\nMin value: {train_y.min()}\\nMax value: {train_y.max()}\")","metadata":{"id":"LrGDDsHvmFrf","outputId":"df2fec31-16f5-4777-ade6-c3e79a8eabf6","execution":{"iopub.status.busy":"2024-02-28T10:04:15.034716Z","iopub.execute_input":"2024-02-28T10:04:15.035067Z","iopub.status.idle":"2024-02-28T10:04:15.766626Z","shell.execute_reply.started":"2024-02-28T10:04:15.035028Z","shell.execute_reply":"2024-02-28T10:04:15.765688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visulaising train batch\n'''montage_rgb = lambda x: np.stack([montage(x[:, :, :, i]) for i in range(x.shape[3])], -1)\nbatch_rgb = montage_rgb(train_x)                                                   # Create montage of img\nbatch_seg = montage(train_y[:, :, :, 0])                                           # Create montafe of msk\nbatch_overlap = mark_boundaries(batch_rgb, batch_seg.astype(int))                  # Create bounding box around ships in img\ntitles = [\"Images\", \"Segmentations\", \"Bounding Boxes on ships in Images\"]          # Titles for subplot\ncolors = ['g', 'm', 'b']                                                           # Colors to be used for title\ndisplay = [batch_rgb, batch_seg, batch_overlap]                                    # What to display in subplot\nplt.figure(figsize=(25,10))                                                        # Generate figure \nfor i in range(3):                                                                 # For i = 0, 1, 2, 3                           \n    plt.subplot(1, 3, i+1)                                                         # Create subplot\n    plt.imshow(display[i])                                                         # Display \n    plt.title(titles[i], fontsize = 18, color = colors[i])                         # Title \n    plt.axis('off')                                                                # Turn off the axis\nplt.suptitle(\"Batch Visualizations\", fontsize = 20, color = 'r', weight = 'bold')  # Add suptitle\nplt.tight_layout()'''                                                                 # Layout for subplot","metadata":{"id":"YeeCgg_PmFrf","outputId":"dca020dc-5290-4100-b206-6f277a1e0506","execution":{"iopub.status.busy":"2024-02-28T10:04:15.7678Z","iopub.execute_input":"2024-02-28T10:04:15.768153Z","iopub.status.idle":"2024-02-28T10:04:15.775384Z","shell.execute_reply.started":"2024-02-28T10:04:15.768121Z","shell.execute_reply":"2024-02-28T10:04:15.774551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Prepare validation data\nvalid_x, valid_y = next(make_image_gen(valid_df, VALID_IMG_COUNT))\nprint(f\"valid_x ~\\nShape: {valid_x.shape}\\nMin value: {valid_x.min()}\\nMax value: {valid_x.max()}\")\nprint(f\"\\nvalid_y ~\\nShape: {valid_y.shape}\\nMin value: {valid_y.min()}\\nMax value: {valid_y.max()}\")","metadata":{"id":"aUXAhwWKmFrg","outputId":"4c4f0973-9452-495c-b696-b0dbe07debcd","execution":{"iopub.status.busy":"2024-02-28T10:04:15.776326Z","iopub.execute_input":"2024-02-28T10:04:15.776568Z","iopub.status.idle":"2024-02-28T10:04:25.105183Z","shell.execute_reply.started":"2024-02-28T10:04:15.776547Z","shell.execute_reply":"2024-02-28T10:04:25.104273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Augmenting Data using ImageDataGenerator\nfrom keras.preprocessing.image import ImageDataGenerator\n\n# Preparing image data generator arguments\ndg_args = dict(rotation_range = 15,            # Degree range for random rotations\n               horizontal_flip = True,         # Randomly flips the inputs horizontally\n               vertical_flip = True,           # Randomly flips the inputs vertically\n               data_format = 'channels_last')  # channels_last refer to (batch, height, width, channels)","metadata":{"id":"Ru71AFFwmFrg","execution":{"iopub.status.busy":"2024-02-28T10:04:25.106183Z","iopub.execute_input":"2024-02-28T10:04:25.106453Z","iopub.status.idle":"2024-02-28T10:04:36.571848Z","shell.execute_reply.started":"2024-02-28T10:04:25.10643Z","shell.execute_reply":"2024-02-28T10:04:36.570897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Click [here](https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/image/ImageDataGenerator) to find more information on ImageDataGenerator and its arguments.","metadata":{"id":"SZcUkxocmFrg"}},{"cell_type":"code","source":"image_gen = ImageDataGenerator(**dg_args)\nlabel_gen = ImageDataGenerator(**dg_args)\n\ndef create_aug_gen(in_gen, seed = None):\n    '''\n    Takes in -\n    in_gen - train data generator, seed value\n    '''\n    np.random.seed(seed if seed is not None else np.random.choice(range(9999)))  # Randomly assign seed value if not provided\n    for in_x, in_y in in_gen:                                                    # For imgs and msks in train data generator\n        seed = 12                                                                # Seed value for imgs and msks must be same else augmentation won't be same\n        \n        # Create augmented imgs\n        g_x = image_gen.flow(255*in_x,                                           # Inverse scaling on imgs for augmentation                                       \n                             batch_size = in_x.shape[0],                         # batch_size = 3\n                             seed = seed,                                        # Seed\n                             shuffle=True)                                       # Shuffle the data\n        \n        # Create augmented masks\n        g_y = label_gen.flow(in_y,\n                             batch_size = in_x.shape[0],                       \n                             seed = seed,                                         \n                             shuffle=True)                                       \n        \n        '''Yeilds - augmented scaled imgs and msks array'''\n        yield next(g_x)/255.0, next(g_y)","metadata":{"id":"HQvlfWvjmFrg","execution":{"iopub.status.busy":"2024-02-28T10:04:36.573271Z","iopub.execute_input":"2024-02-28T10:04:36.57385Z","iopub.status.idle":"2024-02-28T10:04:36.581438Z","shell.execute_reply.started":"2024-02-28T10:04:36.573822Z","shell.execute_reply":"2024-02-28T10:04:36.580592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Augment the train data\ncur_gen = create_aug_gen(train_gen, seed = 42)\nt_x, t_y = next(cur_gen)\nprint('x', t_x.shape, t_x.dtype, t_x.min(), t_x.max())\nprint('y', t_y.shape, t_y.dtype, t_y.min(), t_y.max())","metadata":{"id":"YChZ_zOqmFrg","outputId":"a941718e-7822-48f5-ee72-d37771ecec6a","execution":{"iopub.status.busy":"2024-02-28T10:04:36.582526Z","iopub.execute_input":"2024-02-28T10:04:36.582864Z","iopub.status.idle":"2024-02-28T10:04:36.818005Z","shell.execute_reply.started":"2024-02-28T10:04:36.582839Z","shell.execute_reply":"2024-02-28T10:04:36.817085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Final display before passing data into model\n'''fig, (ax1, ax2, ax3) = plt.subplots(1, 3, figsize = (25, 10))\nax1.imshow(montage_rgb(t_x), cmap='gray')\nax1.set_title('Images', fontsize = 18, color = 'g')\nax1.axis('off')\nax2.imshow(montage(t_y[:, :, :, 0]), cmap='Blues_r')\nax2.set_title('Masks', fontsize = 18, color = 'r')\nax2.axis('off')\nax3.imshow(mark_boundaries(montage_rgb(t_x), montage(t_y[:, :, :, 0].astype(int))))\nax3.set_title('Bounding Box', fontsize = 18, color = 'b')\nax3.axis('off')\nplt.tight_layout()'''","metadata":{"id":"Tg7MNJdOmFrg","outputId":"c8c40493-674e-49d1-e9fa-bd281b33470c","execution":{"iopub.status.busy":"2024-02-28T10:04:36.819571Z","iopub.execute_input":"2024-02-28T10:04:36.819937Z","iopub.status.idle":"2024-02-28T10:04:36.826261Z","shell.execute_reply.started":"2024-02-28T10:04:36.819906Z","shell.execute_reply":"2024-02-28T10:04:36.82534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gc.collect() # Block all the garbage that has been generated","metadata":{"id":"H2GlhnR-mFrh","outputId":"0a87d09b-5b3a-45d2-fa71-7afdf43b8721","execution":{"iopub.status.busy":"2024-02-28T10:04:36.82737Z","iopub.execute_input":"2024-02-28T10:04:36.827695Z","iopub.status.idle":"2024-02-28T10:04:37.102568Z","shell.execute_reply.started":"2024-02-28T10:04:36.827664Z","shell.execute_reply":"2024-02-28T10:04:37.101577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## A brief introduction on U-NET architecture <a class=\"anchor\"  id=\"h6\"></a>\n\n<img src=\"https://miro.medium.com/max/720/1*f7YOaE4TWubwaFF7Z1fzNw.png\" width = 60%>\n\n- The name U-NET itself is due to the shape of its architecture.\n- Each blue box corresponds to a multi-channel feature map.\n- The number of channels are denoted on top of the box.\n- The x-y size is provided at the lower left edge of the box.\n- The arrows shows the respective operations as mentioned on the bottom right of the image.\n- This architecture consists of three sections: The contraction, The bottleneck, and the expansion section. \n- But the heart of this architecture lies in the expansion section. \n- This action would ensure that the features that are learned while contracting the image will be used to reconstruct it.\n\n*For in-depth understanding do read this amazing [line by line explanation](https://towardsdatascience.com/unet-line-by-line-explanation-9b191c76baf5).*\n\n## Building and Training U-NET from Scratch <a class=\"anchor\"  id=\"h7\"></a>\n\n<img src = \"https://i.stack.imgur.com/o5TBk.png\" width = 45%>       \n\n<img src = \"https://www.researchgate.net/publication/333593451/figure/fig2/AS:765890261966848@1559613876098/Illustration-of-Max-Pooling-and-Average-Pooling-Figure-2-above-shows-an-example-of-max.png\" width = 45% height = 30%>","metadata":{"id":"bzJgfZStmFrh"}},{"cell_type":"code","source":"# Build U-Net model\nfrom keras import models, layers\n\n# Conv2DTranspose upsampling\ndef upsample_conv(filters, kernel_size, strides, padding):\n    return layers.Conv2DTranspose(filters, kernel_size, strides=strides, padding=padding)\n# Upsampling without Conv2DTranspose\ndef upsample_simple(filters, kernel_size, strides, padding):\n    return layers.UpSampling2D(strides)\n\n# Upsampling method choice\nif UPSAMPLE_MODE=='DECONV':\n    upsample=upsample_conv\nelse:\n    upsample=upsample_simple\n\n# Building the layers of UNET\ninput_img = layers.Input(t_x.shape[1:], name = 'RGB_Input')\npp_in_layer = input_img\n\n# If NET_SCALING is defined then do the next step else continue ahead\nif NET_SCALING is not None:\n    pp_in_layer = layers.AvgPool2D(NET_SCALING)(pp_in_layer)\n\n# To avoid overfitting and fastening the process of training\npp_in_layer = layers.GaussianNoise(GAUSSIAN_NOISE)(pp_in_layer)                       # Useful to mitigate overfitting\npp_in_layer = layers.BatchNormalization()(pp_in_layer)                                # Allows using higher learning rate without causing problems with gradients\n\n\n## Downsample (C-->C-->MP)\n\nc1 = layers.Conv2D(8, (3, 3), activation='relu', padding='same') (pp_in_layer)\nc1 = layers.Conv2D(8, (3, 3), activation='relu', padding='same') (c1)\np1 = layers.MaxPooling2D((2, 2)) (c1)\n\nc2 = layers.Conv2D(16, (3, 3), activation='relu', padding='same') (p1)\nc2 = layers.Conv2D(16, (3, 3), activation='relu', padding='same') (c2)\np2 = layers.MaxPooling2D((2, 2)) (c2)\n\nc3 = layers.Conv2D(32, (3, 3), activation='relu', padding='same') (p2)\nc3 = layers.Conv2D(32, (3, 3), activation='relu', padding='same') (c3)\np3 = layers.MaxPooling2D((2, 2)) (c3)\n\nc4 = layers.Conv2D(64, (3, 3), activation='relu', padding='same') (p3)\nc4 = layers.Conv2D(64, (3, 3), activation='relu', padding='same') (c4)\np4 = layers.MaxPooling2D(pool_size=(2, 2)) (c4)\n\n\nc5 = layers.Conv2D(128, (3, 3), activation='relu', padding='same') (p4)\nc5 = layers.Conv2D(128, (3, 3), activation='relu', padding='same') (c5)\n\n## Upsample (U --> Concat --> C --> C)\n\nu6 = upsample(64, (2, 2), strides=(2, 2), padding='same') (c5)\nu6 = layers.concatenate([u6, c4])\nc6 = layers.Conv2D(64, (3, 3), activation='relu', padding='same') (u6)\nc6 = layers.Conv2D(64, (3, 3), activation='relu', padding='same') (c6)\n\nu7 = upsample(32, (2, 2), strides=(2, 2), padding='same') (c6)\nu7 = layers.concatenate([u7, c3])\nc7 = layers.Conv2D(32, (3, 3), activation='relu', padding='same') (u7)\nc7 = layers.Conv2D(32, (3, 3), activation='relu', padding='same') (c7)\n\nu8 = upsample(16, (2, 2), strides=(2, 2), padding='same') (c7)\nu8 = layers.concatenate([u8, c2])\nc8 = layers.Conv2D(16, (3, 3), activation='relu', padding='same') (u8)\nc8 = layers.Conv2D(16, (3, 3), activation='relu', padding='same') (c8)\n\nu9 = upsample(8, (2, 2), strides=(2, 2), padding='same') (c8)\nu9 = layers.concatenate([u9, c1], axis=3)\nc9 = layers.Conv2D(8, (3, 3), activation='relu', padding='same') (u9)\nc9 = layers.Conv2D(8, (3, 3), activation='relu', padding='same') (c9)\n\nd = layers.Conv2D(1, (1, 1), activation='sigmoid') (c9)\nd = layers.Cropping2D((EDGE_CROP, EDGE_CROP))(d)\nd = layers.ZeroPadding2D((EDGE_CROP, EDGE_CROP))(d)\n\nif NET_SCALING is not None:\n    d = layers.UpSampling2D(NET_SCALING)(d)\n\nseg_model = models.Model(inputs=[input_img], outputs=[d])\n\nseg_model.summary()","metadata":{"id":"Wb1YIcPTmFrh","outputId":"3b67adaa-872e-4943-c435-79b23331b253","execution":{"iopub.status.busy":"2024-02-28T10:04:37.103968Z","iopub.execute_input":"2024-02-28T10:04:37.104647Z","iopub.status.idle":"2024-02-28T10:04:38.16967Z","shell.execute_reply.started":"2024-02-28T10:04:37.104613Z","shell.execute_reply":"2024-02-28T10:04:38.168807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Calculating the output shape of the feature map depending on strides, kernel size, input size, padding\n\n<img src = \"https://i.stack.imgur.com/qPKxm.png\" width = 45%>\n\n#### Metric and Loss for compiling the model\n\n**Dice Coeffiient:-**        \n<img src = \"https://cdn-images-1.medium.com/max/1600/0*HuENmnLgplFLg7Xv\" width = 45%>\n\n<img src = \"https://drive.google.com/uc?id=1XoP-1kwuScIj2Ee7rYrARylF8DRyJDF6\" width = 45%>       \n\n<img src = \"https://pbs.twimg.com/media/FBmVmdHWQAAU7gq.png\" width = 45% height = 20%>\n\n<img src = \"https://drive.google.com/uc?id=1oFWisqT_z0AKXvp1-JQ3LSjcwWrxjz2J\" width = 45%>\n\n\n*[Here](https://arxiv.org/pdf/2006.14822.pdf), you can find a survey of Loss functions for semantic segmentations.*        \n*More info can be found [here](https://www2.cs.sfu.ca/~hamarneh/ecopy/cmig2019.pdf)*","metadata":{"id":"PDgP-RPCmFrh"}},{"cell_type":"code","source":"# Compute dice coefficient, loss with BCE and compile the model\nimport tensorflow as tf\nimport keras.backend as K\nfrom tensorflow.keras.optimizers import Adam\nfrom keras.losses import binary_crossentropy\n\n# Dice coeff\ndef dice_coef(y_true, y_pred, smooth=1):\n    intersection = K.sum(y_true * y_pred, axis=[1,2,3]  )                         # int = y_true ∩ y_pred\n    union = K.sum(y_true, axis=[1,2,3]) + K.sum(y_pred, axis=[1,2,3])           # un = y_true_flatten ed ∪ y_pred_flattened\n    return K.mean( (2. * intersection + smooth) / (union + smooth), axis=0)     # dice = 2 * int + 1 / un + 1\n\n# Dice with BCE\ndef dice_p_bce(y_true, y_pred):\n    '''\n    Compute this function based on the explanation\n    - use alpha = 1e-3\n    '''\n    \n    combo_loss = \"Something\"\n         \n    return combo_loss\n\n# Compile the model\nseg_model.compile(optimizer=tf.keras.optimizers.legacy.Adam(1e-4, decay=1e-6), loss=dice_p_bce, metrics=[dice_coef])","metadata":{"id":"esAWI3D_mFrh","execution":{"iopub.status.busy":"2024-02-28T11:20:27.980051Z","iopub.execute_input":"2024-02-28T11:20:27.980757Z","iopub.status.idle":"2024-02-28T11:20:27.995386Z","shell.execute_reply.started":"2024-02-28T11:20:27.980725Z","shell.execute_reply":"2024-02-28T11:20:27.994539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Preparing Callbacks \nfrom keras.callbacks import ModelCheckpoint, LearningRateScheduler, EarlyStopping, ReduceLROnPlateau\n\n# Best model weights\nweight_path=\"{}_weights.best.hdf5\".format('seg_model')\n\n# Monitor validation dice coeff and save the best model weights\ncheckpoint = ModelCheckpoint(weight_path, monitor='val_dice_coef', verbose=1, \n                             save_best_only=True, mode='max', save_weights_only = True)\n\n# Reduce Learning Rate on Plateau\nreduceLROnPlat = ReduceLROnPlateau(monitor='val_dice_coef', factor=0.5, \n                                   patience=3, \n                                   verbose=1, mode='max', epsilon=0.0001, cooldown=2, min_lr=1e-6)\n\n# Stop training once there is no improvement seen in the model\nearly = EarlyStopping(monitor=\"val_dice_coef\", \n                      mode=\"max\", \n                      patience=15) # probably needs to be more patient, but kaggle time is limited\n\n# Callbacks ready\ncallbacks_list = [checkpoint, early, reduceLROnPlat]","metadata":{"id":"Oy1k7BWEmFrh","execution":{"iopub.status.busy":"2024-02-28T11:20:32.679886Z","iopub.execute_input":"2024-02-28T11:20:32.680596Z","iopub.status.idle":"2024-02-28T11:20:32.6872Z","shell.execute_reply.started":"2024-02-28T11:20:32.680567Z","shell.execute_reply":"2024-02-28T11:20:32.686268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- *Just in case you are not aware about callbacks we use in keras you can learn more about it [here](https://keras.io/api/callbacks/).*","metadata":{"id":"qH2NrpmNs5zL"}},{"cell_type":"code","source":"# Finalizing steps per epoch\n'''step_count = min(MAX_TRAIN_STEPS, balanced_train_df.shape[0]//BATCH_SIZE)\n\n# Final augmented data being used in training\naug_gen = create_aug_gen(make_image_gen(balanced_train_df))\n\n# Save loss history while training\nloss_history = [seg_model.fit_generator(aug_gen, \n                             steps_per_epoch=step_count, \n                             epochs=NB_EPOCHS, \n                             validation_data=(valid_x, valid_y),\n                             callbacks=callbacks_list,\n                            workers=1)] '''","metadata":{"id":"4QNJtGrRmFrj","execution":{"iopub.status.busy":"2024-02-28T11:27:49.760184Z","iopub.execute_input":"2024-02-28T11:27:49.760886Z","iopub.status.idle":"2024-02-28T11:27:49.767282Z","shell.execute_reply.started":"2024-02-28T11:27:49.760854Z","shell.execute_reply":"2024-02-28T11:27:49.766284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Finalizing steps per epoch\n'''step_count = min(MAX_TRAIN_STEPS, balanced_train_df.shape[0] // BATCH_SIZE)\n\n# Final augmented data being used in training\naug_gen = create_aug_gen(make_image_gen(balanced_train_df))\n\n# Save loss history while training\nloss_history = seg_model.fit(\n    aug_gen,\n    steps_per_epoch=step_count,\n    epochs=NB_EPOCHS,\n    validation_data=(valid_x, valid_y),\n    callbacks=callbacks_list,\n    workers=1\n)\n'''","metadata":{"execution":{"iopub.status.busy":"2024-02-28T11:23:41.587076Z","iopub.execute_input":"2024-02-28T11:23:41.587795Z","iopub.status.idle":"2024-02-28T11:23:43.279104Z","shell.execute_reply.started":"2024-02-28T11:23:41.587747Z","shell.execute_reply":"2024-02-28T11:23:43.277672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save the weights to load it later for test data \nseg_model.load_weights(weight_path)\nseg_model.save('seg_model.h5')","metadata":{"id":"puTYhSN0mFrj","execution":{"iopub.status.busy":"2024-02-28T10:04:40.594921Z","iopub.status.idle":"2024-02-28T10:04:40.595246Z","shell.execute_reply.started":"2024-02-28T10:04:40.59509Z","shell.execute_reply":"2024-02-28T10:04:40.595103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Tasks for you <mark>(Complete these tasks and get certificate as a reward for your achievement)</mark><a class=\"anchor\"  id=\"h8\"></a>\n\n**We will be providing the link of this notebook in the description of this video.**\n- Download this notebook, compute the combo loss as discussed in the video and complete the training.\n- After training, write down your conclusions and observations.\n- Based on your conclusions you do the changes in this notebook in order to make your model more efficient.\n- Now apply this saved model on the test data.\n\n- **The submission notebook must contain:-**        \n    - *Implementation of combo loss (BCE Dice)*                              \n    - *Conclusions and observations after re-training this model based on your changes to the provided notebook*.                     \n    - *Applying this model to the test data and your conclusions based on the final output.* \n    \n***Submit your final notebook by clicking [here](https://forms.gle/gVW3Spv148dMzamo9).***\n\n---\n\n\n# THE END\n","metadata":{"id":"U0tL3mZemFrj"}}]}