{
  "id": 396178,
  "title": "Introduction To Semantic Segmentation",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/396178",
  "author_name": "",
  "post_date": "2023-03-20T16:53:32.080868700Z",
  "votes": 22,
  "comment_count": 2,
  "views": 0,
  "content": "<h3>Table of Contents</h3>\n<ol>\n<li>Introduction</li>\n<li>We need to understand Object Detection first!</li>\n<li>Slow R-CNN</li>\n<li>Fast R-CNN</li>\n<li>Faster R-CNN</li>\n<li>Finally… Semantic Segmentation!</li>\n<li>Mask R-CNN</li>\n<li>Where can I find segmentation models?</li>\n<li>Conclusion</li>\n</ol>\n<h3>Introduction</h3>\n<p>We will present an introduction to the semantic segmentation task. But before we understand semantic segmentation we first need to understand object detection as semantic segmentation is object detection \"with an extra step\".</p>\n<p>There are many different tasks in Computer Vision, some of which you may already be familiary with:</p>\n<p><img src=\"https://i.imgur.com/xeXdDYC.png\"></p>\n<p>For example, some Computer Vision tasks are:</p>\n<ul>\n<li><strong>Image classification</strong> consists of predicting whether a class is present or not in an image. </li>\n<li><strong>Multi-class classification:</strong> multi-class classification is the extension of the <strong>binary classification</strong> problem. In this type of problem each image can belong to only one category, but the amount of categories is more than two. This means that class labels or class membership are mutually exclusive. For example, predicting if the animal present in an image is a cat, a dog or a mouse.</li>\n<li><strong>Multi-label classification:</strong> in multi-label classification one single image can have multiple labels present. This means that class labels or class membership are not mutually exclusive. In multi-label classification, zero or more labels are required as output for each input sample. For example, predicting which animals are present in an image.</li>\n<li><strong>Object detection:</strong> object detection involves drawing a bounding box around one or more objects in an image. In this way, object detection does not only detect if a class is present in an image or not but it also tells where it is located providing a box surrounding the object.</li>\n<li><strong>Semantic segmentation</strong>: labels all parts of an image with a label. Each pixel in an image is assigned a label. The goal of semantic image segmentation is to label each pixel of an image with a corresponding class of what is being represented. Because we're predicting for every pixel in the image, this task is commonly referred to as dense prediction.</li>\n<li><strong>Instance segmentation:</strong> instance segmentation deals with detecting instances of objects and demarcating their boundaries. It is similar to semantic segmentation in the sense that it clearly outputs the object boundaries (irregular boundaries), but it is different in the sense that it does not segment all the image but rather the objects of interest. It is similar to object detection in the sense that it does not segment all the image, but it is different in the sense that it provides a more fine grained boundary than just a bounding box.</li>\n</ul>\n<h3>We need to understand Object Detection first!</h3>\n<p>As said before, semantic segmentation is like an \"extension\" of object detection. So let's get first a clear understanding of the object detection task.</p>\n<p><strong>What's the input and the output?</strong></p>\n<ul>\n<li><strong>Input:</strong> Single RGB image. The shape is then <code>(height, width, channels)</code>.</li>\n<li><strong>Output:</strong> a set of detected objects. For each object predict:<ol>\n<li>Category label, from a fixed known set of categories, like \"cat\", \"dog\", \"mouse\", etc.</li>\n<li>Bounding box (four numbers: x, y, width, height).</li></ol></li>\n</ul>\n<p><strong>What is the notation for bounding boxes?</strong></p>\n<p>These are some common types of notations:</p>\n<ol>\n<li><strong>Upper left corner to lower right corner:</strong> we provide four numbers <code>(x1, y1, x2, y2)</code> where <code>(x1, y1)</code> correspond to the upper left corner coordinates of the box and <code>(x2, y2)</code> correspond to the lower right corner coordinates of the box.</li>\n<li><strong>Center box:</strong> we provide four numbers <code>(x, y, width, height)</code> where <code>(x, y)</code> are the coordinates of the center of the box and the height and width of the box.</li>\n<li><strong>COCO format:</strong> A bounding box is described by the pixel coordinate <code>(x_min, y_min)</code> of its lower left corner within the image together with its width and height in pixels. </li>\n</ol>\n<h3>Detecting a single object</h3>\n<p>What would be the approach? We can do it with a simple NN architecture which uses common backbones (like VGG, AlexNet, ResNet, etc) and attach to them:</p>\n<ul>\n<li>One branch (<strong>\"what\" branch</strong>) of layers that handle the classification problem which consists of some fully connected layers plus a softmax function (to classify which of the possible categories are present) with a Softmax Loss. </li>\n<li>One other branch (<strong>\"where\" branch</strong>) of layers that handle the localization problem which consists of some fully connected layers which end up in 4 nodes (x, y, width, height) with a regression loss, like L2 Loss.</li>\n</ul>\n<p>Since we have two loss functions and we need only one single loss to compute gradient descent we usually just sum up these two loss functions as a weighted sum. We do a weighted sum so that the weights of each loss can be learned and thus one problem does not overcomes the other.</p>\n<h3>Detecting multiple objects: a more challenging challenge</h3>\n<p>Images can have more than one object! What do we do? If we must detect multiple objects in an image we don't know beforehand how many objects will be present and the same object can appear multiple times in the same image too! For example, in a single image:</p>\n<ul>\n<li>One object can appear one or more times (10 dogs in an image).</li>\n<li>Multiple objects can appear and more than one time (2 dogs and 3 cats).</li>\n<li>The objects can appear in different regions.</li>\n</ul>\n<p>One approach to deal with this problem is the <strong>sliding window</strong>, which consists in cropping or dividing the input image in multiple crops. For each crop we apply the CNN and classify each crop as object or background. The obvious problems that rise with this approach are:</p>\n<ul>\n<li><strong>Which should be the shape of the bounding box?</strong></li>\n<li><strong>How many possible boxes are there in an image of size HxW?</strong> There are many! For each image we need to consider all the possible box sizes and ratios which yield the next formula: <br>\n$$ Total:possible:boxes = \\frac{H(H+1)}{2}\\frac{W(W+1)}{2} $$</li>\n</ul>\n<p>If we have a 800x600 image we have ~58M boxes! No way we can evaluate all of them!</p>\n<p>One solution to this problem are <strong>Region Proposal</strong> mechanisms.</p>\n<ul>\n<li>Find a small set of boxes that are likely to cover all objects.</li>\n<li>Often based on heuristics: e.g. look for \"blob-like\" image regions that have high probabilites of containing objects.</li>\n<li>Relatively fast to run: e.g. Selective Search gives ~2000 region proposals in a few seconds on CPU.</li>\n<li>They will end up being replaced by CNN (much faster approach and with learnable parameters).</li>\n<li>A proposal method gives us <strong>regions of interest</strong> (RoI)</li>\n</ul>\n<h3>R-CNN: Region-based CNN (usually nowadays called \"Slow\" R-CNN)</h3>\n<ol>\n<li>We start with an input image and a Region Proposal mechanism which gives us ~2k regions of interest (RoI).</li>\n<li>For each proposed region of interest we warp/resize each region into a fixed size region since the RoI can be of variable heights and widths. The warp size depends on the backbone architecure we use for classifying each warp. Some CNNs require input images to be 224x224 for example.</li>\n<li>For each warped region we run it through a CNN and classify each region. In this way, one image will yield many regions which will be passed through the CNN.</li>\n<li>Once we get a score for each region we try to select a subset of region proposals to output. To explain this better, let's say we have a cat and a dog in an image, but the model outputs ~2000 predictions for all the RoIs, we need a way to summarize all the predicitons so that we just output one single box for the dog and one for the cat. Many choices here: threshold on background or per-category, take top K proposals per image, etc.</li>\n<li>Compare with ground-truth boxes.</li>\n</ol>\n<p><img src=\"https://i.imgur.com/9yLr02y.png\"></p>\n<p>Problems:</p>\n<ul>\n<li>What happen if the proposed regions do not match well to our objects? To overcome this problem, in addition to classifying the class of the region we also perform a bounding box regression which tries to predict or \"transform\" the correct RoI's x, y, height and width. This way we iteratively try to get better RoIs which help us better detect the object in our image.</li>\n</ul>\n<h3>Fast R-CNN</h3>\n<p>As we can see, the R-CNN needs to do ~2000 forward passes for each image since we had ~2000 RoI. This is very expensive computationally and thus slow. The solution that people has come up with is running the CNN <strong>before</strong> warping. This allows us to reshare a lot of computation across image regions. </p>\n<p>This new solution is usually called <strong>Fast R-CNN</strong>. It basically runs the whole input image through a CNN and outputs a feature map or also called image features. This CNN we use is usually called the \"backbone\" network (like AlexNet, VGG, ResNet, etc). Then, the regions of interest are applied to the image features. Next, we warp/crop/resize the RoIs and input them into a light-weight CNN that outputs categories and box transforms per crop.</p>\n<p>This is much much faster since most of the computation happens in the backbone network (this saves work for overlapping region proposals) and the per-region network is relatively light-weight.</p>\n<p>Fast R-CNN can be up to 10 times faster than Slow R-CNN at training time and up to 20 times faster at test time.</p>\n<p><img src=\"https://i.imgur.com/mQeWzlJ.png\"></p>\n<h3>Faster R-CNN: learning region proposals.</h3>\n<p>Even though that Fast R-CNN is much faster than Slow R-CNN most of the compute time is dominated by region proposals! Region proposals are computed by an heuristic (Selective search algorithm) which runs on CPU. We could learn region proposals with a CNN instead!</p>\n<p>Faster R-CNN inserts a <strong>Region Proposal Network (RPN)</strong> to predict proposals from features. Otherwise, it is the same as Fast R-CNN. RPNs use \"anchor boxes\" of fixed size at each point in the feature map. For each anchor box we predict whether the corresponding anchor contains an object or not and a box transform to regress from anchor box to object box.</p>\n<p>Faster R-CNN has 4 train losses!</p>\n<ol>\n<li>RPN classification: anchor box is object or not object.</li>\n<li>RPN regression: predict transform from anchor box to proposal box.</li>\n<li>Object classification: classify proposals as background or object class.</li>\n<li>Object regression: predict transform from proposal box to object box.</li>\n</ol>\n<p>On a GPU Faster R-CNN is up to 10 times faster than Fast R-CNN.</p>\n<p><img src=\"https://i.imgur.com/bYlP0RJ.jpeg\"></p>\n<h3>Finally… Semantic Segmentation!</h3>\n<p>Semantic segmentation's objective is to label each pixel in the image with a category label. It does not care about differentiating instances, it only cares about pixels. What this means, for example, is that if we have to adjacent categories next to each other (e.g: two cows next to each other) they will not be classified as two different instances but rather one single cloud of points that belong to the same category (one cloud/blob of points which belong to the class \"cow\" but does not tell us which pixels belong to which cow).</p>\n<h3>Semantic segmentation: fully convolutional networks</h3>\n<p>The architecture used in Semantic Segmentation is very much the same as Faster RCNN but with one more branch at the end called the <strong>mask branch</strong>.</p>\n<p>What is commonly used in semantic segmentations mask branchs are Fully Convolutional Networks. This is a convolutional network architecture that does not have any fully connected layers nor pooling layers but rather only have convolutional layers. In this way an input image with a fixed size is passed through all these convolutional layers until on the final layer we perfom a <code>softmax()</code> on each pixel that gives us a set of class scores for every pixel followed by an <code>argmax()</code> that classifies each pixel to a category. We can train this using a cross entropy loss on each pixel.</p>\n<p>In practice, what is used is convolutional networks with downsampling and upsampling, like the famous U-Net. These architectures have high resolution layers followed by mid-resolution layers followed by low-resolution layers on the downsampling stage and the other way round on the upsampling stage.</p>\n<p><img src=\"https://i.imgur.com/zYpAleW.png\"></p>\n<h3>Mask R-CNN</h3>\n<p>Mask R-CNN is just like Faster R-CNN but with an extra branch that predicts a mask.</p>\n<p><img src=\"https://i.imgur.com/8DyNisQ.png\"></p>\n<p>So the way this works in summary is:</p>\n<ol>\n<li>We input our RGB image into the CNN backbone (this can be any CNN, like EfficientNet) which outputs image-level features.</li>\n<li>We pass the image-level features to the Region Proposal Network (RPN) to get the region proposals.</li>\n<li>For each region proposal we pass it through a small semantic segmentation network (like Unet) to get the prediction masks.</li>\n<li>There are many branches involved!<ul>\n<li>Two for the RPN: one to detect the ROIs bounding boxes and another to detect if an object is present in the ROI.</li>\n<li>One for classification to detect which object is present.</li>\n<li>One to detect the bounding box.</li>\n<li>One to detect the mask.</li></ul></li>\n</ol>\n<h3>Where can I find segmentation models?</h3>\n<p>Some popular Python libraries are <a href=\"https://github.com/qubvel/segmentation_models.pytorch\" target=\"_blank\">PyTorch Segmentation Models</a> and <a href=\"https://github.com/open-mmlab/mmsegmentation\" target=\"_blank\">OpenMMLab Semantic Segmentation</a>. </p>\n<p>Example of a Segmentation Network in <code>segmentation_models_pytorch</code>:</p>\n<pre><code> segmentation_models_pytorch  smp\n\nmodel = smp.Unet(\n    encoder_name=,        \n    encoder_weights=,     \n    in_channels=,                  \n    classes=,                      \n)\n</code></pre>\n<h3>Conclusion</h3>\n<p>Semantic Segmentation is a Computer Vision task where we want to classify each pixel in an image to detect an object and its shape/mask. It's similar to Object Detection but with one extra step!</p>\n<p>In this competition our input is an RGB image and its corresponding mask: an image with the same dimension as the original but with 1s and 0s (1 to indicate where the object is present and 0 where it's not). The output is the predicted mask.</p>\n<p>Some popular architectures are UNets and you can find models implemented in PyTorch in <code>segmentation_models_pytorch</code>.</p>\n<p>Some common loss functions used in Segmentation tasks are <a href=\"https://smp.readthedocs.io/en/latest/losses.html#diceloss\" target=\"_blank\">Dice Loss</a> and <a href=\"https://smp.readthedocs.io/en/latest/losses.html#jaccardloss\" target=\"_blank\">Jaccard Loss</a>. Some common metrics used are <a href=\"https://pbs.twimg.com/media/EWci84bWsAIBWvV?format=png&amp;name=small\" target=\"_blank\">Dice Coefficient</a> and <a href=\"https://miro.medium.com/v2/resize:fit:1400/1*FSauyBPV0fiVa_BUVutVTg.png\" target=\"_blank\">IoU Coefficient</a>.</p>\n<p>Hope you liked it! 🍀</p>\n<h3>References</h3>\n<p>Slides were taken from Justin Johnson's amazing Computer Vision course. You can find the course <a href=\"https://www.youtube.com/playlist?list=PL5-TkQAfAZFbzxjBHtzdVCWE0Zbhomg7r\" target=\"_blank\">here</a>. </p>",
  "messages": [
    {
      "id": "2189652",
      "postDate": "03/20/2023 16:53:32",
      "content": "<h3>Table of Contents</h3>\n<ol>\n<li>Introduction</li>\n<li>We need to understand Object Detection first!</li>\n<li>Slow R-CNN</li>\n<li>Fast R-CNN</li>\n<li>Faster R-CNN</li>\n<li>Finally… Semantic Segmentation!</li>\n<li>Mask R-CNN</li>\n<li>Where can I find segmentation models?</li>\n<li>Conclusion</li>\n</ol>\n<h3>Introduction</h3>\n<p>We will present an introduction to the semantic segmentation task. But before we understand semantic segmentation we first need to understand object detection as semantic segmentation is object detection \"with an extra step\".</p>\n<p>There are many different tasks in Computer Vision, some of which you may already be familiary with:</p>\n<p><img src=\"https://i.imgur.com/xeXdDYC.png\"></p>\n<p>For example, some Computer Vision tasks are:</p>\n<ul>\n<li><strong>Image classification</strong> consists of predicting whether a class is present or not in an image. </li>\n<li><strong>Multi-class classification:</strong> multi-class classification is the extension of the <strong>binary classification</strong> problem. In this type of problem each image can belong to only one category, but the amount of categories is more than two. This means that class labels or class membership are mutually exclusive. For example, predicting if the animal present in an image is a cat, a dog or a mouse.</li>\n<li><strong>Multi-label classification:</strong> in multi-label classification one single image can have multiple labels present. This means that class labels or class membership are not mutually exclusive. In multi-label classification, zero or more labels are required as output for each input sample. For example, predicting which animals are present in an image.</li>\n<li><strong>Object detection:</strong> object detection involves drawing a bounding box around one or more objects in an image. In this way, object detection does not only detect if a class is present in an image or not but it also tells where it is located providing a box surrounding the object.</li>\n<li><strong>Semantic segmentation</strong>: labels all parts of an image with a label. Each pixel in an image is assigned a label. The goal of semantic image segmentation is to label each pixel of an image with a corresponding class of what is being represented. Because we're predicting for every pixel in the image, this task is commonly referred to as dense prediction.</li>\n<li><strong>Instance segmentation:</strong> instance segmentation deals with detecting instances of objects and demarcating their boundaries. It is similar to semantic segmentation in the sense that it clearly outputs the object boundaries (irregular boundaries), but it is different in the sense that it does not segment all the image but rather the objects of interest. It is similar to object detection in the sense that it does not segment all the image, but it is different in the sense that it provides a more fine grained boundary than just a bounding box.</li>\n</ul>\n<h3>We need to understand Object Detection first!</h3>\n<p>As said before, semantic segmentation is like an \"extension\" of object detection. So let's get first a clear understanding of the object detection task.</p>\n<p><strong>What's the input and the output?</strong></p>\n<ul>\n<li><strong>Input:</strong> Single RGB image. The shape is then <code>(height, width, channels)</code>.</li>\n<li><strong>Output:</strong> a set of detected objects. For each object predict:<ol>\n<li>Category label, from a fixed known set of categories, like \"cat\", \"dog\", \"mouse\", etc.</li>\n<li>Bounding box (four numbers: x, y, width, height).</li></ol></li>\n</ul>\n<p><strong>What is the notation for bounding boxes?</strong></p>\n<p>These are some common types of notations:</p>\n<ol>\n<li><strong>Upper left corner to lower right corner:</strong> we provide four numbers <code>(x1, y1, x2, y2)</code> where <code>(x1, y1)</code> correspond to the upper left corner coordinates of the box and <code>(x2, y2)</code> correspond to the lower right corner coordinates of the box.</li>\n<li><strong>Center box:</strong> we provide four numbers <code>(x, y, width, height)</code> where <code>(x, y)</code> are the coordinates of the center of the box and the height and width of the box.</li>\n<li><strong>COCO format:</strong> A bounding box is described by the pixel coordinate <code>(x_min, y_min)</code> of its lower left corner within the image together with its width and height in pixels. </li>\n</ol>\n<h3>Detecting a single object</h3>\n<p>What would be the approach? We can do it with a simple NN architecture which uses common backbones (like VGG, AlexNet, ResNet, etc) and attach to them:</p>\n<ul>\n<li>One branch (<strong>\"what\" branch</strong>) of layers that handle the classification problem which consists of some fully connected layers plus a softmax function (to classify which of the possible categories are present) with a Softmax Loss. </li>\n<li>One other branch (<strong>\"where\" branch</strong>) of layers that handle the localization problem which consists of some fully connected layers which end up in 4 nodes (x, y, width, height) with a regression loss, like L2 Loss.</li>\n</ul>\n<p>Since we have two loss functions and we need only one single loss to compute gradient descent we usually just sum up these two loss functions as a weighted sum. We do a weighted sum so that the weights of each loss can be learned and thus one problem does not overcomes the other.</p>\n<h3>Detecting multiple objects: a more challenging challenge</h3>\n<p>Images can have more than one object! What do we do? If we must detect multiple objects in an image we don't know beforehand how many objects will be present and the same object can appear multiple times in the same image too! For example, in a single image:</p>\n<ul>\n<li>One object can appear one or more times (10 dogs in an image).</li>\n<li>Multiple objects can appear and more than one time (2 dogs and 3 cats).</li>\n<li>The objects can appear in different regions.</li>\n</ul>\n<p>One approach to deal with this problem is the <strong>sliding window</strong>, which consists in cropping or dividing the input image in multiple crops. For each crop we apply the CNN and classify each crop as object or background. The obvious problems that rise with this approach are:</p>\n<ul>\n<li><strong>Which should be the shape of the bounding box?</strong></li>\n<li><strong>How many possible boxes are there in an image of size HxW?</strong> There are many! For each image we need to consider all the possible box sizes and ratios which yield the next formula: <br>\n$$ Total:possible:boxes = \\frac{H(H+1)}{2}\\frac{W(W+1)}{2} $$</li>\n</ul>\n<p>If we have a 800x600 image we have ~58M boxes! No way we can evaluate all of them!</p>\n<p>One solution to this problem are <strong>Region Proposal</strong> mechanisms.</p>\n<ul>\n<li>Find a small set of boxes that are likely to cover all objects.</li>\n<li>Often based on heuristics: e.g. look for \"blob-like\" image regions that have high probabilites of containing objects.</li>\n<li>Relatively fast to run: e.g. Selective Search gives ~2000 region proposals in a few seconds on CPU.</li>\n<li>They will end up being replaced by CNN (much faster approach and with learnable parameters).</li>\n<li>A proposal method gives us <strong>regions of interest</strong> (RoI)</li>\n</ul>\n<h3>R-CNN: Region-based CNN (usually nowadays called \"Slow\" R-CNN)</h3>\n<ol>\n<li>We start with an input image and a Region Proposal mechanism which gives us ~2k regions of interest (RoI).</li>\n<li>For each proposed region of interest we warp/resize each region into a fixed size region since the RoI can be of variable heights and widths. The warp size depends on the backbone architecure we use for classifying each warp. Some CNNs require input images to be 224x224 for example.</li>\n<li>For each warped region we run it through a CNN and classify each region. In this way, one image will yield many regions which will be passed through the CNN.</li>\n<li>Once we get a score for each region we try to select a subset of region proposals to output. To explain this better, let's say we have a cat and a dog in an image, but the model outputs ~2000 predictions for all the RoIs, we need a way to summarize all the predicitons so that we just output one single box for the dog and one for the cat. Many choices here: threshold on background or per-category, take top K proposals per image, etc.</li>\n<li>Compare with ground-truth boxes.</li>\n</ol>\n<p><img src=\"https://i.imgur.com/9yLr02y.png\"></p>\n<p>Problems:</p>\n<ul>\n<li>What happen if the proposed regions do not match well to our objects? To overcome this problem, in addition to classifying the class of the region we also perform a bounding box regression which tries to predict or \"transform\" the correct RoI's x, y, height and width. This way we iteratively try to get better RoIs which help us better detect the object in our image.</li>\n</ul>\n<h3>Fast R-CNN</h3>\n<p>As we can see, the R-CNN needs to do ~2000 forward passes for each image since we had ~2000 RoI. This is very expensive computationally and thus slow. The solution that people has come up with is running the CNN <strong>before</strong> warping. This allows us to reshare a lot of computation across image regions. </p>\n<p>This new solution is usually called <strong>Fast R-CNN</strong>. It basically runs the whole input image through a CNN and outputs a feature map or also called image features. This CNN we use is usually called the \"backbone\" network (like AlexNet, VGG, ResNet, etc). Then, the regions of interest are applied to the image features. Next, we warp/crop/resize the RoIs and input them into a light-weight CNN that outputs categories and box transforms per crop.</p>\n<p>This is much much faster since most of the computation happens in the backbone network (this saves work for overlapping region proposals) and the per-region network is relatively light-weight.</p>\n<p>Fast R-CNN can be up to 10 times faster than Slow R-CNN at training time and up to 20 times faster at test time.</p>\n<p><img src=\"https://i.imgur.com/mQeWzlJ.png\"></p>\n<h3>Faster R-CNN: learning region proposals.</h3>\n<p>Even though that Fast R-CNN is much faster than Slow R-CNN most of the compute time is dominated by region proposals! Region proposals are computed by an heuristic (Selective search algorithm) which runs on CPU. We could learn region proposals with a CNN instead!</p>\n<p>Faster R-CNN inserts a <strong>Region Proposal Network (RPN)</strong> to predict proposals from features. Otherwise, it is the same as Fast R-CNN. RPNs use \"anchor boxes\" of fixed size at each point in the feature map. For each anchor box we predict whether the corresponding anchor contains an object or not and a box transform to regress from anchor box to object box.</p>\n<p>Faster R-CNN has 4 train losses!</p>\n<ol>\n<li>RPN classification: anchor box is object or not object.</li>\n<li>RPN regression: predict transform from anchor box to proposal box.</li>\n<li>Object classification: classify proposals as background or object class.</li>\n<li>Object regression: predict transform from proposal box to object box.</li>\n</ol>\n<p>On a GPU Faster R-CNN is up to 10 times faster than Fast R-CNN.</p>\n<p><img src=\"https://i.imgur.com/bYlP0RJ.jpeg\"></p>\n<h3>Finally… Semantic Segmentation!</h3>\n<p>Semantic segmentation's objective is to label each pixel in the image with a category label. It does not care about differentiating instances, it only cares about pixels. What this means, for example, is that if we have to adjacent categories next to each other (e.g: two cows next to each other) they will not be classified as two different instances but rather one single cloud of points that belong to the same category (one cloud/blob of points which belong to the class \"cow\" but does not tell us which pixels belong to which cow).</p>\n<h3>Semantic segmentation: fully convolutional networks</h3>\n<p>The architecture used in Semantic Segmentation is very much the same as Faster RCNN but with one more branch at the end called the <strong>mask branch</strong>.</p>\n<p>What is commonly used in semantic segmentations mask branchs are Fully Convolutional Networks. This is a convolutional network architecture that does not have any fully connected layers nor pooling layers but rather only have convolutional layers. In this way an input image with a fixed size is passed through all these convolutional layers until on the final layer we perfom a <code>softmax()</code> on each pixel that gives us a set of class scores for every pixel followed by an <code>argmax()</code> that classifies each pixel to a category. We can train this using a cross entropy loss on each pixel.</p>\n<p>In practice, what is used is convolutional networks with downsampling and upsampling, like the famous U-Net. These architectures have high resolution layers followed by mid-resolution layers followed by low-resolution layers on the downsampling stage and the other way round on the upsampling stage.</p>\n<p><img src=\"https://i.imgur.com/zYpAleW.png\"></p>\n<h3>Mask R-CNN</h3>\n<p>Mask R-CNN is just like Faster R-CNN but with an extra branch that predicts a mask.</p>\n<p><img src=\"https://i.imgur.com/8DyNisQ.png\"></p>\n<p>So the way this works in summary is:</p>\n<ol>\n<li>We input our RGB image into the CNN backbone (this can be any CNN, like EfficientNet) which outputs image-level features.</li>\n<li>We pass the image-level features to the Region Proposal Network (RPN) to get the region proposals.</li>\n<li>For each region proposal we pass it through a small semantic segmentation network (like Unet) to get the prediction masks.</li>\n<li>There are many branches involved!<ul>\n<li>Two for the RPN: one to detect the ROIs bounding boxes and another to detect if an object is present in the ROI.</li>\n<li>One for classification to detect which object is present.</li>\n<li>One to detect the bounding box.</li>\n<li>One to detect the mask.</li></ul></li>\n</ol>\n<h3>Where can I find segmentation models?</h3>\n<p>Some popular Python libraries are <a href=\"https://github.com/qubvel/segmentation_models.pytorch\" target=\"_blank\">PyTorch Segmentation Models</a> and <a href=\"https://github.com/open-mmlab/mmsegmentation\" target=\"_blank\">OpenMMLab Semantic Segmentation</a>. </p>\n<p>Example of a Segmentation Network in <code>segmentation_models_pytorch</code>:</p>\n<pre><code> segmentation_models_pytorch  smp\n\nmodel = smp.Unet(\n    encoder_name=,        \n    encoder_weights=,     \n    in_channels=,                  \n    classes=,                      \n)\n</code></pre>\n<h3>Conclusion</h3>\n<p>Semantic Segmentation is a Computer Vision task where we want to classify each pixel in an image to detect an object and its shape/mask. It's similar to Object Detection but with one extra step!</p>\n<p>In this competition our input is an RGB image and its corresponding mask: an image with the same dimension as the original but with 1s and 0s (1 to indicate where the object is present and 0 where it's not). The output is the predicted mask.</p>\n<p>Some popular architectures are UNets and you can find models implemented in PyTorch in <code>segmentation_models_pytorch</code>.</p>\n<p>Some common loss functions used in Segmentation tasks are <a href=\"https://smp.readthedocs.io/en/latest/losses.html#diceloss\" target=\"_blank\">Dice Loss</a> and <a href=\"https://smp.readthedocs.io/en/latest/losses.html#jaccardloss\" target=\"_blank\">Jaccard Loss</a>. Some common metrics used are <a href=\"https://pbs.twimg.com/media/EWci84bWsAIBWvV?format=png&amp;name=small\" target=\"_blank\">Dice Coefficient</a> and <a href=\"https://miro.medium.com/v2/resize:fit:1400/1*FSauyBPV0fiVa_BUVutVTg.png\" target=\"_blank\">IoU Coefficient</a>.</p>\n<p>Hope you liked it! 🍀</p>\n<h3>References</h3>\n<p>Slides were taken from Justin Johnson's amazing Computer Vision course. You can find the course <a href=\"https://www.youtube.com/playlist?list=PL5-TkQAfAZFbzxjBHtzdVCWE0Zbhomg7r\" target=\"_blank\">here</a>. </p>",
      "rawMarkdown": "### Table of Contents\n\n1. Introduction\n2. We need to understand Object Detection first!\n3. Slow R-CNN\n4. Fast R-CNN\n5. Faster R-CNN\n6. Finally... Semantic Segmentation!\n7. Mask R-CNN\n8. Where can I find segmentation models?\n9. Conclusion\n\n\n### Introduction\n\nWe will present an introduction to the semantic segmentation task. But before we understand semantic segmentation we first need to understand object detection as semantic segmentation is object detection \"with an extra step\".\n\nThere are many different tasks in Computer Vision, some of which you may already be familiary with:\n\n<img src=\"https://i.imgur.com/xeXdDYC.png\" >\n\nFor example, some Computer Vision tasks are:\n\n- **Image classification** consists of predicting whether a class is present or not in an image. \n- **Multi-class classification:** multi-class classification is the extension of the **binary classification** problem. In this type of problem each image can belong to only one category, but the amount of categories is more than two. This means that class labels or class membership are mutually exclusive. For example, predicting if the animal present in an image is a cat, a dog or a mouse.\n- **Multi-label classification:** in multi-label classification one single image can have multiple labels present. This means that class labels or class membership are not mutually exclusive. In multi-label classification, zero or more labels are required as output for each input sample. For example, predicting which animals are present in an image.\n- **Object detection:** object detection involves drawing a bounding box around one or more objects in an image. In this way, object detection does not only detect if a class is present in an image or not but it also tells where it is located providing a box surrounding the object.\n- **Semantic segmentation**: labels all parts of an image with a label. Each pixel in an image is assigned a label. The goal of semantic image segmentation is to label each pixel of an image with a corresponding class of what is being represented. Because we're predicting for every pixel in the image, this task is commonly referred to as dense prediction.\n- **Instance segmentation:** instance segmentation deals with detecting instances of objects and demarcating their boundaries. It is similar to semantic segmentation in the sense that it clearly outputs the object boundaries (irregular boundaries), but it is different in the sense that it does not segment all the image but rather the objects of interest. It is similar to object detection in the sense that it does not segment all the image, but it is different in the sense that it provides a more fine grained boundary than just a bounding box.\n\n### We need to understand Object Detection first!\n\nAs said before, semantic segmentation is like an \"extension\" of object detection. So let's get first a clear understanding of the object detection task.\n\n**What's the input and the output?**\n\n- **Input:** Single RGB image. The shape is then `(height, width, channels)`.\n- **Output:** a set of detected objects. For each object predict:\n    1. Category label, from a fixed known set of categories, like \"cat\", \"dog\", \"mouse\", etc.\n    2. Bounding box (four numbers: x, y, width, height).\n    \n**What is the notation for bounding boxes?**\n\nThese are some common types of notations:\n1. **Upper left corner to lower right corner:** we provide four numbers `(x1, y1, x2, y2)` where `(x1, y1)` correspond to the upper left corner coordinates of the box and `(x2, y2)` correspond to the lower right corner coordinates of the box.\n2. **Center box:** we provide four numbers `(x, y, width, height)` where `(x, y)` are the coordinates of the center of the box and the height and width of the box.\n3. **COCO format:** A bounding box is described by the pixel coordinate `(x_min, y_min)` of its lower left corner within the image together with its width and height in pixels. \n\n### Detecting a single object\n\nWhat would be the approach? We can do it with a simple NN architecture which uses common backbones (like VGG, AlexNet, ResNet, etc) and attach to them:\n- One branch (**\"what\" branch**) of layers that handle the classification problem which consists of some fully connected layers plus a softmax function (to classify which of the possible categories are present) with a Softmax Loss. \n- One other branch (**\"where\" branch**) of layers that handle the localization problem which consists of some fully connected layers which end up in 4 nodes (x, y, width, height) with a regression loss, like L2 Loss.\n\nSince we have two loss functions and we need only one single loss to compute gradient descent we usually just sum up these two loss functions as a weighted sum. We do a weighted sum so that the weights of each loss can be learned and thus one problem does not overcomes the other.\n\n### Detecting multiple objects: a more challenging challenge\n\nImages can have more than one object! What do we do? If we must detect multiple objects in an image we don't know beforehand how many objects will be present and the same object can appear multiple times in the same image too! For example, in a single image:\n- One object can appear one or more times (10 dogs in an image).\n- Multiple objects can appear and more than one time (2 dogs and 3 cats).\n- The objects can appear in different regions.\n\nOne approach to deal with this problem is the **sliding window**, which consists in cropping or dividing the input image in multiple crops. For each crop we apply the CNN and classify each crop as object or background. The obvious problems that rise with this approach are:\n- **Which should be the shape of the bounding box?**\n- **How many possible boxes are there in an image of size HxW?** There are many! For each image we need to consider all the possible box sizes and ratios which yield the next formula: \n$$ Total\\:possible\\:boxes = \\frac{H(H+1)}{2}\\frac{W(W+1)}{2} $$\n\nIf we have a 800x600 image we have ~58M boxes! No way we can evaluate all of them!\n\nOne solution to this problem are **Region Proposal** mechanisms.\n\n- Find a small set of boxes that are likely to cover all objects.\n- Often based on heuristics: e.g. look for \"blob-like\" image regions that have high probabilites of containing objects.\n- Relatively fast to run: e.g. Selective Search gives ~2000 region proposals in a few seconds on CPU.\n- They will end up being replaced by CNN (much faster approach and with learnable parameters).\n- A proposal method gives us **regions of interest** (RoI)\n\n### R-CNN: Region-based CNN (usually nowadays called \"Slow\" R-CNN)\n\n\n\n1. We start with an input image and a Region Proposal mechanism which gives us ~2k regions of interest (RoI).\n2. For each proposed region of interest we warp/resize each region into a fixed size region since the RoI can be of variable heights and widths. The warp size depends on the backbone architecure we use for classifying each warp. Some CNNs require input images to be 224x224 for example.\n3. For each warped region we run it through a CNN and classify each region. In this way, one image will yield many regions which will be passed through the CNN.\n4. Once we get a score for each region we try to select a subset of region proposals to output. To explain this better, let's say we have a cat and a dog in an image, but the model outputs ~2000 predictions for all the RoIs, we need a way to summarize all the predicitons so that we just output one single box for the dog and one for the cat. Many choices here: threshold on background or per-category, take top K proposals per image, etc.\n5. Compare with ground-truth boxes.\n\n<img src=\"https://i.imgur.com/9yLr02y.png\" width=\"350px\">\n\nProblems:\n- What happen if the proposed regions do not match well to our objects? To overcome this problem, in addition to classifying the class of the region we also perform a bounding box regression which tries to predict or \"transform\" the correct RoI's x, y, height and width. This way we iteratively try to get better RoIs which help us better detect the object in our image.\n    \n### Fast R-CNN\n\nAs we can see, the R-CNN needs to do ~2000 forward passes for each image since we had ~2000 RoI. This is very expensive computationally and thus slow. The solution that people has come up with is running the CNN **before** warping. This allows us to reshare a lot of computation across image regions. \n\nThis new solution is usually called **Fast R-CNN**. It basically runs the whole input image through a CNN and outputs a feature map or also called image features. This CNN we use is usually called the \"backbone\" network (like AlexNet, VGG, ResNet, etc). Then, the regions of interest are applied to the image features. Next, we warp/crop/resize the RoIs and input them into a light-weight CNN that outputs categories and box transforms per crop.\n\nThis is much much faster since most of the computation happens in the backbone network (this saves work for overlapping region proposals) and the per-region network is relatively light-weight.\n\nFast R-CNN can be up to 10 times faster than Slow R-CNN at training time and up to 20 times faster at test time.\n\n<img src=\"https://i.imgur.com/mQeWzlJ.png\" width=\"450px\">\n\n### Faster R-CNN: learning region proposals.\n\nEven though that Fast R-CNN is much faster than Slow R-CNN most of the compute time is dominated by region proposals! Region proposals are computed by an heuristic (Selective search algorithm) which runs on CPU. We could learn region proposals with a CNN instead!\n\n\nFaster R-CNN inserts a **Region Proposal Network (RPN)** to predict proposals from features. Otherwise, it is the same as Fast R-CNN. RPNs use \"anchor boxes\" of fixed size at each point in the feature map. For each anchor box we predict whether the corresponding anchor contains an object or not and a box transform to regress from anchor box to object box.\n\nFaster R-CNN has 4 train losses!\n\n1. RPN classification: anchor box is object or not object.\n2. RPN regression: predict transform from anchor box to proposal box.\n3. Object classification: classify proposals as background or object class.\n4. Object regression: predict transform from proposal box to object box.\n\nOn a GPU Faster R-CNN is up to 10 times faster than Fast R-CNN.\n\n<img src=\"https://i.imgur.com/bYlP0RJ.jpeg\" width=\"450px\">\n\n\n### Finally... Semantic Segmentation!\n\nSemantic segmentation's objective is to label each pixel in the image with a category label. It does not care about differentiating instances, it only cares about pixels. What this means, for example, is that if we have to adjacent categories next to each other (e.g: two cows next to each other) they will not be classified as two different instances but rather one single cloud of points that belong to the same category (one cloud/blob of points which belong to the class \"cow\" but does not tell us which pixels belong to which cow).\n\n\n### Semantic segmentation: fully convolutional networks\n\nThe architecture used in Semantic Segmentation is very much the same as Faster RCNN but with one more branch at the end called the **mask branch**.\n\nWhat is commonly used in semantic segmentations mask branchs are Fully Convolutional Networks. This is a convolutional network architecture that does not have any fully connected layers nor pooling layers but rather only have convolutional layers. In this way an input image with a fixed size is passed through all these convolutional layers until on the final layer we perfom a `softmax()` on each pixel that gives us a set of class scores for every pixel followed by an `argmax()` that classifies each pixel to a category. We can train this using a cross entropy loss on each pixel.\n\nIn practice, what is used is convolutional networks with downsampling and upsampling, like the famous U-Net. These architectures have high resolution layers followed by mid-resolution layers followed by low-resolution layers on the downsampling stage and the other way round on the upsampling stage.\n\n<img src=\"https://i.imgur.com/zYpAleW.png\" width=600>\n\n\n### Mask R-CNN\n\nMask R-CNN is just like Faster R-CNN but with an extra branch that predicts a mask.\n\n</center><img src=\"https://i.imgur.com/8DyNisQ.png\" width=600></center>\n\nSo the way this works in summary is:\n1. We input our RGB image into the CNN backbone (this can be any CNN, like EfficientNet) which outputs image-level features.\n2. We pass the image-level features to the Region Proposal Network (RPN) to get the region proposals.\n3. For each region proposal we pass it through a small semantic segmentation network (like Unet) to get the prediction masks.\n4. There are many branches involved!\n    - Two for the RPN: one to detect the ROIs bounding boxes and another to detect if an object is present in the ROI.\n    - One for classification to detect which object is present.\n    - One to detect the bounding box.\n    - One to detect the mask.\n    \n### Where can I find segmentation models?\n\n\nSome popular Python libraries are [PyTorch Segmentation Models](https://github.com/qubvel/segmentation_models.pytorch) and [OpenMMLab Semantic Segmentation](https://github.com/open-mmlab/mmsegmentation). \n\nExample of a Segmentation Network in `segmentation_models_pytorch`:\n```python\nimport segmentation_models_pytorch as smp\n\nmodel = smp.Unet(\n    encoder_name=\"resnet34\",        # choose encoder, e.g. mobilenet_v2 or efficientnet-b7\n    encoder_weights=\"imagenet\",     # use `imagenet` pre-trained weights for encoder initialization\n    in_channels=1,                  # model input channels (1 for gray-scale images, 3 for RGB, etc.)\n    classes=3,                      # model output channels (number of classes in your dataset)\n)\n```\n\n### Conclusion\n\nSemantic Segmentation is a Computer Vision task where we want to classify each pixel in an image to detect an object and its shape/mask. It's similar to Object Detection but with one extra step!\n\nIn this competition our input is an RGB image and its corresponding mask: an image with the same dimension as the original but with 1s and 0s (1 to indicate where the object is present and 0 where it's not). The output is the predicted mask.\n\nSome popular architectures are UNets and you can find models implemented in PyTorch in `segmentation_models_pytorch`.\n\nSome common loss functions used in Segmentation tasks are [Dice Loss](https://smp.readthedocs.io/en/latest/losses.html#diceloss) and [Jaccard Loss](https://smp.readthedocs.io/en/latest/losses.html#jaccardloss). Some common metrics used are [Dice Coefficient](https://pbs.twimg.com/media/EWci84bWsAIBWvV?format=png&name=small) and [IoU Coefficient](https://miro.medium.com/v2/resize:fit:1400/1*FSauyBPV0fiVa_BUVutVTg.png).\n\nHope you liked it! 🍀\n\n\n### References\n\nSlides were taken from Justin Johnson's amazing Computer Vision course. You can find the course [here](https://www.youtube.com/playlist?list=PL5-TkQAfAZFbzxjBHtzdVCWE0Zbhomg7r).",
      "votes": null
    },
    {
      "id": "2193696",
      "postDate": "03/23/2023 12:29:47",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a>  for sharing a very useful source of knowledge. Nice elaboration full of insights👍</p>",
      "rawMarkdown": "Thanks @alejopaullier  for sharing a very useful source of knowledge. Nice elaboration full of insights👍",
      "votes": null
    },
    {
      "id": "2264021",
      "postDate": "05/18/2023 05:10:32",
      "content": "<p>Excellent explanation <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> </p>",
      "rawMarkdown": "Excellent explanation @alejopaullier",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2193696,
      "author_name": "tariqbashir",
      "author_url": "",
      "post_date": "03/23/2023 12:29:47",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a>  for sharing a very useful source of knowledge. Nice elaboration full of insights👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2264021,
      "author_name": "tamannaakterswarna",
      "author_url": "",
      "post_date": "05/18/2023 05:10:32",
      "content": "<p>Excellent explanation <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2189652": "### Table of Contents\n\n1. Introduction\n2. We need to understand Object Detection first!\n3. Slow R-CNN\n4. Fast R-CNN\n5. Faster R-CNN\n6. Finally... Semantic Segmentation!\n7. Mask R-CNN\n8. Where can I find segmentation models?\n9. Conclusion\n\n\n### Introduction\n\nWe will present an introduction to the semantic segmentation task. But before we understand semantic segmentation we first need to understand object detection as semantic segmentation is object detection \"with an extra step\".\n\nThere are many different tasks in Computer Vision, some of which you may already be familiary with:\n\n<img src=\"https://i.imgur.com/xeXdDYC.png\" >\n\nFor example, some Computer Vision tasks are:\n\n- **Image classification** consists of predicting whether a class is present or not in an image. \n- **Multi-class classification:** multi-class classification is the extension of the **binary classification** problem. In this type of problem each image can belong to only one category, but the amount of categories is more than two. This means that class labels or class membership are mutually exclusive. For example, predicting if the animal present in an image is a cat, a dog or a mouse.\n- **Multi-label classification:** in multi-label classification one single image can have multiple labels present. This means that class labels or class membership are not mutually exclusive. In multi-label classification, zero or more labels are required as output for each input sample. For example, predicting which animals are present in an image.\n- **Object detection:** object detection involves drawing a bounding box around one or more objects in an image. In this way, object detection does not only detect if a class is present in an image or not but it also tells where it is located providing a box surrounding the object.\n- **Semantic segmentation**: labels all parts of an image with a label. Each pixel in an image is assigned a label. The goal of semantic image segmentation is to label each pixel of an image with a corresponding class of what is being represented. Because we're predicting for every pixel in the image, this task is commonly referred to as dense prediction.\n- **Instance segmentation:** instance segmentation deals with detecting instances of objects and demarcating their boundaries. It is similar to semantic segmentation in the sense that it clearly outputs the object boundaries (irregular boundaries), but it is different in the sense that it does not segment all the image but rather the objects of interest. It is similar to object detection in the sense that it does not segment all the image, but it is different in the sense that it provides a more fine grained boundary than just a bounding box.\n\n### We need to understand Object Detection first!\n\nAs said before, semantic segmentation is like an \"extension\" of object detection. So let's get first a clear understanding of the object detection task.\n\n**What's the input and the output?**\n\n- **Input:** Single RGB image. The shape is then `(height, width, channels)`.\n- **Output:** a set of detected objects. For each object predict:\n    1. Category label, from a fixed known set of categories, like \"cat\", \"dog\", \"mouse\", etc.\n    2. Bounding box (four numbers: x, y, width, height).\n    \n**What is the notation for bounding boxes?**\n\nThese are some common types of notations:\n1. **Upper left corner to lower right corner:** we provide four numbers `(x1, y1, x2, y2)` where `(x1, y1)` correspond to the upper left corner coordinates of the box and `(x2, y2)` correspond to the lower right corner coordinates of the box.\n2. **Center box:** we provide four numbers `(x, y, width, height)` where `(x, y)` are the coordinates of the center of the box and the height and width of the box.\n3. **COCO format:** A bounding box is described by the pixel coordinate `(x_min, y_min)` of its lower left corner within the image together with its width and height in pixels. \n\n### Detecting a single object\n\nWhat would be the approach? We can do it with a simple NN architecture which uses common backbones (like VGG, AlexNet, ResNet, etc) and attach to them:\n- One branch (**\"what\" branch**) of layers that handle the classification problem which consists of some fully connected layers plus a softmax function (to classify which of the possible categories are present) with a Softmax Loss. \n- One other branch (**\"where\" branch**) of layers that handle the localization problem which consists of some fully connected layers which end up in 4 nodes (x, y, width, height) with a regression loss, like L2 Loss.\n\nSince we have two loss functions and we need only one single loss to compute gradient descent we usually just sum up these two loss functions as a weighted sum. We do a weighted sum so that the weights of each loss can be learned and thus one problem does not overcomes the other.\n\n### Detecting multiple objects: a more challenging challenge\n\nImages can have more than one object! What do we do? If we must detect multiple objects in an image we don't know beforehand how many objects will be present and the same object can appear multiple times in the same image too! For example, in a single image:\n- One object can appear one or more times (10 dogs in an image).\n- Multiple objects can appear and more than one time (2 dogs and 3 cats).\n- The objects can appear in different regions.\n\nOne approach to deal with this problem is the **sliding window**, which consists in cropping or dividing the input image in multiple crops. For each crop we apply the CNN and classify each crop as object or background. The obvious problems that rise with this approach are:\n- **Which should be the shape of the bounding box?**\n- **How many possible boxes are there in an image of size HxW?** There are many! For each image we need to consider all the possible box sizes and ratios which yield the next formula: \n$$ Total\\:possible\\:boxes = \\frac{H(H+1)}{2}\\frac{W(W+1)}{2} $$\n\nIf we have a 800x600 image we have ~58M boxes! No way we can evaluate all of them!\n\nOne solution to this problem are **Region Proposal** mechanisms.\n\n- Find a small set of boxes that are likely to cover all objects.\n- Often based on heuristics: e.g. look for \"blob-like\" image regions that have high probabilites of containing objects.\n- Relatively fast to run: e.g. Selective Search gives ~2000 region proposals in a few seconds on CPU.\n- They will end up being replaced by CNN (much faster approach and with learnable parameters).\n- A proposal method gives us **regions of interest** (RoI)\n\n### R-CNN: Region-based CNN (usually nowadays called \"Slow\" R-CNN)\n\n\n\n1. We start with an input image and a Region Proposal mechanism which gives us ~2k regions of interest (RoI).\n2. For each proposed region of interest we warp/resize each region into a fixed size region since the RoI can be of variable heights and widths. The warp size depends on the backbone architecure we use for classifying each warp. Some CNNs require input images to be 224x224 for example.\n3. For each warped region we run it through a CNN and classify each region. In this way, one image will yield many regions which will be passed through the CNN.\n4. Once we get a score for each region we try to select a subset of region proposals to output. To explain this better, let's say we have a cat and a dog in an image, but the model outputs ~2000 predictions for all the RoIs, we need a way to summarize all the predicitons so that we just output one single box for the dog and one for the cat. Many choices here: threshold on background or per-category, take top K proposals per image, etc.\n5. Compare with ground-truth boxes.\n\n<img src=\"https://i.imgur.com/9yLr02y.png\" width=\"350px\">\n\nProblems:\n- What happen if the proposed regions do not match well to our objects? To overcome this problem, in addition to classifying the class of the region we also perform a bounding box regression which tries to predict or \"transform\" the correct RoI's x, y, height and width. This way we iteratively try to get better RoIs which help us better detect the object in our image.\n    \n### Fast R-CNN\n\nAs we can see, the R-CNN needs to do ~2000 forward passes for each image since we had ~2000 RoI. This is very expensive computationally and thus slow. The solution that people has come up with is running the CNN **before** warping. This allows us to reshare a lot of computation across image regions. \n\nThis new solution is usually called **Fast R-CNN**. It basically runs the whole input image through a CNN and outputs a feature map or also called image features. This CNN we use is usually called the \"backbone\" network (like AlexNet, VGG, ResNet, etc). Then, the regions of interest are applied to the image features. Next, we warp/crop/resize the RoIs and input them into a light-weight CNN that outputs categories and box transforms per crop.\n\nThis is much much faster since most of the computation happens in the backbone network (this saves work for overlapping region proposals) and the per-region network is relatively light-weight.\n\nFast R-CNN can be up to 10 times faster than Slow R-CNN at training time and up to 20 times faster at test time.\n\n<img src=\"https://i.imgur.com/mQeWzlJ.png\" width=\"450px\">\n\n### Faster R-CNN: learning region proposals.\n\nEven though that Fast R-CNN is much faster than Slow R-CNN most of the compute time is dominated by region proposals! Region proposals are computed by an heuristic (Selective search algorithm) which runs on CPU. We could learn region proposals with a CNN instead!\n\n\nFaster R-CNN inserts a **Region Proposal Network (RPN)** to predict proposals from features. Otherwise, it is the same as Fast R-CNN. RPNs use \"anchor boxes\" of fixed size at each point in the feature map. For each anchor box we predict whether the corresponding anchor contains an object or not and a box transform to regress from anchor box to object box.\n\nFaster R-CNN has 4 train losses!\n\n1. RPN classification: anchor box is object or not object.\n2. RPN regression: predict transform from anchor box to proposal box.\n3. Object classification: classify proposals as background or object class.\n4. Object regression: predict transform from proposal box to object box.\n\nOn a GPU Faster R-CNN is up to 10 times faster than Fast R-CNN.\n\n<img src=\"https://i.imgur.com/bYlP0RJ.jpeg\" width=\"450px\">\n\n\n### Finally... Semantic Segmentation!\n\nSemantic segmentation's objective is to label each pixel in the image with a category label. It does not care about differentiating instances, it only cares about pixels. What this means, for example, is that if we have to adjacent categories next to each other (e.g: two cows next to each other) they will not be classified as two different instances but rather one single cloud of points that belong to the same category (one cloud/blob of points which belong to the class \"cow\" but does not tell us which pixels belong to which cow).\n\n\n### Semantic segmentation: fully convolutional networks\n\nThe architecture used in Semantic Segmentation is very much the same as Faster RCNN but with one more branch at the end called the **mask branch**.\n\nWhat is commonly used in semantic segmentations mask branchs are Fully Convolutional Networks. This is a convolutional network architecture that does not have any fully connected layers nor pooling layers but rather only have convolutional layers. In this way an input image with a fixed size is passed through all these convolutional layers until on the final layer we perfom a `softmax()` on each pixel that gives us a set of class scores for every pixel followed by an `argmax()` that classifies each pixel to a category. We can train this using a cross entropy loss on each pixel.\n\nIn practice, what is used is convolutional networks with downsampling and upsampling, like the famous U-Net. These architectures have high resolution layers followed by mid-resolution layers followed by low-resolution layers on the downsampling stage and the other way round on the upsampling stage.\n\n<img src=\"https://i.imgur.com/zYpAleW.png\" width=600>\n\n\n### Mask R-CNN\n\nMask R-CNN is just like Faster R-CNN but with an extra branch that predicts a mask.\n\n</center><img src=\"https://i.imgur.com/8DyNisQ.png\" width=600></center>\n\nSo the way this works in summary is:\n1. We input our RGB image into the CNN backbone (this can be any CNN, like EfficientNet) which outputs image-level features.\n2. We pass the image-level features to the Region Proposal Network (RPN) to get the region proposals.\n3. For each region proposal we pass it through a small semantic segmentation network (like Unet) to get the prediction masks.\n4. There are many branches involved!\n    - Two for the RPN: one to detect the ROIs bounding boxes and another to detect if an object is present in the ROI.\n    - One for classification to detect which object is present.\n    - One to detect the bounding box.\n    - One to detect the mask.\n    \n### Where can I find segmentation models?\n\n\nSome popular Python libraries are [PyTorch Segmentation Models](https://github.com/qubvel/segmentation_models.pytorch) and [OpenMMLab Semantic Segmentation](https://github.com/open-mmlab/mmsegmentation). \n\nExample of a Segmentation Network in `segmentation_models_pytorch`:\n```python\nimport segmentation_models_pytorch as smp\n\nmodel = smp.Unet(\n    encoder_name=\"resnet34\",        # choose encoder, e.g. mobilenet_v2 or efficientnet-b7\n    encoder_weights=\"imagenet\",     # use `imagenet` pre-trained weights for encoder initialization\n    in_channels=1,                  # model input channels (1 for gray-scale images, 3 for RGB, etc.)\n    classes=3,                      # model output channels (number of classes in your dataset)\n)\n```\n\n### Conclusion\n\nSemantic Segmentation is a Computer Vision task where we want to classify each pixel in an image to detect an object and its shape/mask. It's similar to Object Detection but with one extra step!\n\nIn this competition our input is an RGB image and its corresponding mask: an image with the same dimension as the original but with 1s and 0s (1 to indicate where the object is present and 0 where it's not). The output is the predicted mask.\n\nSome popular architectures are UNets and you can find models implemented in PyTorch in `segmentation_models_pytorch`.\n\nSome common loss functions used in Segmentation tasks are [Dice Loss](https://smp.readthedocs.io/en/latest/losses.html#diceloss) and [Jaccard Loss](https://smp.readthedocs.io/en/latest/losses.html#jaccardloss). Some common metrics used are [Dice Coefficient](https://pbs.twimg.com/media/EWci84bWsAIBWvV?format=png&name=small) and [IoU Coefficient](https://miro.medium.com/v2/resize:fit:1400/1*FSauyBPV0fiVa_BUVutVTg.png).\n\nHope you liked it! 🍀\n\n\n### References\n\nSlides were taken from Justin Johnson's amazing Computer Vision course. You can find the course [here](https://www.youtube.com/playlist?list=PL5-TkQAfAZFbzxjBHtzdVCWE0Zbhomg7r).",
    "2193696": "Thanks @alejopaullier  for sharing a very useful source of knowledge. Nice elaboration full of insights👍",
    "2264021": "Excellent explanation @alejopaullier"
  },
  "source": "meta"
}