{
  "id": 290242,
  "title": "🔥🔥 Tutorial and Paper on Object Detection-(github link)",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/290242",
  "author_name": "Gaju Ahmed",
  "post_date": "2021-11-23T17:38:26.010000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong>Papers</strong><br>\n <a href=\"https://postimg.cc/G4wLHFjZ\" target=\"_blank\"><img src=\"https://i.postimg.cc/g2rnt4Lk/examples.gif\" alt=\"examples.gif\"></a><br>\n<a href=\"https://openreview.net/pdf?id=i8kfkuiCJCI\">Dense Unsupervised Learning for Video Segmentation</a>:  We present a novel approach to unsupervised learning for video object segmentation (VOS). Unlike previous work, our formulation allows to learn dense feature representations directly in a fully convolutional regime. We rely on uniform grid sampling to extract a set of anchors and train our model to disambiguate between them on both inter- and intra-video levels. However, a naive scheme to train such a model results in a degenerate solution. We propose to prevent this with a simple regularisation scheme, accommodating the equivariance property of the segmentation task to similarity transformations. Our training objective admits efficient implementation and exhibits fast training convergence. On established VOS benchmarks, our approach exceeds the segmentation accuracy of previous work despite using significantly less training data and compute power.</p>\n<p><strong>Code</strong> <a href=\"https://github.com/visinf/dense-ulearn-vos\">github link</a><a></a></p>\n<p><a href=\"https://arxiv.org/pdf/1708.02002.pdf\">Focal Loss for Dense Object Detection</a>: The highest accuracy object detectors to date are based on a two-stage approach popularized by R-CNN, where a classifier is applied to a sparse set of candidate object locations. In contrast, one-stage detectors that are appliedover a regular, dense sampling of possible object locations have the potential to be faster and simpler, but have trailedthe accuracy of two-stage detectors thus far. In this paper, we investigate why this is the case. We discover that the extreme foreground-background class imbalance encountered during training of dense detectors is the central cause. We propose to address this class imbalance by reshaping the standard cross entropy loss such that it down-weights the loss assigned to well-classified examples</p>\n<p><strong>Tutorials</strong><br>\n<a href=\"https://machinelearningknowledge.ai/a-brief-history-of-yolo-object-detection-models/\"> A Brief History of YOLO Object Detection Models From YOLOv1 to YOLOv5</a>: Any computer vision enthusiast has surely heard of YOLO models for object detection. Ever since the first YOLOv1 was introduced in 2015, it garnered too much popularity within the computer vision community. Subsequently, multiple versions of YOLOv2, YOLOv3, YOLOv4, and YOLOv5 have been released albeit by different people. In this article, we will give a brief background about all the object detection models of the YOLO family from YOLOv1 to YOLOv5.</p>\n<p><a href=\"https://www.section.io/engineering-education/introduction-to-yolo-algorithm-for-object-detection/\"> Introduction to YOLO Algorithm for Object Detection</a>: YOLO is an algorithm that uses neural networks to provide real-time object detection. This algorithm is popular because of its speed and accuracy. It has been used in various applications to detect traffic signals, people, parking meters, and animals.This article introduces readers to the YOLO algorithm for object detection and explains how it works. It also highlights some of its real-life applications.</p>\n<p><a href=\"https://tryolabs.com/blog/2018/01/18/faster-r-cnn-down-the-rabbit-hole-of-modern-object-detection\">Faster R-CNN: Down the rabbit hole of modern object detection</a>: In 2017, we decided to get into Faster R-CNN, reading the original paper, and all the referenced papers (and so on and on) until we got a clear understanding of how it works and how to implement it.</p>\n<p><a href=\"https://medium.com/@smallfishbigsea/faster-r-cnn-explained-864d4fb7e3f8\">Faster R-CNN Explained</a>: Faster R-CNN has two networks: region proposal network (RPN) for generating region proposals and a network using these proposals to detect objects. The main different here with Fast R-CNN is that the later uses selective search to generate region proposals.</p>\n<p><a href=\"https://www.pyimagesearch.com/2018/11/19/mask-r-cnn-with-opencv/\"> Mask R-CNN with OpenCV</a>:  In the first part of this tutorial, we’ll discuss the difference between image classification, object detection, instance segmentation, and semantic segmentation. From there we’ll briefly review the Mask R-CNN architecture and its connections to Faster R-CNN.</p>",
  "messages": [
    {
      "id": 1593185,
      "postDate": "2021-11-23T17:38:26.010Z",
      "content": "<p><strong>Papers</strong><br>\n <a href=\"https://postimg.cc/G4wLHFjZ\" target=\"_blank\"><img src=\"https://i.postimg.cc/g2rnt4Lk/examples.gif\" alt=\"examples.gif\"></a><br>\n<a href=\"https://openreview.net/pdf?id=i8kfkuiCJCI\">Dense Unsupervised Learning for Video Segmentation</a>:  We present a novel approach to unsupervised learning for video object segmentation (VOS). Unlike previous work, our formulation allows to learn dense feature representations directly in a fully convolutional regime. We rely on uniform grid sampling to extract a set of anchors and train our model to disambiguate between them on both inter- and intra-video levels. However, a naive scheme to train such a model results in a degenerate solution. We propose to prevent this with a simple regularisation scheme, accommodating the equivariance property of the segmentation task to similarity transformations. Our training objective admits efficient implementation and exhibits fast training convergence. On established VOS benchmarks, our approach exceeds the segmentation accuracy of previous work despite using significantly less training data and compute power.</p>\n<p><strong>Code</strong> <a href=\"https://github.com/visinf/dense-ulearn-vos\">github link</a><a></a></p>\n<p><a href=\"https://arxiv.org/pdf/1708.02002.pdf\">Focal Loss for Dense Object Detection</a>: The highest accuracy object detectors to date are based on a two-stage approach popularized by R-CNN, where a classifier is applied to a sparse set of candidate object locations. In contrast, one-stage detectors that are appliedover a regular, dense sampling of possible object locations have the potential to be faster and simpler, but have trailedthe accuracy of two-stage detectors thus far. In this paper, we investigate why this is the case. We discover that the extreme foreground-background class imbalance encountered during training of dense detectors is the central cause. We propose to address this class imbalance by reshaping the standard cross entropy loss such that it down-weights the loss assigned to well-classified examples</p>\n<p><strong>Tutorials</strong><br>\n<a href=\"https://machinelearningknowledge.ai/a-brief-history-of-yolo-object-detection-models/\"> A Brief History of YOLO Object Detection Models From YOLOv1 to YOLOv5</a>: Any computer vision enthusiast has surely heard of YOLO models for object detection. Ever since the first YOLOv1 was introduced in 2015, it garnered too much popularity within the computer vision community. Subsequently, multiple versions of YOLOv2, YOLOv3, YOLOv4, and YOLOv5 have been released albeit by different people. In this article, we will give a brief background about all the object detection models of the YOLO family from YOLOv1 to YOLOv5.</p>\n<p><a href=\"https://www.section.io/engineering-education/introduction-to-yolo-algorithm-for-object-detection/\"> Introduction to YOLO Algorithm for Object Detection</a>: YOLO is an algorithm that uses neural networks to provide real-time object detection. This algorithm is popular because of its speed and accuracy. It has been used in various applications to detect traffic signals, people, parking meters, and animals.This article introduces readers to the YOLO algorithm for object detection and explains how it works. It also highlights some of its real-life applications.</p>\n<p><a href=\"https://tryolabs.com/blog/2018/01/18/faster-r-cnn-down-the-rabbit-hole-of-modern-object-detection\">Faster R-CNN: Down the rabbit hole of modern object detection</a>: In 2017, we decided to get into Faster R-CNN, reading the original paper, and all the referenced papers (and so on and on) until we got a clear understanding of how it works and how to implement it.</p>\n<p><a href=\"https://medium.com/@smallfishbigsea/faster-r-cnn-explained-864d4fb7e3f8\">Faster R-CNN Explained</a>: Faster R-CNN has two networks: region proposal network (RPN) for generating region proposals and a network using these proposals to detect objects. The main different here with Fast R-CNN is that the later uses selective search to generate region proposals.</p>\n<p><a href=\"https://www.pyimagesearch.com/2018/11/19/mask-r-cnn-with-opencv/\"> Mask R-CNN with OpenCV</a>:  In the first part of this tutorial, we’ll discuss the difference between image classification, object detection, instance segmentation, and semantic segmentation. From there we’ll briefly review the Mask R-CNN architecture and its connections to Faster R-CNN.</p>",
      "rawMarkdown": "**Papers**\n [![examples.gif](https://i.postimg.cc/g2rnt4Lk/examples.gif)](https://postimg.cc/G4wLHFjZ)\n<a href=\"https://openreview.net/pdf?id=i8kfkuiCJCI\">Dense Unsupervised Learning for Video Segmentation</a>:  We present a novel approach to unsupervised learning for video object segmentation (VOS). Unlike previous work, our formulation allows to learn dense feature representations directly in a fully convolutional regime. We rely on uniform grid sampling to extract a set of anchors and train our model to disambiguate between them on both inter- and intra-video levels. However, a naive scheme to train such a model results in a degenerate solution. We propose to prevent this with a simple regularisation scheme, accommodating the equivariance property of the segmentation task to similarity transformations. Our training objective admits efficient implementation and exhibits fast training convergence. On established VOS benchmarks, our approach exceeds the segmentation accuracy of previous work despite using significantly less training data and compute power.\n\n**Code** <a href=\"https://github.com/visinf/dense-ulearn-vos\">github link<a>\n\n<a href=\"https://arxiv.org/pdf/1708.02002.pdf\">Focal Loss for Dense Object Detection</a>: The highest accuracy object detectors to date are based on a two-stage approach popularized by R-CNN, where a classifier is applied to a sparse set of candidate object locations. In contrast, one-stage detectors that are appliedover a regular, dense sampling of possible object locations have the potential to be faster and simpler, but have trailedthe accuracy of two-stage detectors thus far. In this paper, we investigate why this is the case. We discover that the extreme foreground-background class imbalance encountered during training of dense detectors is the central cause. We propose to address this class imbalance by reshaping the standard cross entropy loss such that it down-weights the loss assigned to well-classified examples\n\n\n\n**Tutorials**\n<a href=\"https://machinelearningknowledge.ai/a-brief-history-of-yolo-object-detection-models/\"> A Brief History of YOLO Object Detection Models From YOLOv1 to YOLOv5</a>: Any computer vision enthusiast has surely heard of YOLO models for object detection. Ever since the first YOLOv1 was introduced in 2015, it garnered too much popularity within the computer vision community. Subsequently, multiple versions of YOLOv2, YOLOv3, YOLOv4, and YOLOv5 have been released albeit by different people. In this article, we will give a brief background about all the object detection models of the YOLO family from YOLOv1 to YOLOv5.\n\n<a href=\"https://www.section.io/engineering-education/introduction-to-yolo-algorithm-for-object-detection/\"> Introduction to YOLO Algorithm for Object Detection</a>: YOLO is an algorithm that uses neural networks to provide real-time object detection. This algorithm is popular because of its speed and accuracy. It has been used in various applications to detect traffic signals, people, parking meters, and animals.This article introduces readers to the YOLO algorithm for object detection and explains how it works. It also highlights some of its real-life applications.\n\n<a href=\"https://tryolabs.com/blog/2018/01/18/faster-r-cnn-down-the-rabbit-hole-of-modern-object-detection\">Faster R-CNN: Down the rabbit hole of modern object detection</a>: In 2017, we decided to get into Faster R-CNN, reading the original paper, and all the referenced papers (and so on and on) until we got a clear understanding of how it works and how to implement it.\n\n<a href=\"https://medium.com/@smallfishbigsea/faster-r-cnn-explained-864d4fb7e3f8\">Faster R-CNN Explained</a>: Faster R-CNN has two networks: region proposal network (RPN) for generating region proposals and a network using these proposals to detect objects. The main different here with Fast R-CNN is that the later uses selective search to generate region proposals.\n\n<a href=\"https://www.pyimagesearch.com/2018/11/19/mask-r-cnn-with-opencv/\"> Mask R-CNN with OpenCV</a>:  In the first part of this tutorial, we’ll discuss the difference between image classification, object detection, instance segmentation, and semantic segmentation. From there we’ll briefly review the Mask R-CNN architecture and its connections to Faster R-CNN.",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1593185": "**Papers**\n [![examples.gif](https://i.postimg.cc/g2rnt4Lk/examples.gif)](https://postimg.cc/G4wLHFjZ)\n<a href=\"https://openreview.net/pdf?id=i8kfkuiCJCI\">Dense Unsupervised Learning for Video Segmentation</a>:  We present a novel approach to unsupervised learning for video object segmentation (VOS). Unlike previous work, our formulation allows to learn dense feature representations directly in a fully convolutional regime. We rely on uniform grid sampling to extract a set of anchors and train our model to disambiguate between them on both inter- and intra-video levels. However, a naive scheme to train such a model results in a degenerate solution. We propose to prevent this with a simple regularisation scheme, accommodating the equivariance property of the segmentation task to similarity transformations. Our training objective admits efficient implementation and exhibits fast training convergence. On established VOS benchmarks, our approach exceeds the segmentation accuracy of previous work despite using significantly less training data and compute power.\n\n**Code** <a href=\"https://github.com/visinf/dense-ulearn-vos\">github link<a>\n\n<a href=\"https://arxiv.org/pdf/1708.02002.pdf\">Focal Loss for Dense Object Detection</a>: The highest accuracy object detectors to date are based on a two-stage approach popularized by R-CNN, where a classifier is applied to a sparse set of candidate object locations. In contrast, one-stage detectors that are appliedover a regular, dense sampling of possible object locations have the potential to be faster and simpler, but have trailedthe accuracy of two-stage detectors thus far. In this paper, we investigate why this is the case. We discover that the extreme foreground-background class imbalance encountered during training of dense detectors is the central cause. We propose to address this class imbalance by reshaping the standard cross entropy loss such that it down-weights the loss assigned to well-classified examples\n\n\n\n**Tutorials**\n<a href=\"https://machinelearningknowledge.ai/a-brief-history-of-yolo-object-detection-models/\"> A Brief History of YOLO Object Detection Models From YOLOv1 to YOLOv5</a>: Any computer vision enthusiast has surely heard of YOLO models for object detection. Ever since the first YOLOv1 was introduced in 2015, it garnered too much popularity within the computer vision community. Subsequently, multiple versions of YOLOv2, YOLOv3, YOLOv4, and YOLOv5 have been released albeit by different people. In this article, we will give a brief background about all the object detection models of the YOLO family from YOLOv1 to YOLOv5.\n\n<a href=\"https://www.section.io/engineering-education/introduction-to-yolo-algorithm-for-object-detection/\"> Introduction to YOLO Algorithm for Object Detection</a>: YOLO is an algorithm that uses neural networks to provide real-time object detection. This algorithm is popular because of its speed and accuracy. It has been used in various applications to detect traffic signals, people, parking meters, and animals.This article introduces readers to the YOLO algorithm for object detection and explains how it works. It also highlights some of its real-life applications.\n\n<a href=\"https://tryolabs.com/blog/2018/01/18/faster-r-cnn-down-the-rabbit-hole-of-modern-object-detection\">Faster R-CNN: Down the rabbit hole of modern object detection</a>: In 2017, we decided to get into Faster R-CNN, reading the original paper, and all the referenced papers (and so on and on) until we got a clear understanding of how it works and how to implement it.\n\n<a href=\"https://medium.com/@smallfishbigsea/faster-r-cnn-explained-864d4fb7e3f8\">Faster R-CNN Explained</a>: Faster R-CNN has two networks: region proposal network (RPN) for generating region proposals and a network using these proposals to detect objects. The main different here with Fast R-CNN is that the later uses selective search to generate region proposals.\n\n<a href=\"https://www.pyimagesearch.com/2018/11/19/mask-r-cnn-with-opencv/\"> Mask R-CNN with OpenCV</a>:  In the first part of this tutorial, we’ll discuss the difference between image classification, object detection, instance segmentation, and semantic segmentation. From there we’ll briefly review the Mask R-CNN architecture and its connections to Faster R-CNN."
  }
}