{
  "id": 118907,
  "title": "Anchor Free Object Detection",
  "url": "/competitions/pku-autonomous-driving/discussion/118907",
  "author_name": "",
  "post_date": "2019-11-25T10:49:57.007354700Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p><em>The most sucessfull single stage object detection algorithms, e.g., YOLO, SSD, all relies all some anchor to refine to the final detection location. For those algorithms, the anchor are typically defined as the grid on the image coordinates at all possible locations, with different scale and aspect ratio.</em></p>\n\n<p></p>\n\n<p>Though much faster than their two-stage counterparts, single stage algorithms’ speed and performance is still limited by the choice of the anchor boxes: fewer than anchor leads better speed but deteroiates the accuracy. As a result, many new works are trying to design anchor free object detection algorithms.</p>\n\n<p>**The table summarizes the performance of some of the best anchor free methods:</p>\n\n<p>Methods     mAP     FPS     Code</p>\n\n<p>FSAF    42.9    5.3     N.A.\nFCOS    43.2    Nona    <a href=\"https://github.com/tianzhi0549/FCOS\">https://github.com/tianzhi0549/FCOS</a>\nCenterNet：Objects as Points     42.1    7.8     <a href=\"https://github.com/xingyizhou/CenterNet\">https://github.com/xingyizhou/CenterNet</a>\nCenterNet: Keypoint Triplets for Object Detection   44.9    3   <a href=\"https://github.com/Duankaiwen/CenterNet\">https://github.com/Duankaiwen/CenterNet</a>\nAlignDet    44.1    5.6     N.A**</p>\n\n<h1>UnitBox: An Advanced Object Detection Network</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LzOcSkG8Bnd39q5UYWylv8sEayaWWZqIQOpQ7xuVibPCPYt2G9l3JPIg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>UnitBox uses Intersection over Union (IoU) loss function for bounding box prediction.</p>\n\n<h1>DenseBox: Unifying Landmark Localization and Object Detection</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LYc8cvGVeZfACEf8JUwicZohbb8YXfhbJfIwkNZSib0J8shyAuU03uFDA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>DenseBox directly compute the bounding box and its label from the feature map.</p>\n\n<h1>CornerNet: Detecting Objects as Paired Keypoints</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0L3ctVVSmzxIVCkExZNkzg30wGzxHkf06TiaVslj44EF58ycicaRXicDjgA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>In CornerNet, the bounding box is uniquely defined by its top-left corner and bottom-right corner, which is detected by each of the two branches. Corner-pooling is applied to detect the corners, which utilizes the ideas of integral image (see below)</p>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LD86GAmEbLxZxVIbaFJMOhCbppdicrGpWqxftBtLIu7qQBMGSS5kic4Sg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<h1>ExtremeNet: Bottom-up Object Detection by Grouping Extreme and Center Points</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LT64vAia9vmoTmpt0jcBjibFcp3C0odVAiajwoQZQVaicSvsjiaKkz5HDafA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>Similar as CornerNet, it formulates the problem of finding bounding box as finding some corner points. But instead of two corners as in CornerNet, it requires four corner points and one center point, which is computed via peaks of heatmaps of each corner points.</p>\n\n<h1>FSAF: Feature Selective Anchor-Free</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LJCRy91ibAlAsEB7ZNfEMEvxJkMef7893P2x5R3b5UBxlqUIA3wpMhhg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>It is based on feature pyramid network, where the final result is dynamically selected from the optimal resolution.</p>\n\n<h1>FCOS: Fully Convolutional One-Stage</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LvBBpibWrfwdSXa6bbKM1HLzfhMZQIwJhCxoCm5rYSqaghIs4oicLo5MA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>FCOS is anchor-box free, as well as proposal free. FCOS works by predicting a 4D vector (l, t, r, b) encoding the location of a bounding box at each foreground pixel (supervised by ground-truth bounding box information during training).</p>\n\n<p>This done in a per-pixel prediction way, i.e., for each pixel, the network try to predict a bounding box from it, together with the label of class. To counter for the pixel which are far from the ground truth object (center), a centerness score is also predicted which downweights the prediction for those pixels.</p>\n\n<p>If a location falls into multiple bounding boxes, it is considered as an ambiguous sample. For now, we simply choose the bounding box with minimal area as its regression target.</p>\n\n<p>Feature Pyramid Network is used as the backbone.</p>\n\n<h1>FoveaBox: Beyond Anchor-based Object Detector</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0Ljs1AqgtEc9jQpeibLT5RpCyibViccLj4yDK5CDLfOaUeN4fsibQjOaZHdg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>It is very similar to FCOS.</p>\n\n<h1>Region Proposal by Guided Anchoring(GA-RPN)</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0Lg9aBScmc5vbx7eibpICPfndpGpiaUbkzmcicnKZEiabsbq5UOAmgM4NtJQ/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>In GA-RPN, the anchor (defined as a tuple of its location and shape) is learned instead of manually defined. Then feature extraction is then adapted to this computed anchor. CenterNet: Objects as Points</p>\n\n<h1>CenterNet: Objects as Points</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0Lnia89KDm0CMyzL67zXt7qv43vA0mFKpIWLS3ibOvwYZ83CPKATmASNXQ/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>CenterNet defines the bounding box by its center. After the center is computed, its shape and pose can be further computed.</p>\n\n<h1>CenterNet: Object Detection with Keypoint Triplets</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LH5iczkq229yLv31ciaH9AhlWmoGicYBtiacZia1PbAnsf0udVvfbMjGv0Yw/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>It is based on CenterNet but very similar to ExtremeNet or CornerNet, where the bounding box is now defined by a pair of corner points and the label is defined by the response of the center point.</p>\n\n<h1>CornerNet-Lite: Efficient Keypoint Based Object Detection</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LejkaY57fbNONtcBCUl6fuYkFI1jZQiaJt9rOuTgJ8YtBz3qp2AaWdcw/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>CornerNet-Lite：CornerNet-Saccade（attention mechanism）+ CornerNet-Squeeze</p>\n\n<h1>[Center and Scale Prediction: A Box-free Approach for Object Detection]</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LTHBjImc9t9y9mnnEnH1MiaKJEevS5heHA47OTAoicIpFFiawKs1CLQYmg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>As GA-RPN, the bounding box is defined by its center and shape, which is computed from two branches of the neural network.</p>\n\n<h1>Matrix Nets</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_png/KmXPKA19gW9eTcrcwFsib63LiacvgWwib9lZCVgVNXvTYfic0cnRLciap91gibXTiaQYs57M6ibCtCzyhTOuzk4IiabP7nw/640?wx_fmt=png&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>Maxtrix Nets addresses different aspect ratio of the objects. Compared with image pyramid, it generates feature across different scales and aspect ratio, like a matrix. Especially, for feature x_{ij} at layer i and j, feature x_{i+1, j+1} is generated via a 3x3 kernel with stride 2x2, feature x_{i+1, j} is generated via a 3x3 kernel with stride 2x1 and feature x_{i, j+1} is generated via a 3x3 kernel with stride 1x2. Note those three convolutions share the same parameter.</p>\n\n<p>To detect objects, it utilizes the similar idea of center net: the location of top left corner and bottom right corner are detected, the centerness is also computed. Then the result for feature x_{i+k, i+[0:k]} and x_{i+[0:k], i+k} is aggreated via nonmax suppression.</p>\n\n<p><a href=\"https://zhangtemplar.github.io/anchor-free-detection/\">source</a></p>",
  "messages": [
    {
      "id": "680865",
      "postDate": "11/25/2019 10:49:57",
      "content": "<p><em>The most sucessfull single stage object detection algorithms, e.g., YOLO, SSD, all relies all some anchor to refine to the final detection location. For those algorithms, the anchor are typically defined as the grid on the image coordinates at all possible locations, with different scale and aspect ratio.</em></p>\n\n<p></p>\n\n<p>Though much faster than their two-stage counterparts, single stage algorithms’ speed and performance is still limited by the choice of the anchor boxes: fewer than anchor leads better speed but deteroiates the accuracy. As a result, many new works are trying to design anchor free object detection algorithms.</p>\n\n<p>**The table summarizes the performance of some of the best anchor free methods:</p>\n\n<p>Methods     mAP     FPS     Code</p>\n\n<p>FSAF    42.9    5.3     N.A.\nFCOS    43.2    Nona    <a href=\"https://github.com/tianzhi0549/FCOS\">https://github.com/tianzhi0549/FCOS</a>\nCenterNet：Objects as Points     42.1    7.8     <a href=\"https://github.com/xingyizhou/CenterNet\">https://github.com/xingyizhou/CenterNet</a>\nCenterNet: Keypoint Triplets for Object Detection   44.9    3   <a href=\"https://github.com/Duankaiwen/CenterNet\">https://github.com/Duankaiwen/CenterNet</a>\nAlignDet    44.1    5.6     N.A**</p>\n\n<h1>UnitBox: An Advanced Object Detection Network</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LzOcSkG8Bnd39q5UYWylv8sEayaWWZqIQOpQ7xuVibPCPYt2G9l3JPIg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>UnitBox uses Intersection over Union (IoU) loss function for bounding box prediction.</p>\n\n<h1>DenseBox: Unifying Landmark Localization and Object Detection</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LYc8cvGVeZfACEf8JUwicZohbb8YXfhbJfIwkNZSib0J8shyAuU03uFDA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>DenseBox directly compute the bounding box and its label from the feature map.</p>\n\n<h1>CornerNet: Detecting Objects as Paired Keypoints</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0L3ctVVSmzxIVCkExZNkzg30wGzxHkf06TiaVslj44EF58ycicaRXicDjgA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>In CornerNet, the bounding box is uniquely defined by its top-left corner and bottom-right corner, which is detected by each of the two branches. Corner-pooling is applied to detect the corners, which utilizes the ideas of integral image (see below)</p>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LD86GAmEbLxZxVIbaFJMOhCbppdicrGpWqxftBtLIu7qQBMGSS5kic4Sg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<h1>ExtremeNet: Bottom-up Object Detection by Grouping Extreme and Center Points</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LT64vAia9vmoTmpt0jcBjibFcp3C0odVAiajwoQZQVaicSvsjiaKkz5HDafA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>Similar as CornerNet, it formulates the problem of finding bounding box as finding some corner points. But instead of two corners as in CornerNet, it requires four corner points and one center point, which is computed via peaks of heatmaps of each corner points.</p>\n\n<h1>FSAF: Feature Selective Anchor-Free</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LJCRy91ibAlAsEB7ZNfEMEvxJkMef7893P2x5R3b5UBxlqUIA3wpMhhg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>It is based on feature pyramid network, where the final result is dynamically selected from the optimal resolution.</p>\n\n<h1>FCOS: Fully Convolutional One-Stage</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LvBBpibWrfwdSXa6bbKM1HLzfhMZQIwJhCxoCm5rYSqaghIs4oicLo5MA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>FCOS is anchor-box free, as well as proposal free. FCOS works by predicting a 4D vector (l, t, r, b) encoding the location of a bounding box at each foreground pixel (supervised by ground-truth bounding box information during training).</p>\n\n<p>This done in a per-pixel prediction way, i.e., for each pixel, the network try to predict a bounding box from it, together with the label of class. To counter for the pixel which are far from the ground truth object (center), a centerness score is also predicted which downweights the prediction for those pixels.</p>\n\n<p>If a location falls into multiple bounding boxes, it is considered as an ambiguous sample. For now, we simply choose the bounding box with minimal area as its regression target.</p>\n\n<p>Feature Pyramid Network is used as the backbone.</p>\n\n<h1>FoveaBox: Beyond Anchor-based Object Detector</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0Ljs1AqgtEc9jQpeibLT5RpCyibViccLj4yDK5CDLfOaUeN4fsibQjOaZHdg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>It is very similar to FCOS.</p>\n\n<h1>Region Proposal by Guided Anchoring(GA-RPN)</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0Lg9aBScmc5vbx7eibpICPfndpGpiaUbkzmcicnKZEiabsbq5UOAmgM4NtJQ/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>In GA-RPN, the anchor (defined as a tuple of its location and shape) is learned instead of manually defined. Then feature extraction is then adapted to this computed anchor. CenterNet: Objects as Points</p>\n\n<h1>CenterNet: Objects as Points</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0Lnia89KDm0CMyzL67zXt7qv43vA0mFKpIWLS3ibOvwYZ83CPKATmASNXQ/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>CenterNet defines the bounding box by its center. After the center is computed, its shape and pose can be further computed.</p>\n\n<h1>CenterNet: Object Detection with Keypoint Triplets</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LH5iczkq229yLv31ciaH9AhlWmoGicYBtiacZia1PbAnsf0udVvfbMjGv0Yw/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>It is based on CenterNet but very similar to ExtremeNet or CornerNet, where the bounding box is now defined by a pair of corner points and the label is defined by the response of the center point.</p>\n\n<h1>CornerNet-Lite: Efficient Keypoint Based Object Detection</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LejkaY57fbNONtcBCUl6fuYkFI1jZQiaJt9rOuTgJ8YtBz3qp2AaWdcw/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>CornerNet-Lite：CornerNet-Saccade（attention mechanism）+ CornerNet-Squeeze</p>\n\n<h1>[Center and Scale Prediction: A Box-free Approach for Object Detection]</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LTHBjImc9t9y9mnnEnH1MiaKJEevS5heHA47OTAoicIpFFiawKs1CLQYmg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>As GA-RPN, the bounding box is defined by its center and shape, which is computed from two branches of the neural network.</p>\n\n<h1>Matrix Nets</h1>\n\n<p><img src=\"https://mmbiz.qpic.cn/mmbiz_png/KmXPKA19gW9eTcrcwFsib63LiacvgWwib9lZCVgVNXvTYfic0cnRLciap91gibXTiaQYs57M6ibCtCzyhTOuzk4IiabP7nw/640?wx_fmt=png&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1\" alt=\"\"></p>\n\n<p>Maxtrix Nets addresses different aspect ratio of the objects. Compared with image pyramid, it generates feature across different scales and aspect ratio, like a matrix. Especially, for feature x_{ij} at layer i and j, feature x_{i+1, j+1} is generated via a 3x3 kernel with stride 2x2, feature x_{i+1, j} is generated via a 3x3 kernel with stride 2x1 and feature x_{i, j+1} is generated via a 3x3 kernel with stride 1x2. Note those three convolutions share the same parameter.</p>\n\n<p>To detect objects, it utilizes the similar idea of center net: the location of top left corner and bottom right corner are detected, the centerness is also computed. Then the result for feature x_{i+k, i+[0:k]} and x_{i+[0:k], i+k} is aggreated via nonmax suppression.</p>\n\n<p><a href=\"https://zhangtemplar.github.io/anchor-free-detection/\">source</a></p>",
      "rawMarkdown": "*The most sucessfull single stage object detection algorithms, e.g., YOLO, SSD, all relies all some anchor to refine to the final detection location. For those algorithms, the anchor are typically defined as the grid on the image coordinates at all possible locations, with different scale and aspect ratio.*\n\n![](https://cdn-images-1.medium.com/max/1600/1*7heX-no7cdqllky-GwGBfQ.png)\n\nThough much faster than their two-stage counterparts, single stage algorithms’ speed and performance is still limited by the choice of the anchor boxes: fewer than anchor leads better speed but deteroiates the accuracy. As a result, many new works are trying to design anchor free object detection algorithms.\n\n**The table summarizes the performance of some of the best anchor free methods:\n\nMethods \tmAP \tFPS \tCode\n\nFSAF \t42.9 \t5.3 \tN.A.\nFCOS \t43.2 \tNona \thttps://github.com/tianzhi0549/FCOS\nCenterNet：Objects as Points \t42.1 \t7.8 \thttps://github.com/xingyizhou/CenterNet\nCenterNet: Keypoint Triplets for Object Detection \t44.9 \t3 \thttps://github.com/Duankaiwen/CenterNet\nAlignDet \t44.1 \t5.6 \tN.A**\n\n# UnitBox: An Advanced Object Detection Network\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LzOcSkG8Bnd39q5UYWylv8sEayaWWZqIQOpQ7xuVibPCPYt2G9l3JPIg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nUnitBox uses Intersection over Union (IoU) loss function for bounding box prediction.\n\n# DenseBox: Unifying Landmark Localization and Object Detection\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LYc8cvGVeZfACEf8JUwicZohbb8YXfhbJfIwkNZSib0J8shyAuU03uFDA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nDenseBox directly compute the bounding box and its label from the feature map.\n\n# CornerNet: Detecting Objects as Paired Keypoints\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0L3ctVVSmzxIVCkExZNkzg30wGzxHkf06TiaVslj44EF58ycicaRXicDjgA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nIn CornerNet, the bounding box is uniquely defined by its top-left corner and bottom-right corner, which is detected by each of the two branches. Corner-pooling is applied to detect the corners, which utilizes the ideas of integral image (see below)\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LD86GAmEbLxZxVIbaFJMOhCbppdicrGpWqxftBtLIu7qQBMGSS5kic4Sg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\n# ExtremeNet: Bottom-up Object Detection by Grouping Extreme and Center Points\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LT64vAia9vmoTmpt0jcBjibFcp3C0odVAiajwoQZQVaicSvsjiaKkz5HDafA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nSimilar as CornerNet, it formulates the problem of finding bounding box as finding some corner points. But instead of two corners as in CornerNet, it requires four corner points and one center point, which is computed via peaks of heatmaps of each corner points.\n\n# FSAF: Feature Selective Anchor-Free\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LJCRy91ibAlAsEB7ZNfEMEvxJkMef7893P2x5R3b5UBxlqUIA3wpMhhg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nIt is based on feature pyramid network, where the final result is dynamically selected from the optimal resolution.\n\n# FCOS: Fully Convolutional One-Stage\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LvBBpibWrfwdSXa6bbKM1HLzfhMZQIwJhCxoCm5rYSqaghIs4oicLo5MA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nFCOS is anchor-box free, as well as proposal free. FCOS works by predicting a 4D vector (l, t, r, b) encoding the location of a bounding box at each foreground pixel (supervised by ground-truth bounding box information during training).\n\nThis done in a per-pixel prediction way, i.e., for each pixel, the network try to predict a bounding box from it, together with the label of class. To counter for the pixel which are far from the ground truth object (center), a centerness score is also predicted which downweights the prediction for those pixels.\n\nIf a location falls into multiple bounding boxes, it is considered as an ambiguous sample. For now, we simply choose the bounding box with minimal area as its regression target.\n\nFeature Pyramid Network is used as the backbone.\n\n# FoveaBox: Beyond Anchor-based Object Detector\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0Ljs1AqgtEc9jQpeibLT5RpCyibViccLj4yDK5CDLfOaUeN4fsibQjOaZHdg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nIt is very similar to FCOS.\n\n# Region Proposal by Guided Anchoring(GA-RPN)\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0Lg9aBScmc5vbx7eibpICPfndpGpiaUbkzmcicnKZEiabsbq5UOAmgM4NtJQ/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nIn GA-RPN, the anchor (defined as a tuple of its location and shape) is learned instead of manually defined. Then feature extraction is then adapted to this computed anchor. CenterNet: Objects as Points\n\n# CenterNet: Objects as Points\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0Lnia89KDm0CMyzL67zXt7qv43vA0mFKpIWLS3ibOvwYZ83CPKATmASNXQ/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nCenterNet defines the bounding box by its center. After the center is computed, its shape and pose can be further computed.\n\n# CenterNet: Object Detection with Keypoint Triplets\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LH5iczkq229yLv31ciaH9AhlWmoGicYBtiacZia1PbAnsf0udVvfbMjGv0Yw/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nIt is based on CenterNet but very similar to ExtremeNet or CornerNet, where the bounding box is now defined by a pair of corner points and the label is defined by the response of the center point.\n\n# CornerNet-Lite: Efficient Keypoint Based Object Detection\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LejkaY57fbNONtcBCUl6fuYkFI1jZQiaJt9rOuTgJ8YtBz3qp2AaWdcw/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nCornerNet-Lite：CornerNet-Saccade（attention mechanism）+ CornerNet-Squeeze\n\n# [Center and Scale Prediction: A Box-free Approach for Object Detection]\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LTHBjImc9t9y9mnnEnH1MiaKJEevS5heHA47OTAoicIpFFiawKs1CLQYmg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nAs GA-RPN, the bounding box is defined by its center and shape, which is computed from two branches of the neural network.\n\n# Matrix Nets\n\n![](https://mmbiz.qpic.cn/mmbiz_png/KmXPKA19gW9eTcrcwFsib63LiacvgWwib9lZCVgVNXvTYfic0cnRLciap91gibXTiaQYs57M6ibCtCzyhTOuzk4IiabP7nw/640?wx_fmt=png&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nMaxtrix Nets addresses different aspect ratio of the objects. Compared with image pyramid, it generates feature across different scales and aspect ratio, like a matrix. Especially, for feature x_{ij} at layer i and j, feature x_{i+1, j+1} is generated via a 3x3 kernel with stride 2x2, feature x_{i+1, j} is generated via a 3x3 kernel with stride 2x1 and feature x_{i, j+1} is generated via a 3x3 kernel with stride 1x2. Note those three convolutions share the same parameter.\n\nTo detect objects, it utilizes the similar idea of center net: the location of top left corner and bottom right corner are detected, the centerness is also computed. Then the result for feature x_{i+k, i+[0:k]} and x_{i+[0:k], i+k} is aggreated via nonmax suppression.\n\n[source](https://zhangtemplar.github.io/anchor-free-detection/)",
      "votes": null
    },
    {
      "id": "680875",
      "postDate": "11/25/2019 10:59:48",
      "content": "<p>This is very helpful Brother . </p>",
      "rawMarkdown": "This is very helpful Brother .",
      "votes": null
    },
    {
      "id": "680879",
      "postDate": "11/25/2019 11:04:59",
      "content": "<p>yes brother,we need to implement some of those architectures for this competition :)</p>",
      "rawMarkdown": "yes brother,we need to implement some of those architectures for this competition :)",
      "votes": null
    },
    {
      "id": "681431",
      "postDate": "11/26/2019 05:22:38",
      "content": "<p>Very Informative &amp; Helpful\nThanks <a href=\"/mobassir\">@mobassir</a> </p>",
      "rawMarkdown": "Very Informative &amp; Helpful\nThanks @mobassir",
      "votes": null
    },
    {
      "id": "681437",
      "postDate": "11/26/2019 05:30:43",
      "content": "<p>My pleasure </p>",
      "rawMarkdown": "My pleasure",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 680875,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "11/25/2019 10:59:48",
      "content": "<p>This is very helpful Brother . </p>",
      "votes": null,
      "replies": [
        {
          "id": 680879,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "11/25/2019 11:04:59",
          "content": "<p>yes brother,we need to implement some of those architectures for this competition :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 681431,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "11/26/2019 05:22:38",
      "content": "<p>Very Informative &amp; Helpful\nThanks <a href=\"/mobassir\">@mobassir</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 681437,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "11/26/2019 05:30:43",
          "content": "<p>My pleasure </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "680865": "*The most sucessfull single stage object detection algorithms, e.g., YOLO, SSD, all relies all some anchor to refine to the final detection location. For those algorithms, the anchor are typically defined as the grid on the image coordinates at all possible locations, with different scale and aspect ratio.*\n\n![](https://cdn-images-1.medium.com/max/1600/1*7heX-no7cdqllky-GwGBfQ.png)\n\nThough much faster than their two-stage counterparts, single stage algorithms’ speed and performance is still limited by the choice of the anchor boxes: fewer than anchor leads better speed but deteroiates the accuracy. As a result, many new works are trying to design anchor free object detection algorithms.\n\n**The table summarizes the performance of some of the best anchor free methods:\n\nMethods \tmAP \tFPS \tCode\n\nFSAF \t42.9 \t5.3 \tN.A.\nFCOS \t43.2 \tNona \thttps://github.com/tianzhi0549/FCOS\nCenterNet：Objects as Points \t42.1 \t7.8 \thttps://github.com/xingyizhou/CenterNet\nCenterNet: Keypoint Triplets for Object Detection \t44.9 \t3 \thttps://github.com/Duankaiwen/CenterNet\nAlignDet \t44.1 \t5.6 \tN.A**\n\n# UnitBox: An Advanced Object Detection Network\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LzOcSkG8Bnd39q5UYWylv8sEayaWWZqIQOpQ7xuVibPCPYt2G9l3JPIg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nUnitBox uses Intersection over Union (IoU) loss function for bounding box prediction.\n\n# DenseBox: Unifying Landmark Localization and Object Detection\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LYc8cvGVeZfACEf8JUwicZohbb8YXfhbJfIwkNZSib0J8shyAuU03uFDA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nDenseBox directly compute the bounding box and its label from the feature map.\n\n# CornerNet: Detecting Objects as Paired Keypoints\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0L3ctVVSmzxIVCkExZNkzg30wGzxHkf06TiaVslj44EF58ycicaRXicDjgA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nIn CornerNet, the bounding box is uniquely defined by its top-left corner and bottom-right corner, which is detected by each of the two branches. Corner-pooling is applied to detect the corners, which utilizes the ideas of integral image (see below)\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LD86GAmEbLxZxVIbaFJMOhCbppdicrGpWqxftBtLIu7qQBMGSS5kic4Sg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\n# ExtremeNet: Bottom-up Object Detection by Grouping Extreme and Center Points\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LT64vAia9vmoTmpt0jcBjibFcp3C0odVAiajwoQZQVaicSvsjiaKkz5HDafA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nSimilar as CornerNet, it formulates the problem of finding bounding box as finding some corner points. But instead of two corners as in CornerNet, it requires four corner points and one center point, which is computed via peaks of heatmaps of each corner points.\n\n# FSAF: Feature Selective Anchor-Free\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LJCRy91ibAlAsEB7ZNfEMEvxJkMef7893P2x5R3b5UBxlqUIA3wpMhhg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nIt is based on feature pyramid network, where the final result is dynamically selected from the optimal resolution.\n\n# FCOS: Fully Convolutional One-Stage\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LvBBpibWrfwdSXa6bbKM1HLzfhMZQIwJhCxoCm5rYSqaghIs4oicLo5MA/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nFCOS is anchor-box free, as well as proposal free. FCOS works by predicting a 4D vector (l, t, r, b) encoding the location of a bounding box at each foreground pixel (supervised by ground-truth bounding box information during training).\n\nThis done in a per-pixel prediction way, i.e., for each pixel, the network try to predict a bounding box from it, together with the label of class. To counter for the pixel which are far from the ground truth object (center), a centerness score is also predicted which downweights the prediction for those pixels.\n\nIf a location falls into multiple bounding boxes, it is considered as an ambiguous sample. For now, we simply choose the bounding box with minimal area as its regression target.\n\nFeature Pyramid Network is used as the backbone.\n\n# FoveaBox: Beyond Anchor-based Object Detector\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0Ljs1AqgtEc9jQpeibLT5RpCyibViccLj4yDK5CDLfOaUeN4fsibQjOaZHdg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nIt is very similar to FCOS.\n\n# Region Proposal by Guided Anchoring(GA-RPN)\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0Lg9aBScmc5vbx7eibpICPfndpGpiaUbkzmcicnKZEiabsbq5UOAmgM4NtJQ/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nIn GA-RPN, the anchor (defined as a tuple of its location and shape) is learned instead of manually defined. Then feature extraction is then adapted to this computed anchor. CenterNet: Objects as Points\n\n# CenterNet: Objects as Points\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0Lnia89KDm0CMyzL67zXt7qv43vA0mFKpIWLS3ibOvwYZ83CPKATmASNXQ/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nCenterNet defines the bounding box by its center. After the center is computed, its shape and pose can be further computed.\n\n# CenterNet: Object Detection with Keypoint Triplets\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LH5iczkq229yLv31ciaH9AhlWmoGicYBtiacZia1PbAnsf0udVvfbMjGv0Yw/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nIt is based on CenterNet but very similar to ExtremeNet or CornerNet, where the bounding box is now defined by a pair of corner points and the label is defined by the response of the center point.\n\n# CornerNet-Lite: Efficient Keypoint Based Object Detection\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LejkaY57fbNONtcBCUl6fuYkFI1jZQiaJt9rOuTgJ8YtBz3qp2AaWdcw/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nCornerNet-Lite：CornerNet-Saccade（attention mechanism）+ CornerNet-Squeeze\n\n# [Center and Scale Prediction: A Box-free Approach for Object Detection]\n\n![](https://mmbiz.qpic.cn/mmbiz_jpg/yNnalkXE7oVtmWMCMAYFgyQHf8Yn8V0LTHBjImc9t9y9mnnEnH1MiaKJEevS5heHA47OTAoicIpFFiawKs1CLQYmg/640?wx_fmt=jpeg&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nAs GA-RPN, the bounding box is defined by its center and shape, which is computed from two branches of the neural network.\n\n# Matrix Nets\n\n![](https://mmbiz.qpic.cn/mmbiz_png/KmXPKA19gW9eTcrcwFsib63LiacvgWwib9lZCVgVNXvTYfic0cnRLciap91gibXTiaQYs57M6ibCtCzyhTOuzk4IiabP7nw/640?wx_fmt=png&amp;tp=webp&amp;wxfrom=5&amp;wx_lazy=1&amp;wx_co=1)\n\nMaxtrix Nets addresses different aspect ratio of the objects. Compared with image pyramid, it generates feature across different scales and aspect ratio, like a matrix. Especially, for feature x_{ij} at layer i and j, feature x_{i+1, j+1} is generated via a 3x3 kernel with stride 2x2, feature x_{i+1, j} is generated via a 3x3 kernel with stride 2x1 and feature x_{i, j+1} is generated via a 3x3 kernel with stride 1x2. Note those three convolutions share the same parameter.\n\nTo detect objects, it utilizes the similar idea of center net: the location of top left corner and bottom right corner are detected, the centerness is also computed. Then the result for feature x_{i+k, i+[0:k]} and x_{i+[0:k], i+k} is aggreated via nonmax suppression.\n\n[source](https://zhangtemplar.github.io/anchor-free-detection/)",
    "680875": "This is very helpful Brother .",
    "680879": "yes brother,we need to implement some of those architectures for this competition :)",
    "681431": "Very Informative &amp; Helpful\nThanks @mobassir",
    "681437": "My pleasure"
  },
  "source": "meta"
}