{
  "id": 307707,
  "title": "3rd place solution - Team Hydrogen",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/307707",
  "author_name": "Psi",
  "post_date": "2022-02-15T10:35:42.740000",
  "votes": 130,
  "comment_count": 31,
  "views": 0,
  "content": "<p>Thanks to hosts and all participants for this interesting competition. This is a joint writeup with  <a href=\"https://www.kaggle.com/ybabakhin\" target=\"_blank\">@ybabakhin</a> (big, big thanks for the awesome team work in this competition) from Team Hydrogen.</p>\n<h4>Summary</h4>\n<p>Our solution is a blend of five different object detection model types with tracking.</p>\n<h4>Cross-Validation</h4>\n<p>For cross-validation we used two strategies: one is simple split by video, and the other is making 5-folds based on subsequences. We were mostly looking at subsequence Cross-Validation scores, time-to-time checking the video folds performance. This strategy gave us a good correlation between CV and Public LB scores throughout the competition, however it failed for the Private LB. You can read our thoughts about such a behavior below in this post.</p>\n<h4>Models</h4>\n<p>Our final solution is a blend of five different model types. All our models are trained on a resolution of 1152x2048 and only using images with boxes. For augmentations we employ a mix of random flips, scale, shift, rotate, brightness, contrats, cutout, cutmix.</p>\n<p>An additional augmentation we utilize in most of our models is what we call “background-mix”. Here, we randomly mix an image with boxes, with a random image without boxes. This guarantees us to not overlap / distort any boxes, but generalizes better to unseen backgrounds and also randomizes the intensity of objects. Also some models benefit from using cut&amp;paste augmentation by simply cutting and pasting objects to other frames.</p>\n<p><strong>CenterNet</strong><br>\nWe took an idea from the original CenterNet and substituted backbone with HRNet. The object detection task in CenterNet architecture is treated as a keypoint estimation, so we used x4 downsampled labels (288x512) for predicting heatmaps of: objects centers, width and height, and center offsets. Final submission includes two backbones: hrnet_w18 and hrnet_w32 that produced similar local performance. The models were trained for 10 epochs and were using cut&amp;paste augmentation.</p>\n<p><strong>FasterRCNN</strong><br>\nWe implemented FasterRCNN to work with any timm backbone and found that NFNet models work particularly well due to generalizability and also easiness to train on small batch sizes. Our final models use eca_nfnet_l1 and dm_nfnet_f0 backbones. We use 4 feature layers and 256 FPN out channels. Models were trained for 10 epochs.</p>\n<p><strong>FCOS</strong><br>\nSimilar to FasterRCNN, we integrated FCOS into our pipeline to work with any timm backbone and use the same backbones for training. Here, we use 3 feature layers and also 256 FPN out channels. Again, models are trained for 10 epochs.</p>\n<p><strong>EfficientDet</strong><br>\nWe train EfficientDet models using efficientdetv2_dt backbone which has a great balance between complexity, memory usage, and runtime. Models are trained for 20 epochs.</p>\n<p><strong>YoloV5</strong><br>\nWe trained yolov5l-6 using rectangular training. We adapted the original code to directly optimize F2 score, and also for rectangular training to properly work with random shuffling, mixup, cutout and other granularities. Models were trained using Adam and for 40 epochs. Inference is done on training resolution without tta.</p>\n<h4>Blending &amp; Tracking</h4>\n<p>For blending, we employ average WBF blending which is particularly useful here as it automatically incorporates a voting mechanism of boxes to downweight boxes that are not present in all models. </p>\n<p>For tracking, the issue with adding missing boxes using Kalman Filter is quite tricky, as also there it is tough to estimate width and height, and the public kernel solutions do not work well when models get better. So we use tracking in a slightly different way. We use euclidean distance on the center points to find the tracks, but we keep low confidence boxes for finding the tracks. Then, we are setting the confidence of low confidence boxes to the maximum of the confidences of the track. So let’s assume we have a track with confidences 0.8, 0.7, 0.9, 0.15 - in this case we increase the confidence of the 0.15 box to 0.9. The models are quite good in finding the boxes, but sometimes they are just not confident enough, so this is clearly better than estimating the new boxes from the track with Kalman filter or similar estimates. This tracking brings about 0.01 locally.</p>\n<p>Another tracking technique that we explored in the last days of the competition was Optical Flow estimation. We used the OpenCV implementation of goodFeaturesToTrack() and calcOpticalFlowPyrLK() methods. It produced pretty nice estimates of diver/camera movements that allowed to better predict the position of starfish across the frames. Unfortunately, it gave only 0.001 improvement in both CV and LB.</p>\n<h4>Label tightness &amp; Public LB</h4>\n<p>As probably for most people, our public LB was significantly lower than our CV for the first half of this competition. However, we quite quickly noticed that there has to be something different in public LB vs. training data, mostly because people with higher LB has so much worse CV reported than what we saw locally. After a while, the discussion about higher resolution came up, and we just blindly tried it on LB and it immediately pushed individual CenterNet and RCNN models above 0.70. </p>\n<p>So then some funny theories emerged on the forums and obviously we also tried to understand how something like this could happen. After several failed theories, we finally found that higher resolution than training, produces boxes that are more tight (exact) around the COTS. In training, labels are very imprecise, and also inconsistent across frames, it even looked semi-automatic. And the size of the boxes plays a huge role in this evaluation metric setting, as 0.8 IOU threshold has a particularly large impact.</p>\n<p>To check this theory, we just submitted manual pixel adjustment on public LB, i.e. decreasing the size of boxes by fixed pixels, and after tuning this a bit, it even worked better than the upscaling, and was much faster obviously. So our theory emerged that test data is labeled better after collecting the following additional evidence:</p>\n<ul>\n<li>Higher resolution inference produces tighter boxes, it works better on Public LB</li>\n<li>Manual pixel adjustment (making boxes smaller) works better on Public LB</li>\n<li>None of these adjustments works on ANY part in training data</li>\n<li>Visually and manually checking a few training examples, we found that boxes on training are on average 3px too large in all dimensions. This is the EXACT same amount that worked well for us on public LB. </li>\n<li>Hand-labeling a small portion of train, finetuning on it, boosted LB significantly.</li>\n<li>The public LB had sequences from both videos. We checked if the manual adjustment works on both videos, and it did, further strengthening our suspicion that whole test is different.</li>\n</ul>\n<p>So we were very confident that it should hold for whole test, but as we know, it did not. So either we completely misinterpreted some of this evidence, or private test is again labeled differently. In general, we personally feel like that public LB should better represent private LB, and in best case it should be similar to training data. This would lead for competitors to overall focus more on improving models, vs. trying to understand the splits. </p>",
  "messages": [
    {
      "id": 1691293,
      "postDate": "2022-02-15T10:35:42.740Z",
      "content": "<p>Thanks to hosts and all participants for this interesting competition. This is a joint writeup with  <a href=\"https://www.kaggle.com/ybabakhin\" target=\"_blank\">@ybabakhin</a> (big, big thanks for the awesome team work in this competition) from Team Hydrogen.</p>\n<h4>Summary</h4>\n<p>Our solution is a blend of five different object detection model types with tracking.</p>\n<h4>Cross-Validation</h4>\n<p>For cross-validation we used two strategies: one is simple split by video, and the other is making 5-folds based on subsequences. We were mostly looking at subsequence Cross-Validation scores, time-to-time checking the video folds performance. This strategy gave us a good correlation between CV and Public LB scores throughout the competition, however it failed for the Private LB. You can read our thoughts about such a behavior below in this post.</p>\n<h4>Models</h4>\n<p>Our final solution is a blend of five different model types. All our models are trained on a resolution of 1152x2048 and only using images with boxes. For augmentations we employ a mix of random flips, scale, shift, rotate, brightness, contrats, cutout, cutmix.</p>\n<p>An additional augmentation we utilize in most of our models is what we call “background-mix”. Here, we randomly mix an image with boxes, with a random image without boxes. This guarantees us to not overlap / distort any boxes, but generalizes better to unseen backgrounds and also randomizes the intensity of objects. Also some models benefit from using cut&amp;paste augmentation by simply cutting and pasting objects to other frames.</p>\n<p><strong>CenterNet</strong><br>\nWe took an idea from the original CenterNet and substituted backbone with HRNet. The object detection task in CenterNet architecture is treated as a keypoint estimation, so we used x4 downsampled labels (288x512) for predicting heatmaps of: objects centers, width and height, and center offsets. Final submission includes two backbones: hrnet_w18 and hrnet_w32 that produced similar local performance. The models were trained for 10 epochs and were using cut&amp;paste augmentation.</p>\n<p><strong>FasterRCNN</strong><br>\nWe implemented FasterRCNN to work with any timm backbone and found that NFNet models work particularly well due to generalizability and also easiness to train on small batch sizes. Our final models use eca_nfnet_l1 and dm_nfnet_f0 backbones. We use 4 feature layers and 256 FPN out channels. Models were trained for 10 epochs.</p>\n<p><strong>FCOS</strong><br>\nSimilar to FasterRCNN, we integrated FCOS into our pipeline to work with any timm backbone and use the same backbones for training. Here, we use 3 feature layers and also 256 FPN out channels. Again, models are trained for 10 epochs.</p>\n<p><strong>EfficientDet</strong><br>\nWe train EfficientDet models using efficientdetv2_dt backbone which has a great balance between complexity, memory usage, and runtime. Models are trained for 20 epochs.</p>\n<p><strong>YoloV5</strong><br>\nWe trained yolov5l-6 using rectangular training. We adapted the original code to directly optimize F2 score, and also for rectangular training to properly work with random shuffling, mixup, cutout and other granularities. Models were trained using Adam and for 40 epochs. Inference is done on training resolution without tta.</p>\n<h4>Blending &amp; Tracking</h4>\n<p>For blending, we employ average WBF blending which is particularly useful here as it automatically incorporates a voting mechanism of boxes to downweight boxes that are not present in all models. </p>\n<p>For tracking, the issue with adding missing boxes using Kalman Filter is quite tricky, as also there it is tough to estimate width and height, and the public kernel solutions do not work well when models get better. So we use tracking in a slightly different way. We use euclidean distance on the center points to find the tracks, but we keep low confidence boxes for finding the tracks. Then, we are setting the confidence of low confidence boxes to the maximum of the confidences of the track. So let’s assume we have a track with confidences 0.8, 0.7, 0.9, 0.15 - in this case we increase the confidence of the 0.15 box to 0.9. The models are quite good in finding the boxes, but sometimes they are just not confident enough, so this is clearly better than estimating the new boxes from the track with Kalman filter or similar estimates. This tracking brings about 0.01 locally.</p>\n<p>Another tracking technique that we explored in the last days of the competition was Optical Flow estimation. We used the OpenCV implementation of goodFeaturesToTrack() and calcOpticalFlowPyrLK() methods. It produced pretty nice estimates of diver/camera movements that allowed to better predict the position of starfish across the frames. Unfortunately, it gave only 0.001 improvement in both CV and LB.</p>\n<h4>Label tightness &amp; Public LB</h4>\n<p>As probably for most people, our public LB was significantly lower than our CV for the first half of this competition. However, we quite quickly noticed that there has to be something different in public LB vs. training data, mostly because people with higher LB has so much worse CV reported than what we saw locally. After a while, the discussion about higher resolution came up, and we just blindly tried it on LB and it immediately pushed individual CenterNet and RCNN models above 0.70. </p>\n<p>So then some funny theories emerged on the forums and obviously we also tried to understand how something like this could happen. After several failed theories, we finally found that higher resolution than training, produces boxes that are more tight (exact) around the COTS. In training, labels are very imprecise, and also inconsistent across frames, it even looked semi-automatic. And the size of the boxes plays a huge role in this evaluation metric setting, as 0.8 IOU threshold has a particularly large impact.</p>\n<p>To check this theory, we just submitted manual pixel adjustment on public LB, i.e. decreasing the size of boxes by fixed pixels, and after tuning this a bit, it even worked better than the upscaling, and was much faster obviously. So our theory emerged that test data is labeled better after collecting the following additional evidence:</p>\n<ul>\n<li>Higher resolution inference produces tighter boxes, it works better on Public LB</li>\n<li>Manual pixel adjustment (making boxes smaller) works better on Public LB</li>\n<li>None of these adjustments works on ANY part in training data</li>\n<li>Visually and manually checking a few training examples, we found that boxes on training are on average 3px too large in all dimensions. This is the EXACT same amount that worked well for us on public LB. </li>\n<li>Hand-labeling a small portion of train, finetuning on it, boosted LB significantly.</li>\n<li>The public LB had sequences from both videos. We checked if the manual adjustment works on both videos, and it did, further strengthening our suspicion that whole test is different.</li>\n</ul>\n<p>So we were very confident that it should hold for whole test, but as we know, it did not. So either we completely misinterpreted some of this evidence, or private test is again labeled differently. In general, we personally feel like that public LB should better represent private LB, and in best case it should be similar to training data. This would lead for competitors to overall focus more on improving models, vs. trying to understand the splits. </p>",
      "rawMarkdown": "Thanks to hosts and all participants for this interesting competition. This is a joint writeup with  @ybabakhin (big, big thanks for the awesome team work in this competition) from Team Hydrogen.\n\n#### Summary\nOur solution is a blend of five different object detection model types with tracking.\n\n#### Cross-Validation\nFor cross-validation we used two strategies: one is simple split by video, and the other is making 5-folds based on subsequences. We were mostly looking at subsequence Cross-Validation scores, time-to-time checking the video folds performance. This strategy gave us a good correlation between CV and Public LB scores throughout the competition, however it failed for the Private LB. You can read our thoughts about such a behavior below in this post.\n\n#### Models\nOur final solution is a blend of five different model types. All our models are trained on a resolution of 1152x2048 and only using images with boxes. For augmentations we employ a mix of random flips, scale, shift, rotate, brightness, contrats, cutout, cutmix.\n\nAn additional augmentation we utilize in most of our models is what we call “background-mix”. Here, we randomly mix an image with boxes, with a random image without boxes. This guarantees us to not overlap / distort any boxes, but generalizes better to unseen backgrounds and also randomizes the intensity of objects. Also some models benefit from using cut&paste augmentation by simply cutting and pasting objects to other frames.\n\n**CenterNet**\nWe took an idea from the original CenterNet and substituted backbone with HRNet. The object detection task in CenterNet architecture is treated as a keypoint estimation, so we used x4 downsampled labels (288x512) for predicting heatmaps of: objects centers, width and height, and center offsets. Final submission includes two backbones: hrnet_w18 and hrnet_w32 that produced similar local performance. The models were trained for 10 epochs and were using cut&paste augmentation.\n\n**FasterRCNN**\nWe implemented FasterRCNN to work with any timm backbone and found that NFNet models work particularly well due to generalizability and also easiness to train on small batch sizes. Our final models use eca_nfnet_l1 and dm_nfnet_f0 backbones. We use 4 feature layers and 256 FPN out channels. Models were trained for 10 epochs.\n\n**FCOS**\nSimilar to FasterRCNN, we integrated FCOS into our pipeline to work with any timm backbone and use the same backbones for training. Here, we use 3 feature layers and also 256 FPN out channels. Again, models are trained for 10 epochs.\n\n**EfficientDet**\nWe train EfficientDet models using efficientdetv2_dt backbone which has a great balance between complexity, memory usage, and runtime. Models are trained for 20 epochs.\n\n**YoloV5**\nWe trained yolov5l-6 using rectangular training. We adapted the original code to directly optimize F2 score, and also for rectangular training to properly work with random shuffling, mixup, cutout and other granularities. Models were trained using Adam and for 40 epochs. Inference is done on training resolution without tta.\n\n#### Blending & Tracking\nFor blending, we employ average WBF blending which is particularly useful here as it automatically incorporates a voting mechanism of boxes to downweight boxes that are not present in all models. \n\nFor tracking, the issue with adding missing boxes using Kalman Filter is quite tricky, as also there it is tough to estimate width and height, and the public kernel solutions do not work well when models get better. So we use tracking in a slightly different way. We use euclidean distance on the center points to find the tracks, but we keep low confidence boxes for finding the tracks. Then, we are setting the confidence of low confidence boxes to the maximum of the confidences of the track. So let’s assume we have a track with confidences 0.8, 0.7, 0.9, 0.15 - in this case we increase the confidence of the 0.15 box to 0.9. The models are quite good in finding the boxes, but sometimes they are just not confident enough, so this is clearly better than estimating the new boxes from the track with Kalman filter or similar estimates. This tracking brings about 0.01 locally.\n\nAnother tracking technique that we explored in the last days of the competition was Optical Flow estimation. We used the OpenCV implementation of goodFeaturesToTrack() and calcOpticalFlowPyrLK() methods. It produced pretty nice estimates of diver/camera movements that allowed to better predict the position of starfish across the frames. Unfortunately, it gave only 0.001 improvement in both CV and LB.\n\n#### Label tightness & Public LB\nAs probably for most people, our public LB was significantly lower than our CV for the first half of this competition. However, we quite quickly noticed that there has to be something different in public LB vs. training data, mostly because people with higher LB has so much worse CV reported than what we saw locally. After a while, the discussion about higher resolution came up, and we just blindly tried it on LB and it immediately pushed individual CenterNet and RCNN models above 0.70. \n\nSo then some funny theories emerged on the forums and obviously we also tried to understand how something like this could happen. After several failed theories, we finally found that higher resolution than training, produces boxes that are more tight (exact) around the COTS. In training, labels are very imprecise, and also inconsistent across frames, it even looked semi-automatic. And the size of the boxes plays a huge role in this evaluation metric setting, as 0.8 IOU threshold has a particularly large impact.\n\nTo check this theory, we just submitted manual pixel adjustment on public LB, i.e. decreasing the size of boxes by fixed pixels, and after tuning this a bit, it even worked better than the upscaling, and was much faster obviously. So our theory emerged that test data is labeled better after collecting the following additional evidence:\n\n* Higher resolution inference produces tighter boxes, it works better on Public LB\n* Manual pixel adjustment (making boxes smaller) works better on Public LB\n* None of these adjustments works on ANY part in training data\n* Visually and manually checking a few training examples, we found that boxes on training are on average 3px too large in all dimensions. This is the EXACT same amount that worked well for us on public LB. \n* Hand-labeling a small portion of train, finetuning on it, boosted LB significantly.\n* The public LB had sequences from both videos. We checked if the manual adjustment works on both videos, and it did, further strengthening our suspicion that whole test is different.\n\nSo we were very confident that it should hold for whole test, but as we know, it did not. So either we completely misinterpreted some of this evidence, or private test is again labeled differently. In general, we personally feel like that public LB should better represent private LB, and in best case it should be similar to training data. This would lead for competitors to overall focus more on improving models, vs. trying to understand the splits. \n",
      "votes": 130
    },
    {
      "id": 1692051,
      "postDate": "2022-02-15T19:08:05.757Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and <a href=\"https://www.kaggle.com/ybabakhin\" target=\"_blank\">@ybabakhin</a> for a great competition and for holding the shake up!<br>\nVery Impressed by your solution.</p>",
      "rawMarkdown": "Congratulations @philippsinger and @ybabakhin for a great competition and for holding the shake up!\nVery Impressed by your solution.",
      "votes": 5
    },
    {
      "id": 1691379,
      "postDate": "2022-02-15T11:34:35.093Z",
      "content": "<blockquote>\n  <p>in general, we personally feel like that public LB should better represent private LB, and in best case it should be similar to training data.</p>\n</blockquote>\n<p>I very much agree with this, after first LB commit I spent 2 of three total weeks fighting LB windmills. That didnt contribute to anyone</p>",
      "rawMarkdown": "> in general, we personally feel like that public LB should better represent private LB, and in best case it should be similar to training data.\n\nI very much agree with this, after first LB commit I spent 2 of three total weeks fighting LB windmills. That didnt contribute to anyone",
      "votes": 4
    },
    {
      "id": 1692426,
      "postDate": "2022-02-16T03:12:37.947Z",
      "content": "<p>Congratulations!<br>\nAmazing that you guys use so many different object detection models</p>",
      "rawMarkdown": "Congratulations!\nAmazing that you guys use so many different object detection models",
      "votes": 1
    },
    {
      "id": 1691423,
      "postDate": "2022-02-15T11:59:00.070Z",
      "content": "<p>Congratulations and thanks for sharing! Did you make any modifications to EfficientDet to get it to work effectively? I couldn't get it to work at all for this task…</p>\n<blockquote>\n  <p>In training, labels are very imprecise, and also inconsistent across frames, it even looked semi-automatic.</p>\n</blockquote>\n<p>I believe in the short paper the hosts shared it said that the boxes were first generated from an object detection then manually refined/verified.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! Did you make any modifications to EfficientDet to get it to work effectively? I couldn't get it to work at all for this task...\n\n> In training, labels are very imprecise, and also inconsistent across frames, it even looked semi-automatic.\n\nI believe in the short paper the hosts shared it said that the boxes were first generated from an object detection then manually refined/verified.",
      "votes": 1,
      "replies": [
        {
          "id": 1691435,
          "postDate": "2022-02-15T12:09:24.520Z",
          "content": "<p>We did not make any changes to EfficientDet hyperparameters themselve, even though we tried a lot changing anchors, anchor scales etc. What did work very well there was to increase augmentations quite significantly and adding this background mixing. And then training a bit longer. Also the efficientdetv2 backbones worked very nicely.</p>",
          "rawMarkdown": "We did not make any changes to EfficientDet hyperparameters themselve, even though we tried a lot changing anchors, anchor scales etc. What did work very well there was to increase augmentations quite significantly and adding this background mixing. And then training a bit longer. Also the efficientdetv2 backbones worked very nicely.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1691328,
      "postDate": "2022-02-15T10:52:30.140Z",
      "content": "<p>Congratulations! Outstanding work and great solution. Thank you for sharing. I really like your way of thinking (system thinking). As I can see you during the competition you asked yourself a lot of \"why\" questions? And found answer. Great!</p>",
      "rawMarkdown": "Congratulations! Outstanding work and great solution. Thank you for sharing. I really like your way of thinking (system thinking). As I can see you during the competition you asked yourself a lot of \"why\" questions? And found answer. Great!",
      "votes": 2
    },
    {
      "id": 1691311,
      "postDate": "2022-02-15T10:46:45.797Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and Team Hydrogen! I’ve been looking out for you guys ever since you were the first to achieve 0.7 LB, so congrats on the Gold Medal and the prize winning!</p>",
      "rawMarkdown": "Congratulations @philippsinger and Team Hydrogen! I’ve been looking out for you guys ever since you were the first to achieve 0.7 LB, so congrats on the Gold Medal and the prize winning!",
      "votes": 2
    },
    {
      "id": 1695370,
      "postDate": "2022-02-18T05:26:05.170Z",
      "content": "<p>Congratulations, amazing work and solution. Thanks for sharing 🙌</p>",
      "rawMarkdown": "Congratulations, amazing work and solution. Thanks for sharing 🙌"
    },
    {
      "id": 1693575,
      "postDate": "2022-02-16T18:56:16.557Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 1693469,
      "postDate": "2022-02-16T17:32:44.483Z",
      "content": "<p>congratulation</p>",
      "rawMarkdown": "congratulation"
    },
    {
      "id": 1693233,
      "postDate": "2022-02-16T14:37:01.713Z",
      "content": "<p>Hi and congratulations!</p>\n<p>Just wondering if you will be open-sourcing any of the models for this competition.</p>\n<p>Also, I'm wondering if you used Pytorch or TF or a mix?</p>",
      "rawMarkdown": "Hi and congratulations!\n\nJust wondering if you will be open-sourcing any of the models for this competition.\n\nAlso, I'm wondering if you used Pytorch or TF or a mix?",
      "replies": [
        {
          "id": 1693295,
          "postDate": "2022-02-16T15:14:45.233Z",
          "content": "<p>Pure Pytorch, probably will not open source it.</p>",
          "rawMarkdown": "Pure Pytorch, probably will not open source it.",
          "votes": 1
        },
        {
          "id": 1693321,
          "postDate": "2022-02-16T15:40:50.837Z",
          "content": "<p>Cool! Thanks for the response and congratulations again!</p>",
          "rawMarkdown": "Cool! Thanks for the response and congratulations again!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1692044,
      "postDate": "2022-02-15T18:59:25.363Z",
      "content": "<p>Congratulations on the 3rd place ranking! Impressive solution. </p>\n<p><code>Visually and manually checking a few training examples, we found that boxes on training are on average 3px too large in all dimensions. This is the EXACT same amount that worked well for us on public LB.\n</code><br>\nI am curious, did you try training on tighter boxes by adjusting the GT boxes? The hope would be that model performs similarly in CV and LB. </p>",
      "rawMarkdown": "Congratulations on the 3rd place ranking! Impressive solution. \n\n`Visually and manually checking a few training examples, we found that boxes on training are on average 3px too large in all dimensions. This is the EXACT same amount that worked well for us on public LB.\n`\nI am curious, did you try training on tighter boxes by adjusting the GT boxes? The hope would be that model performs similarly in CV and LB. ",
      "replies": [
        {
          "id": 1692047,
          "postDate": "2022-02-15T19:03:04.257Z",
          "content": "<p>Yes, we did, and it was our big hope for private. It scored high for public, but not private. It can reach ~0.8+ CV, 0.79 LB, but only scores ~0.69 on private, meaning that private is much more inprecise again.</p>",
          "rawMarkdown": "Yes, we did, and it was our big hope for private. It scored high for public, but not private. It can reach ~0.8+ CV, 0.79 LB, but only scores ~0.69 on private, meaning that private is much more inprecise again.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1691716,
      "postDate": "2022-02-15T15:09:38.033Z",
      "content": "<p>Congratulation and Thanks for the solution.</p>",
      "rawMarkdown": "Congratulation and Thanks for the solution."
    },
    {
      "id": 1691661,
      "postDate": "2022-02-15T14:38:19.233Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and <a href=\"https://www.kaggle.com/ybabakhin\" target=\"_blank\">@ybabakhin</a>!</p>\n<p>Just wondering, do you guys plan on releasing the source code in the future?</p>",
      "rawMarkdown": "Congratulations @philippsinger and @ybabakhin!\n\nJust wondering, do you guys plan on releasing the source code in the future?"
    },
    {
      "id": 1691430,
      "postDate": "2022-02-15T12:07:29.390Z",
      "content": "<p>Congratulations, could you give us the GitHub link from where you used the Centernet.<br>\nyour investigative work is awesome, but yes if the LB and training data were similar you guys could have made some better models. The same thing is happening with the new Happy whale comp.</p>",
      "rawMarkdown": "Congratulations, could you give us the GitHub link from where you used the Centernet.\nyour investigative work is awesome, but yes if the LB and training data were similar you guys could have made some better models. The same thing is happening with the new Happy whale comp.\n",
      "replies": [
        {
          "id": 1691437,
          "postDate": "2022-02-15T12:10:01.833Z",
          "content": "<p>BTW, on average how many hours did you dedicate to the comp daily, looking at your work I think it was close to 10-15 hours ;)</p>",
          "rawMarkdown": "\nBTW, on average how many hours did you dedicate to the comp daily, looking at your work I think it was close to 10-15 hours ;)"
        },
        {
          "id": 1691438,
          "postDate": "2022-02-15T12:10:05.127Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1691459,
          "postDate": "2022-02-15T12:23:11.120Z",
          "content": "<p>We implemented pipelines for all model types ourselves. <br>\nAnd yeah, spent quite some time on it :)</p>",
          "rawMarkdown": "We implemented pipelines for all model types ourselves. \nAnd yeah, spent quite some time on it :)"
        },
        {
          "id": 1691492,
          "postDate": "2022-02-15T12:44:08.233Z",
          "content": "<p><code>time-to-time checking the video folds performance.</code><br>\ndoes this mean you don't track your validation scores while training and see the CV scores after a few epochs trained?<br>\nif not, what all metrics are you tracking while training? as I found it difficult to code to find the CV in between the epochs and track</p>",
          "rawMarkdown": "`time-to-time checking the video folds performance.`\ndoes this mean you don't track your validation scores while training and see the CV scores after a few epochs trained?\nif not, what all metrics are you tracking while training? as I found it difficult to code to find the CV in between the epochs and track"
        },
        {
          "id": 1691500,
          "postDate": "2022-02-15T12:47:13.900Z",
          "content": "<p>No, what we mean with that is that we track our subsequence CV, and whenever we get clearly better models, we also check them on video CV setup to make sure they generalize across videos. Implementing the metric and tracking it at every step (even across thresholds) possible for all models is the first thing we always add to the pipeline.</p>",
          "rawMarkdown": "No, what we mean with that is that we track our subsequence CV, and whenever we get clearly better models, we also check them on video CV setup to make sure they generalize across videos. Implementing the metric and tracking it at every step (even across thresholds) possible for all models is the first thing we always add to the pipeline.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1691410,
      "postDate": "2022-02-15T11:55:06.597Z",
      "content": "<p>“Box tightness could affect public LB” — nice observations. Thank you for sharing this.</p>\n<p>I also relabeled the boxes to improve missing labels, and as a side effect, the newly created boxes happened to be more tighter ones, it might affected to pushes me up the private LB. However, for me, training with relabeled data didn’t much contributed on public LB.</p>",
      "rawMarkdown": "“Box tightness could affect public LB” — nice observations. Thank you for sharing this.\n\nI also relabeled the boxes to improve missing labels, and as a side effect, the newly created boxes happened to be more tighter ones, it might affected to pushes me up the private LB. However, for me, training with relabeled data didn’t much contributed on public LB.",
      "replies": [
        {
          "id": 1691462,
          "postDate": "2022-02-15T12:23:43.530Z",
          "content": "<p>Did you still infer on larger resolution after training on relabeled data? It might be the reason why it didnt contribute.</p>",
          "rawMarkdown": "Did you still infer on larger resolution after training on relabeled data? It might be the reason why it didnt contribute."
        },
        {
          "id": 1692517,
          "postDate": "2022-02-16T04:59:24.337Z",
          "content": "<blockquote>\n  <p>Did you still infer on larger resolution after training on relabeled data?</p>\n</blockquote>\n<p>No, the final submission is infer size(x2), which is the same as the train scale. However, the score for this model is only 0.660.</p>\n<ul>\n<li>Public LB: 0.660</li>\n<li>Private LB: 0.723</li>\n</ul>\n<p>The size of the box was just seen in the image, no statistics were taken (the average size may be about the same). Or it could be some other reason that pushed up the LB.</p>",
          "rawMarkdown": "> Did you still infer on larger resolution after training on relabeled data?\n\nNo, the final submission is infer size(x2), which is the same as the train scale. However, the score for this model is only 0.660.\n\n* Public LB: 0.660\n* Private LB: 0.723\n\nThe size of the box was just seen in the image, no statistics were taken (the average size may be about the same). Or it could be some other reason that pushed up the LB."
        },
        {
          "id": 1693007,
          "postDate": "2022-02-16T11:49:12.470Z",
          "content": "<p>yeah, private lb seems to be also quite random</p>",
          "rawMarkdown": "yeah, private lb seems to be also quite random"
        }
      ]
    },
    {
      "id": 1691370,
      "postDate": "2022-02-15T11:27:26.647Z",
      "content": "<p>Congratulations and thank you for sharing!</p>",
      "rawMarkdown": "Congratulations and thank you for sharing!"
    },
    {
      "id": 1691355,
      "postDate": "2022-02-15T11:20:30.490Z",
      "content": "<p>Wow! great solution, congratulations for podium! :)</p>",
      "rawMarkdown": "Wow! great solution, congratulations for podium! :)"
    },
    {
      "id": 1691346,
      "postDate": "2022-02-15T11:07:30.890Z",
      "content": "<p>Thanks for the write-up and congratulations.</p>\n<blockquote>\n  <p>we found that boxes on training are on average 3px too large in all dimensions</p>\n</blockquote>\n<p>How did you jump to the conclusion that the bboxes were on average 3px too large ? Was this by done by manually visualizing the boxes ?</p>",
      "rawMarkdown": "Thanks for the write-up and congratulations.\n> we found that boxes on training are on average 3px too large in all dimensions\n\nHow did you jump to the conclusion that the bboxes were on average 3px too large ? Was this by done by manually visualizing the boxes ?",
      "replies": [
        {
          "id": 1691347,
          "postDate": "2022-02-15T11:11:33.713Z",
          "content": "<p>Manually looking and adjusting a subset of frames/boxes to properly fit the actual shape of the COTS and then calculating adjustments from original labels.</p>",
          "rawMarkdown": "Manually looking and adjusting a subset of frames/boxes to properly fit the actual shape of the COTS and then calculating adjustments from original labels.",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1692051,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2022-02-15T19:08:05.757000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and <a href=\"https://www.kaggle.com/ybabakhin\" target=\"_blank\">@ybabakhin</a> for a great competition and for holding the shake up!<br>\nVery Impressed by your solution.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1691379,
      "author_name": "Gleb",
      "author_url": "",
      "post_date": "2022-02-15T11:34:35.093000",
      "content": "<blockquote>\n  <p>in general, we personally feel like that public LB should better represent private LB, and in best case it should be similar to training data.</p>\n</blockquote>\n<p>I very much agree with this, after first LB commit I spent 2 of three total weeks fighting LB windmills. That didnt contribute to anyone</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1692426,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2022-02-16T03:12:37.947000",
      "content": "<p>Congratulations!<br>\nAmazing that you guys use so many different object detection models</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1691423,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2022-02-15T11:59:00.070000",
      "content": "<p>Congratulations and thanks for sharing! Did you make any modifications to EfficientDet to get it to work effectively? I couldn't get it to work at all for this task…</p>\n<blockquote>\n  <p>In training, labels are very imprecise, and also inconsistent across frames, it even looked semi-automatic.</p>\n</blockquote>\n<p>I believe in the short paper the hosts shared it said that the boxes were first generated from an object detection then manually refined/verified.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1691435,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-02-15T12:09:24.520000",
          "content": "<p>We did not make any changes to EfficientDet hyperparameters themselve, even though we tried a lot changing anchors, anchor scales etc. What did work very well there was to increase augmentations quite significantly and adding this background mixing. And then training a bit longer. Also the efficientdetv2 backbones worked very nicely.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1691328,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-02-15T10:52:30.140000",
      "content": "<p>Congratulations! Outstanding work and great solution. Thank you for sharing. I really like your way of thinking (system thinking). As I can see you during the competition you asked yourself a lot of \"why\" questions? And found answer. Great!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1691311,
      "author_name": "Coderrexe",
      "author_url": "",
      "post_date": "2022-02-15T10:46:45.797000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and Team Hydrogen! I’ve been looking out for you guys ever since you were the first to achieve 0.7 LB, so congrats on the Gold Medal and the prize winning!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1695370,
      "author_name": "Niek van der Zwaag",
      "author_url": "",
      "post_date": "2022-02-18T05:26:05.170000",
      "content": "<p>Congratulations, amazing work and solution. Thanks for sharing 🙌</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1693575,
      "author_name": "Philip Joseph",
      "author_url": "",
      "post_date": "2022-02-16T18:56:16.557000",
      "content": "<p>Congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1693469,
      "author_name": "Abdur Rakib",
      "author_url": "",
      "post_date": "2022-02-16T17:32:44.483000",
      "content": "<p>congratulation</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1693233,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2022-02-16T14:37:01.713000",
      "content": "<p>Hi and congratulations!</p>\n<p>Just wondering if you will be open-sourcing any of the models for this competition.</p>\n<p>Also, I'm wondering if you used Pytorch or TF or a mix?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1693295,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-02-16T15:14:45.233000",
          "content": "<p>Pure Pytorch, probably will not open source it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1693321,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2022-02-16T15:40:50.837000",
          "content": "<p>Cool! Thanks for the response and congratulations again!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1692044,
      "author_name": "Trushant Kalyanpur",
      "author_url": "",
      "post_date": "2022-02-15T18:59:25.363000",
      "content": "<p>Congratulations on the 3rd place ranking! Impressive solution. </p>\n<p><code>Visually and manually checking a few training examples, we found that boxes on training are on average 3px too large in all dimensions. This is the EXACT same amount that worked well for us on public LB.\n</code><br>\nI am curious, did you try training on tighter boxes by adjusting the GT boxes? The hope would be that model performs similarly in CV and LB. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1692047,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-02-15T19:03:04.257000",
          "content": "<p>Yes, we did, and it was our big hope for private. It scored high for public, but not private. It can reach ~0.8+ CV, 0.79 LB, but only scores ~0.69 on private, meaning that private is much more inprecise again.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1691716,
      "author_name": "Gaju Ahmed",
      "author_url": "",
      "post_date": "2022-02-15T15:09:38.033000",
      "content": "<p>Congratulation and Thanks for the solution.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1691661,
      "author_name": "kagglemaniak",
      "author_url": "",
      "post_date": "2022-02-15T14:38:19.233000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> and <a href=\"https://www.kaggle.com/ybabakhin\" target=\"_blank\">@ybabakhin</a>!</p>\n<p>Just wondering, do you guys plan on releasing the source code in the future?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1691430,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2022-02-15T12:07:29.390000",
      "content": "<p>Congratulations, could you give us the GitHub link from where you used the Centernet.<br>\nyour investigative work is awesome, but yes if the LB and training data were similar you guys could have made some better models. The same thing is happening with the new Happy whale comp.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1691437,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2022-02-15T12:10:01.833000",
          "content": "<p>BTW, on average how many hours did you dedicate to the comp daily, looking at your work I think it was close to 10-15 hours ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1691438,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-15T12:10:05.127000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1691459,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-02-15T12:23:11.120000",
          "content": "<p>We implemented pipelines for all model types ourselves. <br>\nAnd yeah, spent quite some time on it :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1691492,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2022-02-15T12:44:08.233000",
          "content": "<p><code>time-to-time checking the video folds performance.</code><br>\ndoes this mean you don't track your validation scores while training and see the CV scores after a few epochs trained?<br>\nif not, what all metrics are you tracking while training? as I found it difficult to code to find the CV in between the epochs and track</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1691500,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-02-15T12:47:13.900000",
          "content": "<p>No, what we mean with that is that we track our subsequence CV, and whenever we get clearly better models, we also check them on video CV setup to make sure they generalize across videos. Implementing the metric and tracking it at every step (even across thresholds) possible for all models is the first thing we always add to the pipeline.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1691410,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-02-15T11:55:06.597000",
      "content": "<p>“Box tightness could affect public LB” — nice observations. Thank you for sharing this.</p>\n<p>I also relabeled the boxes to improve missing labels, and as a side effect, the newly created boxes happened to be more tighter ones, it might affected to pushes me up the private LB. However, for me, training with relabeled data didn’t much contributed on public LB.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1691462,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-02-15T12:23:43.530000",
          "content": "<p>Did you still infer on larger resolution after training on relabeled data? It might be the reason why it didnt contribute.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1692517,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-16T04:59:24.337000",
          "content": "<blockquote>\n  <p>Did you still infer on larger resolution after training on relabeled data?</p>\n</blockquote>\n<p>No, the final submission is infer size(x2), which is the same as the train scale. However, the score for this model is only 0.660.</p>\n<ul>\n<li>Public LB: 0.660</li>\n<li>Private LB: 0.723</li>\n</ul>\n<p>The size of the box was just seen in the image, no statistics were taken (the average size may be about the same). Or it could be some other reason that pushed up the LB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1693007,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-02-16T11:49:12.470000",
          "content": "<p>yeah, private lb seems to be also quite random</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1691370,
      "author_name": "Vivo Vinco",
      "author_url": "",
      "post_date": "2022-02-15T11:27:26.647000",
      "content": "<p>Congratulations and thank you for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1691355,
      "author_name": "Lukasz Borecki",
      "author_url": "",
      "post_date": "2022-02-15T11:20:30.490000",
      "content": "<p>Wow! great solution, congratulations for podium! :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1691346,
      "author_name": "Liam Nguyen",
      "author_url": "",
      "post_date": "2022-02-15T11:07:30.890000",
      "content": "<p>Thanks for the write-up and congratulations.</p>\n<blockquote>\n  <p>we found that boxes on training are on average 3px too large in all dimensions</p>\n</blockquote>\n<p>How did you jump to the conclusion that the bboxes were on average 3px too large ? Was this by done by manually visualizing the boxes ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1691347,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-02-15T11:11:33.713000",
          "content": "<p>Manually looking and adjusting a subset of frames/boxes to properly fit the actual shape of the COTS and then calculating adjustments from original labels.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1691293": "Thanks to hosts and all participants for this interesting competition. This is a joint writeup with  @ybabakhin (big, big thanks for the awesome team work in this competition) from Team Hydrogen.\n\n#### Summary\nOur solution is a blend of five different object detection model types with tracking.\n\n#### Cross-Validation\nFor cross-validation we used two strategies: one is simple split by video, and the other is making 5-folds based on subsequences. We were mostly looking at subsequence Cross-Validation scores, time-to-time checking the video folds performance. This strategy gave us a good correlation between CV and Public LB scores throughout the competition, however it failed for the Private LB. You can read our thoughts about such a behavior below in this post.\n\n#### Models\nOur final solution is a blend of five different model types. All our models are trained on a resolution of 1152x2048 and only using images with boxes. For augmentations we employ a mix of random flips, scale, shift, rotate, brightness, contrats, cutout, cutmix.\n\nAn additional augmentation we utilize in most of our models is what we call “background-mix”. Here, we randomly mix an image with boxes, with a random image without boxes. This guarantees us to not overlap / distort any boxes, but generalizes better to unseen backgrounds and also randomizes the intensity of objects. Also some models benefit from using cut&paste augmentation by simply cutting and pasting objects to other frames.\n\n**CenterNet**\nWe took an idea from the original CenterNet and substituted backbone with HRNet. The object detection task in CenterNet architecture is treated as a keypoint estimation, so we used x4 downsampled labels (288x512) for predicting heatmaps of: objects centers, width and height, and center offsets. Final submission includes two backbones: hrnet_w18 and hrnet_w32 that produced similar local performance. The models were trained for 10 epochs and were using cut&paste augmentation.\n\n**FasterRCNN**\nWe implemented FasterRCNN to work with any timm backbone and found that NFNet models work particularly well due to generalizability and also easiness to train on small batch sizes. Our final models use eca_nfnet_l1 and dm_nfnet_f0 backbones. We use 4 feature layers and 256 FPN out channels. Models were trained for 10 epochs.\n\n**FCOS**\nSimilar to FasterRCNN, we integrated FCOS into our pipeline to work with any timm backbone and use the same backbones for training. Here, we use 3 feature layers and also 256 FPN out channels. Again, models are trained for 10 epochs.\n\n**EfficientDet**\nWe train EfficientDet models using efficientdetv2_dt backbone which has a great balance between complexity, memory usage, and runtime. Models are trained for 20 epochs.\n\n**YoloV5**\nWe trained yolov5l-6 using rectangular training. We adapted the original code to directly optimize F2 score, and also for rectangular training to properly work with random shuffling, mixup, cutout and other granularities. Models were trained using Adam and for 40 epochs. Inference is done on training resolution without tta.\n\n#### Blending & Tracking\nFor blending, we employ average WBF blending which is particularly useful here as it automatically incorporates a voting mechanism of boxes to downweight boxes that are not present in all models. \n\nFor tracking, the issue with adding missing boxes using Kalman Filter is quite tricky, as also there it is tough to estimate width and height, and the public kernel solutions do not work well when models get better. So we use tracking in a slightly different way. We use euclidean distance on the center points to find the tracks, but we keep low confidence boxes for finding the tracks. Then, we are setting the confidence of low confidence boxes to the maximum of the confidences of the track. So let’s assume we have a track with confidences 0.8, 0.7, 0.9, 0.15 - in this case we increase the confidence of the 0.15 box to 0.9. The models are quite good in finding the boxes, but sometimes they are just not confident enough, so this is clearly better than estimating the new boxes from the track with Kalman filter or similar estimates. This tracking brings about 0.01 locally.\n\nAnother tracking technique that we explored in the last days of the competition was Optical Flow estimation. We used the OpenCV implementation of goodFeaturesToTrack() and calcOpticalFlowPyrLK() methods. It produced pretty nice estimates of diver/camera movements that allowed to better predict the position of starfish across the frames. Unfortunately, it gave only 0.001 improvement in both CV and LB.\n\n#### Label tightness & Public LB\nAs probably for most people, our public LB was significantly lower than our CV for the first half of this competition. However, we quite quickly noticed that there has to be something different in public LB vs. training data, mostly because people with higher LB has so much worse CV reported than what we saw locally. After a while, the discussion about higher resolution came up, and we just blindly tried it on LB and it immediately pushed individual CenterNet and RCNN models above 0.70. \n\nSo then some funny theories emerged on the forums and obviously we also tried to understand how something like this could happen. After several failed theories, we finally found that higher resolution than training, produces boxes that are more tight (exact) around the COTS. In training, labels are very imprecise, and also inconsistent across frames, it even looked semi-automatic. And the size of the boxes plays a huge role in this evaluation metric setting, as 0.8 IOU threshold has a particularly large impact.\n\nTo check this theory, we just submitted manual pixel adjustment on public LB, i.e. decreasing the size of boxes by fixed pixels, and after tuning this a bit, it even worked better than the upscaling, and was much faster obviously. So our theory emerged that test data is labeled better after collecting the following additional evidence:\n\n* Higher resolution inference produces tighter boxes, it works better on Public LB\n* Manual pixel adjustment (making boxes smaller) works better on Public LB\n* None of these adjustments works on ANY part in training data\n* Visually and manually checking a few training examples, we found that boxes on training are on average 3px too large in all dimensions. This is the EXACT same amount that worked well for us on public LB. \n* Hand-labeling a small portion of train, finetuning on it, boosted LB significantly.\n* The public LB had sequences from both videos. We checked if the manual adjustment works on both videos, and it did, further strengthening our suspicion that whole test is different.\n\nSo we were very confident that it should hold for whole test, but as we know, it did not. So either we completely misinterpreted some of this evidence, or private test is again labeled differently. In general, we personally feel like that public LB should better represent private LB, and in best case it should be similar to training data. This would lead for competitors to overall focus more on improving models, vs. trying to understand the splits. \n",
    "1692051": "Congratulations @philippsinger and @ybabakhin for a great competition and for holding the shake up!\nVery Impressed by your solution.",
    "1691379": "> in general, we personally feel like that public LB should better represent private LB, and in best case it should be similar to training data.\n\nI very much agree with this, after first LB commit I spent 2 of three total weeks fighting LB windmills. That didnt contribute to anyone",
    "1692426": "Congratulations!\nAmazing that you guys use so many different object detection models",
    "1691423": "Congratulations and thanks for sharing! Did you make any modifications to EfficientDet to get it to work effectively? I couldn't get it to work at all for this task...\n\n> In training, labels are very imprecise, and also inconsistent across frames, it even looked semi-automatic.\n\nI believe in the short paper the hosts shared it said that the boxes were first generated from an object detection then manually refined/verified.",
    "1691328": "Congratulations! Outstanding work and great solution. Thank you for sharing. I really like your way of thinking (system thinking). As I can see you during the competition you asked yourself a lot of \"why\" questions? And found answer. Great!",
    "1691311": "Congratulations @philippsinger and Team Hydrogen! I’ve been looking out for you guys ever since you were the first to achieve 0.7 LB, so congrats on the Gold Medal and the prize winning!",
    "1695370": "Congratulations, amazing work and solution. Thanks for sharing 🙌",
    "1693575": "Congratulations!",
    "1693469": "congratulation",
    "1693233": "Hi and congratulations!\n\nJust wondering if you will be open-sourcing any of the models for this competition.\n\nAlso, I'm wondering if you used Pytorch or TF or a mix?",
    "1692044": "Congratulations on the 3rd place ranking! Impressive solution. \n\n`Visually and manually checking a few training examples, we found that boxes on training are on average 3px too large in all dimensions. This is the EXACT same amount that worked well for us on public LB.\n`\nI am curious, did you try training on tighter boxes by adjusting the GT boxes? The hope would be that model performs similarly in CV and LB. ",
    "1691716": "Congratulation and Thanks for the solution.",
    "1691661": "Congratulations @philippsinger and @ybabakhin!\n\nJust wondering, do you guys plan on releasing the source code in the future?",
    "1691430": "Congratulations, could you give us the GitHub link from where you used the Centernet.\nyour investigative work is awesome, but yes if the LB and training data were similar you guys could have made some better models. The same thing is happening with the new Happy whale comp.\n",
    "1691410": "“Box tightness could affect public LB” — nice observations. Thank you for sharing this.\n\nI also relabeled the boxes to improve missing labels, and as a side effect, the newly created boxes happened to be more tighter ones, it might affected to pushes me up the private LB. However, for me, training with relabeled data didn’t much contributed on public LB.",
    "1691370": "Congratulations and thank you for sharing!",
    "1691355": "Wow! great solution, congratulations for podium! :)",
    "1691346": "Thanks for the write-up and congratulations.\n> we found that boxes on training are on average 3px too large in all dimensions\n\nHow did you jump to the conclusion that the bboxes were on average 3px too large ? Was this by done by manually visualizing the boxes ?"
  }
}